Finding and Choosing Models
On this page
llamadart supports two model artifact families:
- GGUF (GGML Unified Format), the standard format used by
llama.cpp. .litertlmbundles, used by LiteRT-LM.
You cannot use raw PyTorch (.bin or .safetensors) models directly; they
must be converted into a runtime-ready artifact first. For most open model
workflows that means a quantized GGUF. For LiteRT-LM deployments, use a
published .litertlm bundle that matches the LiteRT-LM runtime.
Fortunately, thousands of pre-converted GGUF models are readily available. LiteRT-LM bundles are more specialized; use them when your target model is distributed for LiteRT-LM or when you are intentionally benchmarking that runtime path.
Where to find models#
The best place to find GGUF models is Hugging Face.
You can search for any model name followed by gguf (e.g., Llama-3-8B-Instruct-GGUF).
For a deeper dive into the GGUF ecosystem and how quantization works, check out these insightful resources:
For LiteRT-LM, look for model repositories that publish .litertlm files and
document LiteRT-LM compatibility. If both GGUF and .litertlm variants exist,
see Choosing llama.cpp or LiteRT-LM
before
treating them as equivalent artifacts.
Understanding Quantization#
GGUF models usually come in different "quantization" levels, denoted by tags like Q4_K_M
or Q8_0. Quantization reduces the precision of the model's weights to save memory and increase inference speed, at a slight cost to "smartness" (perplexity).
Here is a quick guide to choosing a quantization level:
-
Q4_K_M: The recommended baseline. It offers an excellent balance between small file size, fast generation, and retaining the model's original quality. -
Q5_K_M: Slightly larger and slower than Q4, but retains more quality. Good if you have the RAM to spare. -
Q8_0: Almost indistinguishable from the unquantized raw model, but requires double the RAM of Q4. -
Q2_K/Q3_K: Highly compressed. Useful only if you are severely constrained by RAM (e.g., running on older mobile phones), but expect noticeable degradation in reasoning logic.
Recommended constraints per platform#
When downloading a model, check its file size. Your target device needs enough free RAM (or VRAM for GPU offloading) to load the model, plus a bit extra for the context window.
| Platform | Recommended Model Parameter Size | Target RAM Usage |
|---|---|---|
| Mobile (iOS/Android) | 1B - 3B parameters | 1GB - 2.5GB (e.g., Llama-3.2-1B Q4_K_M) |
| Old Laptop/Desktop | 3B - 8B parameters | 2.5GB - 6GB (e.g., Llama-3.1-8B Q4_K_M) |
| Modern Mac (M1/M2/M3) | 8B - 32B parameters | 6GB - 20GB+ |
Downloading a model#
Once you find a model on Hugging Face:
- Go to the Files and versions tab of the model repository.
-
Look for a file ending in
.gguf(e.g.,model-q4_k_m.gguf) or.litertlm. -
Use the exact repository path with
ModelSource.parse('hf://owner/repo/path/to/model.gguf'), or click the download icon and place the model file in your application's assets or a reachable file path forengine.loadModel(). Use the real file extension in thehf://path soLlamaBackend()can route to the correct runtime.
For package-managed downloads, hf:// defaults to the repository's main
revision. Use hf://owner/repo@tag/model.gguf for simple branch/tag names, or
hf://owner/repo/model.gguf?revision=refs/pr/12 when the revision contains
/.
Private or gated repositories need ModelLoadOptions(bearerToken: hfToken)
(or
custom headers); do not put tokens in source strings. Multimodal repos often
ship a separate mmproj GGUF file—treat it as a separate asset/source. Sharded
GGUF repos are not expanded automatically by llamadart; choose a single-file
GGUF unless you are handling shards yourself.