Documentation for v0.8.15, an older release. Read v0.9.0, the latest release

Finding and Choosing Models

On this page

llamadart supports two model artifact families:

  • GGUF (GGML Unified Format), the standard format used by llama.cpp.
  • .litertlm bundles, used by LiteRT-LM.

You cannot use raw PyTorch (.bin or .safetensors) models directly; they must be converted into a runtime-ready artifact first. For most open model workflows that means a quantized GGUF. For LiteRT-LM deployments, use a published .litertlm bundle that matches the LiteRT-LM runtime.

Fortunately, thousands of pre-converted GGUF models are readily available. LiteRT-LM bundles are more specialized; use them when your target model is distributed for LiteRT-LM or when you are intentionally benchmarking that runtime path.

Where to find models#

The best place to find GGUF models is Hugging Face.

You can search for any model name followed by gguf (e.g., Llama-3-8B-Instruct-GGUF). For a deeper dive into the GGUF ecosystem and how quantization works, check out these insightful resources:

For LiteRT-LM, look for model repositories that publish .litertlm files and document LiteRT-LM compatibility. If both GGUF and .litertlm variants exist, see Choosing llama.cpp or LiteRT-LM before treating them as equivalent artifacts.

Understanding Quantization#

GGUF models usually come in different "quantization" levels, denoted by tags like Q4_K_M or Q8_0. Quantization reduces the precision of the model's weights to save memory and increase inference speed, at a slight cost to "smartness" (perplexity).

Here is a quick guide to choosing a quantization level:

  • Q4_K_M: The recommended baseline. It offers an excellent balance between small file size, fast generation, and retaining the model's original quality.
  • Q5_K_M: Slightly larger and slower than Q4, but retains more quality. Good if you have the RAM to spare.
  • Q8_0: Almost indistinguishable from the unquantized raw model, but requires double the RAM of Q4.
  • Q2_K / Q3_K: Highly compressed. Useful only if you are severely constrained by RAM (e.g., running on older mobile phones), but expect noticeable degradation in reasoning logic.

When downloading a model, check its file size. Your target device needs enough free RAM (or VRAM for GPU offloading) to load the model, plus a bit extra for the context window.

PlatformRecommended Model Parameter SizeTarget RAM Usage
Mobile (iOS/Android) 1B - 3B parameters 1GB - 2.5GB (e.g., Llama-3.2-1B Q4_K_M)
Old Laptop/Desktop 3B - 8B parameters 2.5GB - 6GB (e.g., Llama-3.1-8B Q4_K_M)
Modern Mac (M1/M2/M3)8B - 32B parameters6GB - 20GB+

Downloading a model#

Once you find a model on Hugging Face:

  1. Go to the Files and versions tab of the model repository.
  2. Look for a file ending in .gguf (e.g., model-q4_k_m.gguf) or .litertlm.
  3. Use the exact repository path with ModelSource.parse('hf://owner/repo/path/to/model.gguf'), or click the download icon and place the model file in your application's assets or a reachable file path for engine.loadModel(). Use the real file extension in the hf:// path so LlamaBackend() can route to the correct runtime.

For package-managed downloads, hf:// defaults to the repository's main revision. Use hf://owner/repo@tag/model.gguf for simple branch/tag names, or hf://owner/repo/model.gguf?revision=refs/pr/12 when the revision contains /. Private or gated repositories need ModelLoadOptions(bearerToken: hfToken) (or custom headers); do not put tokens in source strings. Multimodal repos often ship a separate mmproj GGUF file—treat it as a separate asset/source. Sharded GGUF repos are not expanded automatically by llamadart; choose a single-file GGUF unless you are handling shards yourself.

Searches the latest release. Esc to close.