Documentation for v0.8.15, an older release. Read v0.9.0, the latest release

Runtime Parameters

On this page

Runtime behavior is primarily controlled by:

  • ModelParams at model load time.
  • GenerationParams per generation call.

For a strategy-focused walkthrough on how to change these knobs and what to measure, see Performance Tuning.

ModelParams essentials#

await engine.loadModel(
  '/path/to/model.gguf',
  modelParams: const ModelParams(
    contextSize: 4096,
    gpuLayers: ModelParams.maxGpuLayers,
    preferredBackend: GpuBackend.vulkan,
    splitMode: ModelSplitMode.layer,
    mainGpu: 0,
    numberOfThreads: 0,
    numberOfThreadsBatch: 0,
    batchSize: 0,
    microBatchSize: 0,
    maxParallelSequences: 1,
  ),
);

Important fields:

  • contextSize: total context window.
  • gpuLayers: number of layers offloaded to GPU.
  • preferredBackend: backend preference (auto, vulkan, metal, etc).
  • splitMode: model tensor distribution mode passed through to llama.cpp split_mode. Defaults to upstream layer behavior.
  • mainGpu: primary GPU device index passed through to llama.cpp main_gpu. To select one GPU for the full model, use splitMode: ModelSplitMode.none with the desired mainGpu index.
  • batchSize: context logical batch size (n_batch). When left at 0, llamadart uses the effective context size (n_ctx) on native and WebGPU backends, except for model-specific WebGPU safety tuning such as the bundled Qwen3.5-0.8B small-model preset.
  • microBatchSize: context micro-batch size (n_ubatch). When left at 0, llamadart uses the resolved batchSize; explicit values are capped so n_ubatch <= n_batch <= n_ctx.
  • maxParallelSequences: max sequence slots (n_seq_max) for parallel sequence workloads (for example, batched embeddings).
  • chatTemplate: optional template override.
  • preferMemory64 (web/WebGPU only): prefer the 64-bit (wasm64/mem64) bridge core. The default 32-bit core has a 4 GiB address-space limit, but large models need room for KV cache and intermediate buffers. null (default) lets llamadart decide from modelBytesHint using the current wasm32-safe ceiling (about 2 GiB of model bytes); true forces mem64; false forces wasm32. Ignored on non-web backends.
  • modelBytesHint (web/WebGPU only): approximate model size in bytes, used to select the mem64 core up front instead of waiting for an out-of-memory retry. Ignored on non-web backends.
  • liteRtLmActivationDataType, liteRtLmPrefillChunkSize, liteRtLmParallelFileSectionLoading, and liteRtLmDispatchLibDir: opt-in native LiteRT-LM .litertlm engine settings. Leave them unset to preserve runtime defaults; LiteRT-LM web rejects them as native-only.
  • numberOfThreads: honored by native LiteRT-LM; 0 keeps automatic selection.
  • loras: native LiteRT-LM accepts one default-scale initial text adapter at model load. Runtime LoRA control APIs, adapter stacking, and custom scales remain llama.cpp-only.

For runtime LoRA control (setLora, removeLora, clearLoras), see LoRA Adapters.

Embedding-oriented model params#

For high-throughput embedBatch(...), tune context batch fields together:

  • Keep batchSize large enough for total tokens across your average batch.
  • Set microBatchSize close to batchSize unless you need tighter memory bounds.
  • Increase maxParallelSequences above 1 (for example 2, 4, 8) to enable true multi-sequence embedding batching.

See Embeddings for API usage and benchmark scripts.

GenerationParams essentials#

const params = GenerationParams(
  maxTokens: 512,
  temp: 0.7,
  topK: 40,
  topP: 0.9,
  minP: 0.0,
  penalty: 1.1,
  stopSequences: ['</s>'],
  speculativeDecoding: false,
  speculativeDecodingConfig: null,
);

Important fields:

  • maxTokens: generation length cap.
  • temp: randomness.
  • topK, topP, minP: token filtering controls.
  • penalty: repeat penalty.
  • speculativeDecoding / speculativeDecodingConfig: opt-in backend-native speculative decoding. Native LiteRT-LM honors the legacy boolean flag. llama.cpp supports the upstream strategy surface: SpeculativeDecodingConfig.mtp(...), draftSimple(...), draftEagle3(...), draftDflash(...), ngramSimple(...), ngramMapK(...), ngramMapK4v(...), ngramMod(...), ngramCache(...), and mixed(...) for draftless n-gram strategies plus one draft-model strategy. Draft-model strategies can load a separate GGUF through draftModelPath; draftless n-gram strategies use token history or n-gram caches without a draft model. WebGPU and LiteRT-LM web reject speculative decoding until their speculative paths are implemented.
  • seed: deterministic replay when set.
  • grammar: constrained decoding with GBNF.

Practical tuning defaults#

  • Deterministic extraction: lower temp (0.1-0.3) + explicit stops.
  • General chat: temp around 0.6-0.9, topP around 0.9-0.95.
  • Tool calling: stable temp and sufficient maxTokens for call payload.

Searches the latest release. Esc to close.