Runtime Parameters
On this page
Runtime behavior is primarily controlled by:
ModelParamsat model load time.GenerationParamsper generation call.
For a strategy-focused walkthrough on how to change these knobs and what to measure, see Performance Tuning.
ModelParams essentials#
await engine.loadModel(
'/path/to/model.gguf',
modelParams: const ModelParams(
contextSize: 4096,
gpuLayers: ModelParams.maxGpuLayers,
preferredBackend: GpuBackend.vulkan,
splitMode: ModelSplitMode.layer,
mainGpu: 0,
numberOfThreads: 0,
numberOfThreadsBatch: 0,
batchSize: 0,
microBatchSize: 0,
maxParallelSequences: 1,
),
);
Important fields:
contextSize: total context window.gpuLayers: number of layers offloaded to GPU.-
preferredBackend: backend preference (auto,vulkan,metal, etc). -
splitMode: model tensor distribution mode passed through to llama.cppsplit_mode. Defaults to upstreamlayerbehavior. -
mainGpu: primary GPU device index passed through to llama.cppmain_gpu. To select one GPU for the full model, usesplitMode: ModelSplitMode.nonewith the desiredmainGpuindex. -
batchSize: context logical batch size (n_batch). When left at0, llamadart uses the effective context size (n_ctx) on native and WebGPU backends, except for model-specific WebGPU safety tuning such as the bundled Qwen3.5-0.8B small-model preset. -
microBatchSize: context micro-batch size (n_ubatch). When left at0, llamadart uses the resolvedbatchSize; explicit values are capped son_ubatch <= n_batch <= n_ctx. -
maxParallelSequences: max sequence slots (n_seq_max) for parallel sequence workloads (for example, batched embeddings). chatTemplate: optional template override.-
preferMemory64(web/WebGPU only): prefer the 64-bit (wasm64/mem64) bridge core. The default 32-bit core has a 4 GiB address-space limit, but large models need room for KV cache and intermediate buffers.null(default) lets llamadart decide frommodelBytesHintusing the current wasm32-safe ceiling (about 2 GiB of model bytes);trueforces mem64;falseforces wasm32. Ignored on non-web backends. -
modelBytesHint(web/WebGPU only): approximate model size in bytes, used to select the mem64 core up front instead of waiting for an out-of-memory retry. Ignored on non-web backends. -
liteRtLmActivationDataType,liteRtLmPrefillChunkSize,liteRtLmParallelFileSectionLoading, andliteRtLmDispatchLibDir: opt-in native LiteRT-LM.litertlmengine settings. Leave them unset to preserve runtime defaults; LiteRT-LM web rejects them as native-only. -
numberOfThreads: honored by native LiteRT-LM;0keeps automatic selection. -
loras: native LiteRT-LM accepts one default-scale initial text adapter at model load. Runtime LoRA control APIs, adapter stacking, and custom scales remain llama.cpp-only.
For runtime LoRA control (setLora, removeLora, clearLoras), see
LoRA Adapters.
Embedding-oriented model params#
For high-throughput embedBatch(...), tune context batch fields together:
- Keep
batchSizelarge enough for total tokens across your average batch. -
Set
microBatchSizeclose tobatchSizeunless you need tighter memory bounds. -
Increase
maxParallelSequencesabove1(for example2,4,8) to enable true multi-sequence embedding batching.
See Embeddings for API usage and benchmark scripts.
GenerationParams essentials#
const params = GenerationParams(
maxTokens: 512,
temp: 0.7,
topK: 40,
topP: 0.9,
minP: 0.0,
penalty: 1.1,
stopSequences: ['</s>'],
speculativeDecoding: false,
speculativeDecodingConfig: null,
);
Important fields:
maxTokens: generation length cap.temp: randomness.topK,topP,minP: token filtering controls.penalty: repeat penalty.-
speculativeDecoding/speculativeDecodingConfig: opt-in backend-native speculative decoding. Native LiteRT-LM honors the legacy boolean flag. llama.cpp supportsSpeculativeDecodingConfig.mtp(...)for compatible MTP GGUF models, including separate target/draft model pairs throughdraftModelPath. WebGPU and LiteRT-LM web reject speculative decoding until their speculative paths are implemented. seed: deterministic replay when set.grammar: constrained decoding with GBNF.
Practical tuning defaults#
- Deterministic extraction: lower
temp(0.1-0.3) + explicit stops. -
General chat:
temparound0.6-0.9,topParound0.9-0.95. - Tool calling: stable
tempand sufficientmaxTokensfor call payload.