Performance Tuning
On this page
- Pick a tuning goal
- Suggested workflow
- Model load tuning (ModelParams)
- Highest-impact load-time knobs
- Generation tuning (GenerationParams)
- Native LiteRT-LM runtime controls
- Multimodal tuning
- Read the diagnostics you already have
- General heuristics by environment
- Practical diagnostics
- Keep the tuning guide model-agnostic
Performance tuning depends on model size, quantization, backend availability, and context/generation settings.
The most reliable approach is to treat tuning as a measurement problem:
- pick a representative prompt or workload
- record baseline timings
- change one variable at a time
- keep the fastest stable configuration
Pick a tuning goal#
Different knobs help different problems:
-
first-token latency: optimize load-time/runtime setup, prompt size, and prompt evaluation cost. sustained throughput: optimize decode path, backend choice, and batching.-
stability: lower GPU pressure, reduce context, and keep multimodal inputs smaller. -
multimodal responsiveness: reduce image/audio size first, then revisit backend and token budget.
If you do not know which goal matters most, start with latency and stability.
Suggested workflow#
- Benchmark the exact prompt shape you care about.
- Compare
cpuand GPU backends before changing anything else. -
Reduce
contextSizeandmaxTokensto the smallest values that still fit your use case. -
Tune load-time/runtime knobs (
gpuLayers, threads, batch sizes) one at a time. -
Only after runtime is stable, tune sampling (
temp,topK,topP, etc.).
This order matters: sampling changes usually affect output style more than raw runtime cost.
Model load tuning (ModelParams)#
const modelParams = ModelParams(
contextSize: 4096,
gpuLayers: ModelParams.maxGpuLayers,
preferredBackend: GpuBackend.vulkan,
numberOfThreads: 0,
numberOfThreadsBatch: 0,
);
Guidelines:
-
Start with backend choice first. Small/mobile models can be faster on
cputhanvulkan/webgpu, while larger or longer-running workloads may favor GPU acceleration. -
Start with default
gpuLayers, then lower them if stability or latency is worse than CPU. -
Keep
contextSizeonly as large as your use case needs. Oversized context is one of the easiest ways to hurt first-token latency. -
Use explicit
numberOfThreads/numberOfThreadsBatchonly after measuring. Auto-threading is often a good baseline, but some mobile devices prefer fewer threads for lower contention. -
Tune
batchSizeandmicroBatchSizeconservatively on unstable GPU paths. Bigger is not always faster if it increases driver/scheduler overhead. - Use backend preference that matches your actual target runtime, not just the hardware you hope to use.
Highest-impact load-time knobs#
preferredBackend: biggest high-level choice; always measure CPU against GPU.-
gpuLayers: GPU offload depth; can help throughput but may hurt stability or even latency on small models. contextSize: affects prompt evaluation cost and memory footprint directly.-
numberOfThreads,numberOfThreadsBatch: mostly relevant for CPU and hybrid paths. -
batchSize,microBatchSize: scheduler/batching controls for native and web runtimes. -
maxParallelSequences: relevant for embedding or true multi-sequence workloads, not regular single-turn chat.
Generation tuning (GenerationParams)#
const generationParams = GenerationParams(
maxTokens: 256,
temp: 0.7,
topK: 40,
topP: 0.9,
minP: 0.0,
penalty: 1.1,
reusePromptPrefix: true,
streamBatchTokenThreshold: 8,
streamBatchByteThreshold: 512,
);
Guidelines:
- Lower
maxTokensfor latency-sensitive paths. - Lower
tempfor deterministic/extraction tasks. - Adjust
topPandtopKgradually; avoid drastic simultaneous changes. -
Treat
maxTokensas a performance knob as much as a quality knob. If you only need short answers, cap it aggressively. -
penalty,topK,topP, andtempusually do not fix a slow backend; they mainly shape output behavior. -
Native backends can tune stream transport overhead with
streamBatchTokenThresholdandstreamBatchByteThreshold. - Lower stream thresholds improve token-by-token UI granularity, while higher values improve throughput by reducing isolate message overhead.
-
Use speculative decoding only after benchmarking your target model/device; the
default remains off because it is not a universal speedup. Native LiteRT-LM
uses the legacy
speculativeDecodingboolean, while llama.cpp MTP usesSpeculativeDecodingConfig.mtp(...)and can optionally load a separate draft GGUF throughdraftModelPath. -
reusePromptPrefixis enabled by default for native generation; keep it on for multi-turn chats and repeated prompts, and validate parity for your target model/workload. - Native reuse is optimized for evolving prompts with shared prefixes. Exact prompt replays are re-ingested to preserve deterministic parity.
Native LiteRT-LM runtime controls#
Native .litertlm loads expose a small set of LiteRT-LM-specific
ModelParams fields:
-
liteRtLmActivationDataType: override upstream activation data type (float32,float16,int16, orint8). liteRtLmPrefillChunkSize: positive CPU dynamic-model prefill chunk size.-
liteRtLmParallelFileSectionLoading:nullkeeps the native default,falsedisables parallel.litertlmfile-section loading for diagnostics. liteRtLmDispatchLibDir: Android NPU LiteRT dispatch library directory.-
numberOfThreads: generation thread count;0keeps LiteRT-LM automatic selection. -
loras: one default-scale initial text LoRA adapter for native.litertlmloads. Runtime adapter updates, stacking, and scaling remain llama.cpp-only.
Leave these values unset unless you have a target-model reason to change them. Benchmark load time, prefill throughput, decode throughput, and output quality on the deployment device after changing activation type or prefill chunk size. LiteRT-LM web rejects these native-only fields.
The real-model smoke tool accepts matching environment variables and reports the selected values in its JSON result:
LITERT_LM_ACTIVATION_DATA_TYPE=float16 \
LITERT_LM_PREFILL_CHUNK_SIZE=128 \
LITERT_LM_PARALLEL_FILE_SECTION_LOADING=false \
dart run tool/litert_lm_engine_smoke.dart /models/model.litertlm cpu
Multimodal tuning#
- Reduce image size before doing anything else.
- Keep
contextSizeandmaxTokenstighter than your text-only defaults. - If GPU multimodal is unstable, try CPU first to establish a correctness baseline.
- Once CPU multimodal works, revisit GPU/offload settings carefully.
- Treat projector loading and actual multimodal generation as separate stages; one can be healthy while the other is still too slow or unstable.
Read the diagnostics you already have#
Good tuning depends on reading the right signals.
-
first: first-token latency; if this is high, focus on model load, prompt size, context, and prompt evaluation. total: end-to-end wall time.avg: overall throughput across the whole request.decode: steady-state generation speed once output starts.
If your app exposes native llama.cpp timing chips or logs:
-
p_eval: prompt evaluation time. High values usually mean prompt/context overhead, not sampler overhead. -
eval: decode time for generated tokens. High values usually point to backend kernel/scheduler cost. -
sample: token selection overhead. Usually small; if large, inspect runtime overhead or unusual sampling settings. -
reuse: prompt-prefix reuse count. If reuse stays low in multi-turn chat, cached prefix optimization is not helping much.
These numbers help you decide whether to tune prompt size, decode path, or sampling.
General heuristics by environment#
mobile native: test CPU vs GPU early; small models often favor CPU.-
desktop native: GPU is more likely to pay off as model size or response length grows. -
browser: prefer conservative GPU settings first; browser GPU paths usually have tighter stability limits than native. -
multimodal: expect stricter limits than text-only, especially on mobile and browser targets.
Practical diagnostics#
- Measure token throughput with representative prompts.
-
Keep comparisons fair: same model, same prompt, same
contextSize, samemaxTokens, same backend-specific limits. - Record both latency and throughput; a setting that improves one can hurt the other.
- Run prompt-reuse parity checks before relying on prefix reuse in production:
dart run tool/testing/native_prompt_reuse_parity.dart \
--model path/to/model.gguf \
--prompt-file tool/testing/prompts/native_prompt_reuse_parity_prompts.txt \
--max-prompts 8 \
--runs 3 \
--fail-on-mismatch
# Benchmark native generate/create TTFT and throughput
dart run tool/testing/native_inference_benchmark.dart \
--model path/to/model.gguf \
--gpu-layers 0 \
--mode all \
--runs 3 \
--max-tokens 128
# Benchmark embeddings (sequential vs batch)
dart run tool/testing/native_embedding_benchmark.dart \
--model path/to/model.gguf \
--cpu \
--mode both \
--input-count 8 \
--max-seq 8
# Sweep max-seq values and export CSV for plotting
dart run tool/testing/native_embedding_sweep.dart \
--model path/to/model.gguf \
--cpu \
--max-seq-values 1,2,4,8 \
--csv-out embedding_speedup.csv
Compare llama.cpp/GGUF and LiteRT-LM with the bundled fair benchmark scripts:
# macOS native, Gemma 4 E2B artifacts
DECODE_TOKENS=256 tool/macos_fair_litert_vs_llamadart.sh
# Native LiteRT-LM speculative decoding off/on comparison
SPECULATIVE=false DECODE_TOKENS=256 tool/macos_fair_litert_vs_llamadart.sh
SPECULATIVE=true DECODE_TOKENS=256 tool/macos_fair_litert_vs_llamadart.sh
# Web LiteRT-LM; use TARGETS=llamadart to test GGUF WebGPU separately
DOWNLOAD_LITERT_WEB_MODEL=1 \
DECODE_TOKENS=256 \
WARMUPS=1 \
RUNS=3 \
TARGETS=litert_lm \
tool/web_fair_litert_vs_llamadart.sh
# Android / Pixel-style app benchmark
ADB=/path/to/adb
DEVICE=<adb-serial>
"$ADB" -s "$DEVICE" shell svc power stayon true
"$ADB" -s "$DEVICE" shell input keyevent KEYCODE_WAKEUP
DEVICE="$DEVICE" ADB="$ADB" OUTPUT_TOKENS=256 WARMUPS=1 RUNS=3 \
TARGETS=litert_lm,llamadart tool/litert_lm_pixel_benchmark.sh
DEVICE="$DEVICE" ADB="$ADB" TARGETS=litert_lm BACKEND=gpu \
SPECULATIVE=false OUTPUT_TOKENS=256 WARMUPS=1 RUNS=3 \
tool/litert_lm_pixel_benchmark.sh
DEVICE="$DEVICE" ADB="$ADB" TARGETS=litert_lm BACKEND=gpu \
SPECULATIVE=true OUTPUT_TOKENS=256 WARMUPS=1 RUNS=3 \
tool/litert_lm_pixel_benchmark.sh
Current measured Gemma 4 E2B results are recorded in Backend Benchmarks.
- Validate memory behavior with your real context sizes.
- Check runtime backend and VRAM info where available:
final backendName = await engine.getBackendName();
final vram = await engine.getVramInfo();
print('$backendName total=${vram.total} free=${vram.free}');
Keep the tuning guide model-agnostic#
Specific models may need special-case defaults in applications, but the tuning process should stay general:
- define the goal
- measure the baseline
- change one knob at a time
- keep the fastest stable result
That workflow transfers much better than any one model-specific recipe.