Common Issues
On this page
Runtime bundle or native asset load failure#
Symptoms:
- Model fails to load on first run.
- Errors about missing native libs.
Checks:
- Ensure internet connectivity for first runtime bundle resolution.
- Verify your app can access GitHub release endpoints.
- If backend config changed recently, run
flutter cleanonce.
Model path or URL issues#
Symptoms:
Failed to load modelerrors.
Checks:
- Confirm path exists and is readable.
- Confirm file is valid GGUF.
- For URL loading, confirm backend/platform supports URL model load.
Slow generation#
Checks:
- Reduce model size or quantization level.
- Tune
contextSizeand generation length (maxTokens). - Use appropriate backend and GPU offload (
gpuLayers).
DFlash draft model load failures#
Symptoms:
- DFlash speculative decoding fails while loading the draft GGUF.
- Native logs mention
unknown model architecture: 'dflash-draft'. - Native logs mention missing DFlash target-layer metadata.
Checks:
-
Inspect the draft GGUF metadata. Compatible DFlash artifacts use
general.architecture=dflash, notdflash-draft. -
Confirm the draft contains the DFlash metadata block, especially
dflash.target_layers. -
Use an upstream-compatible target/draft pair. One validated public pair is
target
unsloth/Qwen3.5-4B-GGUF(Qwen3.5-4B-Q4_K_M.gguf) with draftEntityDeletr/Qwen3.5-4B-DFlash-GGUF(Qwen3.5-4B-DFlash.gguf). -
If the artifact uses
dflash-draftor lacksdflash.target_layers, replace or reconvert the draft GGUF. llamadart does not patch draft metadata at runtime.
Tool calling seems unstable#
Checks:
- Use
ToolChoice.autobefore forcingrequired. - Lower temperature for tool-calling requests.
- Validate tool schema and required parameters.
- Ensure your loop appends tool result messages correctly.
Web behavior differs from native#
Symptoms:
- The app loads, but model load fails only on web.
- WebGPU falls back to CPU or reports lower GPU layers than requested.
- A hosted build behaves differently from
localhost.
Checks:
-
Confirm bridge runtime is loaded successfully:
window.LlamaWebGpuBridgeshould exist andwindow.__llamadartBridgeLoadErrorshould be empty. -
Verify browser capability: secure context,
navigator.gpu,requestAdapter(), adapter features/limits, and current GPU drivers. -
For large single-file GGUF loads, verify cross-origin isolation:
window.crossOriginIsolated === trueand the app origin sends COOP/COEP headers. -
Distinguish bridge-load failures from model/config pressure. Memory errors,
bad_alloc,memory access out of bounds, or aborts often mean the model, context size, thread count, or GPU-layer count is too large for the current browser. - Validate model URLs, CORS/CORP policy, base href, service-worker cache state, and whether the runtime came from CDN or local assets.
- See WebGPU Bridge for the readiness probe, fallback rules, and Flutter Web smoke-test command.
High log noise#
Use split log levels:
await engine.setDartLogLevel(LlamaLogLevel.warn);
await engine.setNativeLogLevel(LlamaLogLevel.error);