How llamadart works
How llamadart layers one Dart API over llama.cpp and LiteRT-LM, with FFI bindings and worker isolates on native and JavaScript runtimes on the web.
On this page
llamadart exposes one Dart API (LlamaEngine, ChatSession and the typed
speech and decision engines) over two inference runtimes:
- llama.cpp runs GGUF models. It is built on GGML, a tensor library with CPU kernels (NEON, AVX) and GPU backends (Metal, Vulkan, CUDA and more).
-
LiteRT-LM runs
.litertlmbundles. See Choosing llama.cpp or LiteRT-LM.
Architecture overview#
Dart & Flutter application layer
Native llama.cpp & GGML layer
Hardware compute
Native targets#
-
Prebuilt runtimes. During
flutter buildordart run, the package's build hook downloads the prebuilt runtime bundles for the target platform fromllamadart-native(llama.cpp) andlitert-lm-native(LiteRT-LM), so apps never compile C++. See Native build hooks. -
FFI bindings. Dart FFI calls the llama.cpp C API (
llama.h,mtmd.hfor multimodal input, and a thinllamadart-nativewrapper) and the LiteRT-LM C API. llamadart does not use llama.cpp'scommonhelpers: model loading with memory mapping (ModelParams.useMmap), tokenization and sampler chains are all libllama calls. - Worker isolates. Each backend runs native calls in a background isolate, so inference never blocks the UI isolate. Results stream back as Dart streams.
-
Explicit lifecycle. Models and contexts are native memory. Release them
with
await engine.unloadModel()andawait engine.dispose(), typically intry/finally, instead of relying on garbage collection. See Model lifecycle.
Web#
On the web, the same Dart API talks to JavaScript runtimes through interop:
- GGUF models run in the WebGPU bridge, a llama.cpp build for WebGPU with a CPU (WebAssembly) path.
.litertlmmodels run through the official@litert-lm/corebrowser API.
Capabilities differ by runtime; the support matrix lists what each one supports.
Chat templates#
Chat template detection, rendering and output parsing are reimplemented in Dart, in line with llama.cpp, so tool calling and reasoning parsing behave the same on native and web. See Chat templates and output parsing.