llamadart: on-device LLMs for Flutter and Dart
llamadart runs LLMs on-device in Flutter and Dart apps, with one API for GGUF and LiteRT-LM models on six platforms. See what it does and where to start.
On this page
llamadart is a Dart package for on-device inference. It runs GGUF models
through llama.cpp and .litertlm bundles through LiteRT-LM, on Android, iOS,
macOS, Linux, Windows and the web, behind one Dart API.
Who this is for#
- Flutter and Dart developers who want AI features that run on the user's device: prompts stay private, and there is no inference server or API key.
- Apps that must work offline. Native targets need no network once the model is on the device; on the web, the page fetches the runtime and model over the network, then runs inference in the browser.
- Teams that want one codebase across mobile, desktop and web instead of a separate SDK per platform.
- Tools that need an OpenAI-compatible HTTP endpoint backed by a local model: see the OpenAI-compatible server example.
What you can build#
- Streaming chat and text generation, with structured JSON output.
- Tool calling driven by the model's chat template.
- Embeddings for search and retrieval.
- Image and audio input with multimodal models.
- Speech to text and text to speech.
- Decision models that answer typed questions without generating text.
- Runtime LoRA adapters.
Core primitives#
LlamaEngine: stateless generation API.ChatSession: stateful chat wrapper overLlamaEngine.LlamaBackend: platform backend abstraction used by the engine.
Where to go next#
- New to llamadart: Install it, then run the Quickstart.
- Building a Flutter app: follow the Flutter chat app tutorial.
- Choosing a model: Finding models and Model families.
- Answering from your own documents: the RAG tutorial.
- Shipping: the support matrix, performance tuning and troubleshooting.
- Upgrading: the upgrade checklist.