Introduction
Learn what llamadart provides and where to start when building local AI features in Dart and Flutter.
On this page
llamadart is a Dart and Flutter plugin for local LLMs. It runs GGUF models
through llama.cpp across native and web targets, and routes .litertlm
bundles through LiteRT-LM native and web runtimes.
Who this is for#
- App developers building local-first AI features in Dart/Flutter.
- Teams that need OpenAI-style HTTP compatibility from local models.
- Maintainers who need predictable native/web runtime integration.
Core primitives#
LlamaEngine: stateless generation API.ChatSession: stateful chat wrapper overLlamaEngine.LlamaBackend: platform backend abstraction used by the engine.
Read by workflow#
- First setup: Installation
- First inference: Quickstart
- Multi-turn chat: First Chat Session
- Backend choice: Choosing llama.cpp or LiteRT-LM
- Embedding pipelines: Embeddings
- Function calling: Tool Calling
- Template diagnostics: Chat Templates and Parsing
- Template internals: Template Engine Internals
- LoRA runtime workflows: LoRA Adapters
- Performance work: Performance Tuning
- Backend benchmark results: Backend Benchmarks
- Platform/backend planning: Platform & Backend Matrix
- Upgrade planning: Upgrade Checklist
- Maintainer operations: Maintainer Overview