Unreleased documentation for the next version. Read v0.9.0, the latest release

llamadart: on-device LLMs for Flutter and Dart

llamadart runs LLMs on-device in Flutter and Dart apps, with one API for GGUF and LiteRT-LM models on six platforms. See what it does and where to start.

On this page

llamadart is a Dart package for on-device inference. It runs GGUF models through llama.cpp and .litertlm bundles through LiteRT-LM, on Android, iOS, macOS, Linux, Windows and the web, behind one Dart API.

Who this is for#

  • Flutter and Dart developers who want AI features that run on the user's device: prompts stay private, and there is no inference server or API key.
  • Apps that must work offline. Native targets need no network once the model is on the device; on the web, the page fetches the runtime and model over the network, then runs inference in the browser.
  • Teams that want one codebase across mobile, desktop and web instead of a separate SDK per platform.
  • Tools that need an OpenAI-compatible HTTP endpoint backed by a local model: see the OpenAI-compatible server example.

What you can build#

Core primitives#

  • LlamaEngine: stateless generation API.
  • ChatSession: stateful chat wrapper over LlamaEngine.
  • LlamaBackend: platform backend abstraction used by the engine.

Where to go next#

Searches the latest release. Esc to close.