Documentation for v0.10.0, an older release. Read v0.11.1, the latest release

Chat app example

A Flutter chat app with a model library, queued downloads, runtime controls, multimodal input, speech, image generation and streaming chat.

On this page

Path: example/chat_app · Platforms: Android, iOS 16.4+, macOS 14.0+, Windows, Linux, Web

A Flutter chat app that downloads models from a built-in library and streams on-device chat, with runtime controls and diagnostics. Live demo: https://leehack-llamadart.static.hf.space

Run#

cd example/chat_app
flutter pub get
flutter run

Variants:

# Native downloads authenticated with a Hugging Face token
flutter run --dart-define=HF_TOKEN=<your_token>

# Web, from the repository root: build with the bridge assets, then serve
# with cross-origin isolation headers
./scripts/build_chat_app_web.sh
python3 tool/testing/serve_static_with_headers.py \
  --directory example/chat_app/build/web --port 8080

# Web smoke test with a mock bridge and no model download
dart run tool/testing/run_local_e2e.dart --scenario chat-app-web-mock-smoke

What it demonstrates#

  • Streaming chat with per-model sampling presets, thinking output, and copy and regenerate actions (Generation and streaming).
  • Tool-calling toggles and editable tool declarations (Tool calling).
  • A model library with queued, cancellable downloads through ModelDownloadController, cached across launches (Downloads and cache).
  • Image and audio attachments, including clipboard paste, enabled only when the loaded projector or bundle reports the capability (Multimodal).
  • GGUF and .litertlm models through one engine API; the app enables both native runtime families, which an app that ships only GGUF does not need (Backend selection).
  • Backend, GPU layer, context and batch controls, Auto memory planning, and runtime diagnostics (Performance tuning).

Speech and voice#

With a Qwen3-ASR model loaded, the composer transcribes a selected file or a microphone recording of up to 30 seconds. On Android, iOS, macOS and Windows, native chat models can add live English dictation through a separately installed LiteRT model: Moonshine Tiny (54 MB) or Parakeet TDT 0.6B (615 MB). Native Gemma 4 E2B answers a spoken question through Ask with voice, and the Qwen3-TTS preset switches the composer to speech synthesis. See Speech to text and Text to speech.

Image generation#

Image generation in the sidebar opens a Preview text-to-image screen on the opt-in stable_diffusion runtime, which the app bundles. It downloads SDXS-512 (683 MB) or SD-Turbo with TAESD (2.0 GB), then generates at 256 or 512 px with phase progress, cancellation, a reusable seed and PNG save. A model that does not fit in memory shows the engine's refusal, with an offer to unload the chat model. The web and targets without the runtime show why generation is unavailable. See Image generation, including its model licenses.

Test#

cd example/chat_app
flutter test
flutter test --platform chrome test/chat_generation_service_test.dart test/image_generation_screen_test.dart

The second command covers Web-only paths.

Full options: the model catalog, download behavior, settings, platform and validation status, Web and Android notes, and troubleshooting are in the example README.

Searches the latest release. Esc to close.