Documentation for v0.8.17, an older release. Read v0.9.0, the latest release

Quickstart

Load a GGUF or LiteRT-LM model, generate tokens, and try embeddings with the core llamadart APIs in minutes.

On this page

This quickstart uses the core LlamaEngine API.

Minimal generation example#

Start with a model source instead of a machine-specific file path. On native Dart/Flutter targets, loadModelSource(...) downloads the file on first run, stores it in the package-managed model cache, and reuses the cached file on later runs.

import 'package:llamadart/llamadart.dart';

Future<void> main() async {
  final LlamaEngine engine = LlamaEngine(LlamaBackend());

  try {
    await engine.loadModelSource(
      ModelSource.parse(
        'hf://unsloth/SmolLM2-135M-Instruct-GGUF/'
        'SmolLM2-135M-Instruct-Q2_K.gguf',
      ),
      modelParams: const ModelParams(contextSize: 1024, gpuLayers: 0),
      onProgress: (progress) {
        final fraction = progress.fraction;
        if (fraction != null) {
          print('download ${(fraction * 100).toStringAsFixed(1)}%');
        }
      },
    );

    final output = StringBuffer();
    await for (final chunk in engine.create(
      const [
        LlamaChatMessage.fromText(
          role: LlamaChatRole.user,
          text: 'Rewrite professionally: i need this done asap',
        ),
      ],
      params: const GenerationParams(maxTokens: 64, temp: 0.2),
    )) {
      final text = chunk.choices.first.delta.content;
      if (text != null) {
        output.write(text);
      }
    }
    print(output.toString());
  } finally {
    await engine.dispose();
  }
}

This example uses engine.create(...), the stateless chat-completion API: the model's chat template is applied, but no conversation history is stored between calls. Use First Chat Session when you want automatic multi-turn history, or Generation and Streaming when you need to choose between raw prompts, stateless chat, and stateful chat.

The small SmolLM2 GGUF above is intended for copy/paste smoke tests. For a live conference demo, run it once beforehand so the Hugging Face source remains in the code while the actual presentation path uses the local cache instead of conference Wi-Fi.

LiteRT-LM .litertlm bundles load through the same engine. Native targets load local bundle paths, including paths resolved by loadModelSource(...); web targets load web-compatible .litertlm URLs through the @litert-lm/core JavaScript runtime.

await engine.loadModel(
  'path/to/gemma-4-E2B-it.litertlm',
  modelParams: const ModelParams(
    liteRtLmBackend: LiteRtLmBackendPreference.gpu,
  ),
);

LiteRtLmBackendPreference.auto is the default. It chooses GPU on Android, macOS, and web, and CPU on other current LiteRT-LM targets. Android native callers can request LiteRtLmBackendPreference.npu for devices and model bundles that support the LiteRT-LM NPU delegate. Web rejects NPU selection explicitly.

Stateless chat completions#

For OpenAI-style message arrays, use engine.create(...):

final messages = [
  LlamaChatMessage.fromText(
    role: LlamaChatRole.user,
    text: 'Give me three bullet points about Dart.',
  ),
];

await for (final chunk in engine.create(messages)) {
  final text = chunk.choices.first.delta.content;
  if (text != null) {
    print(text);
  }
}

Embeddings (single and batch)#

final single = await engine.embed('hello world');
final batch = await engine.embedBatch([
  'semantic search',
  'document retrieval',
]);

print('single dims=${single.length}');
print('batch size=${batch.length}');

Embeddings are a llama.cpp/GGUF capability in the current package. Check engine.supportsEmbeddings before calling these APIs when your app can switch between GGUF and LiteRT-LM models.

Next steps#

Searches the latest release. Esc to close.