Text generation and streaming
Stream tokens with generate, create and ChatSession; use structured JSON output, thinking budgets, operation observers, cancellation and tokenization helpers.
On this page
llamadart exposes three generation entry points:
engine.generate(prompt)for raw prompt strings.engine.create(messages)for stateless, chat-template aware completions.-
ChatSession.create(parts)for stateful, multi-turn chat with automatic history management.
Choosing the right API#
| API | Template-aware? | Keeps history? | Use when |
|---|---|---|---|
engine.generate(prompt) |
No | No | You already rendered the final raw prompt, or you are benchmarking, testing prefix-cache/state flows, or doing other low-level runtime work. |
engine.create(messages) |
Yes | No |
You have the complete
List<LlamaChatMessage>
for each request, such as an OpenAI-compatible server, a one-shot completion, or an app that owns its transcript.
|
ChatSession.create(parts) |
Yes | Yes | You are building a multi-turn chat UI/CLI and want the SDK to append user/assistant turns, apply the system prompt, and trim history as the context grows. |
For one-shot instructions, prefer engine.create(...): it applies the chat
template without session state, and a follow-up turn sees only the messages you
pass again. For chat apps, prefer ChatSession unless your app already stores
the transcript. session.addMessage(...) restores history or inserts tool
results, and session.reset() starts over. See
First Chat Session for a multi-turn
example.
Generation pipeline (visual)#
sequenceDiagram
autonumber
participant App as App/ChatSession
participant Engine as LlamaEngine
participant Template as Template engine
participant Backend as Native/Web backend
participant Parser as Stream parser
App->>Engine: generate(prompt), create(messages), or ChatSession.create(parts)
alt ChatSession.create(parts)
App->>App: append turn to session history
App->>Engine: create(full session messages)
end
alt create(messages)
Engine->>Template: detect format + render template
Template-->>Engine: prompt + stops + grammar
end
Engine->>Backend: start generation
loop token stream
Backend-->>Engine: token bytes
Engine->>Parser: UTF-8 decode + partial parse
Parser-->>App: streaming chunk delta
end
Backend-->>Engine: generation finished
Engine->>Parser: finalize parse
Parser-->>App: final chunk (finish reason/tool calls)
Low-level generation API#
await for (final token in engine.generate(
'List two advantages of local LLM inference.',
params: const GenerationParams(maxTokens: 64, temp: 0.4),
)) {
print(token);
}
Chat completion API#
final messages = [
LlamaChatMessage.fromText(
role: LlamaChatRole.user,
text: 'Explain top-p in plain language.',
),
];
await for (final chunk in engine.create(
messages,
params: const GenerationParams(maxTokens: 128, topP: 0.95),
)) {
final thinking = chunk.choices.first.delta.thinking;
if (thinking != null) {
print('[thinking] $thinking');
}
final text = chunk.choices.first.delta.content;
if (text != null) {
print(text);
}
}
Token usage and timings#
On native llama.cpp, and on WebGPU with bridge assets v0.1.54+, the final
create chunk carries the request's usage whenever the backend reports it.
The backend can report none, for example for a request cancelled while it is
queued. Usage is null on LiteRT-LM, on older bridge assets and on every
earlier chunk. On WebGPU, completionTokens can include tokens generated after
a stop sequence, before the stop reached the bridge.
final chunks = await engine.create(messages).toList();
final usage = chunks.last.usage;
if (usage != null) {
print('prompt ${usage.promptTokens} '
'(cached ${usage.cachedPromptTokens}), '
'completion ${usage.completionTokens}, '
'first token ${usage.timeToFirstToken}, total ${usage.duration}');
}
The backend times timeToFirstToken and duration from when it starts the
request. They exclude template rendering and time spent queued behind another
request, and timeToFirstToken excludes stream batching.
Observing operations#
Pass observers to LlamaEngine to trace, measure or log its work. An observer
sees chat completions (create, createStructuredJson and
ChatSession.create), generate, embed, embedBatch
and model loads.
final class TimingObserver extends LlamaEngineObserver {
@override
LlamaOperationObserver? onStart(LlamaOperation operation) {
final name = switch (operation) {
LlamaChatOperation() => 'chat',
LlamaTextCompletionOperation() => 'text_completion',
LlamaEmbeddingsOperation() => 'embeddings',
LlamaModelLoadOperation() => 'model_load',
_ => 'other',
};
return _Timing('$name ${operation.model}', Stopwatch()..start());
}
}
final class _Timing extends LlamaOperationObserver {
_Timing(this.name, this.watch);
final String name;
final Stopwatch watch;
@override
void onEnd(LlamaOperationResult result) {
print('$name: ${watch.elapsed}, finish ${result.finishReason}, '
'tokens ${result.usage?.totalTokens}');
}
}
final engine = LlamaEngine(LlamaBackend(), observers: [TimingObserver()]);
-
A
createorgenerateoperation starts when its stream is listened to; the others start when their method is called. Every callback runs in the zone that called the engine method, so a tracer can read its parent context there. -
onChunkreceives eachcreatechunk andonTexteachgeneratepiece. -
onEndruns once, with the error, the cancel, or the finish reason and usage. Usage is reported where the finalcreatechunk carries it. A chat subscription cancelled after the final chunk ends completed, unlesscancelGenerationstopped it first. -
LlamaOperation.modelis the model'sgeneral.namemetadata, or else the last segment of the path or URL it was loaded from. It is null when that segment is empty or contains one of/ \ ? # @ ; & =, so the segment is never a directory path, URL query, fragment or userinfo. On the built-in backends,runtimeisLlamaRuntime.llamaCpporLlamaRuntime.liteRtLmfor operations after a model load, and null for the load itself. - Operations carry copies of the prompts and messages. Record them only when your users opt in.
-
Extend the observer classes rather than implementing them, and give a
switchover operations a default case: later versions may add callbacks and operation types. - An exception an observer throws is reported to the library logger as a warning and never reaches the caller. Without observers the engine does no observation work.
For a runnable observer-to-OpenTelemetry adapter and Langfuse/Grafana recipes, see Observability.
Thinking budget (native llama.cpp)#
For GGUF models with a thinking channel, ThinkingBudget maps to llama.cpp's
reasoning-budget sampler. It counts generated tokens inside each reasoning
block independently of maxTokens, which remains the cap for the entire
completion.
await for (final chunk in engine.create(
messages,
enableThinking: true,
params: const GenerationParams(
maxTokens: 512,
thinkingBudget: ThinkingBudget(maxTokens: 128),
),
)) {
final thinking = chunk.choices.first.delta.thinking;
final text = chunk.choices.first.delta.content;
// Render each channel independently.
}
engine.create(...) fills the start and end delimiters from recognized chat
templates. Raw engine.generate(...) calls must provide both delimiters in
ThinkingBudget. A budget of 0 immediately forces the end delimiter. This
is supported by native llama.cpp text generation only; LiteRT-LM and WebGPU
reject it explicitly, and it cannot be combined with speculative decoding.
Structured JSON output#
Use LlamaStructuredOutput when you want strict JSON plus final validation and
typed decoding. The helper builds the responseFormat map for grammar-capable
backends and validates the completed model output before returning your value.
class TicketClassification {
TicketClassification({required this.priority, required this.category});
final String priority;
final String category;
static TicketClassification fromJson(Map<String, dynamic> json) {
return TicketClassification(
priority: json['priority'] as String,
category: json['category'] as String,
);
}
}
final output = LlamaStructuredOutput<TicketClassification>.jsonSchema(
schema: const {
'type': 'object',
'properties': {
'priority': {
'type': 'string',
'enum': ['low', 'medium', 'high'],
},
'category': {'type': 'string'},
},
'required': ['priority', 'category'],
'additionalProperties': false,
},
decoder: TicketClassification.fromJson,
);
final classification = await engine.createStructuredJson(
[
LlamaChatMessage.fromText(
role: LlamaChatRole.user,
text: 'Classify this ticket: checkout fails with card declined.',
),
],
output: output,
params: const GenerationParams(maxTokens: 96, temp: 0),
);
Without the helper, pass responseFormat: {'type': 'json_object'} or
{'type': 'json_schema', 'json_schema': {'schema': <JSON schema>}}
to
engine.create(...). For live rendering, keep the stream returned by
engine.create(..., responseFormat: output.responseFormat) and finalize it with
await stream.parseStructuredJson(output). Validation is a final-output step
because partial stream chunks are often not valid JSON yet.
Supported schema features match the built-in JSON-schema-to-GBNF subset:
primitive types, objects with properties, required, and
additionalProperties, arrays with items or fixed prefixItems,
enum/const, local $ref, anyOf, oneOf,
allOf, minLength,
maxLength, minItems, and maxItems. Unsupported schemas fail before
generation. Annotation metadata such as title, description, and
default
is preserved but not enforced as a decoding constraint. Backends without
grammar constraints, including current LiteRT-LM native and web paths, still
fail early for strict structured output.
create(...) flow at a glance#
- Build your
List<LlamaChatMessage>. engine.create(...)runs template rendering/parity logic.- Effective stop sequences and grammar are applied to generation params.
- Backend token bytes are decoded and emitted as streaming chunks.
- Final parse resolves tool calls and stop reason.
Cancellation#
engine.cancelGeneration();
This cancels every create, generate and ChatSession.create stream that has
been listened to, including one still rendering its template or checking its
input: that stream ends without generating. A stream listened to after the
call is not affected. How quickly a running generation stops depends on the
backend.
Cancelling a stream's subscription also sends the cancel to its backend at once, even before the first token.
On native llama.cpp, a cancel during text prompt evaluation takes effect at the
next prompt micro-batch (ModelParams.microBatchSize tokens), or at the next
batch (ModelParams.batchSize tokens) with speculative decoding. A generation
started while a cancelled one is still stopping waits for it to stop, then
runs. Starting one while another is running and not cancelled throws
LlamaStateException.
Tokenization helpers#
final tokens = await engine.tokenize('hello world');
final text = await engine.detokenize(tokens);
final count = await engine.getTokenCount('hello world');
These helpers are useful for context budgeting and prompt diagnostics.
Next-token scores#
engine.scoreNextToken(...) evaluates a prompt and returns the
log-probabilities of the token that would follow it, without generating. Ask
for specific token ids with candidates, the most probable tokens with topK,
or both. Reading the probabilities of answer letters turns an
instruction-tuned model into a classifier:
final prompt = (await engine.chatTemplate([
const LlamaChatMessage.fromText(
role: LlamaChatRole.user,
text: 'Is "remind me to call mom at 5" a (A) reminder or (B) search? '
'Answer with the letter only.',
),
], enableThinking: false)).prompt;
final letters = [
for (final letter in ['A', 'B'])
(await engine.tokenize(letter, addSpecial: false)).single,
];
final scores = await engine.scoreNextToken(prompt, candidates: letters);
final probabilities = [for (final t in scores.candidates) t.probability];
The values are a softmax over the raw logits at the last prompt position, the
same as llama-server's n_probs; sampling settings do not apply. The prompt is
tokenized like a generate prompt, and a prefix shared with the previous
prompt is reused unless reusePromptPrefix is false. Check
engine.supportsNextTokenScoring first: native llama.cpp and WebGPU bridge
assets v0.1.52+ support it; LiteRT-LM and older bridge assets report false
and throw LlamaUnsupportedException.