On-device speech to text
Transcribe audio files with Qwen3-ASR or stream live PCM with LiteRT-LM through the experimental typed SpeechToTextEngine API.
On this page
SpeechToTextEngine is the typed, experimental API for speech recognition. It
is separate from LlamaEngine because transcript events, cancellation, and
audio metadata have a different contract from chat-completion tokens.
Recognition quality and language behavior are model dependent. For speech
synthesis, see Text to Speech.
Choose an approach#
A SpeechToTextModel names the model's files, each a ModelSource, and the
adapter that drives it. The adapter picks the runtime.
| Approach | API | Runtimes | Input | Output |
|---|---|---|---|---|
| Qwen3-ASR, whole file | SpeechToTextEngine.load with Qwen3AsrAdapter, or attach |
Native llama.cpp; WebGPU with bridge assets v0.1.30+ |
A complete WAV, MP3, or FLAC file or bytes; bytes only on Web | One final transcript |
| Another audio chat model, whole file | load or attach with your own SpeechToTextPromptAdapter |
As Qwen3-ASR | As Qwen3-ASR | One final transcript |
| LiteRT-LM, live streaming | SpeechToTextEngine.load with LiteRtLmAsrAdapter |
Native LiteRT-LM, CPU only | Mono 16 kHz float PCM, pushed incrementally or as one buffer | Replaceable partial text, then a final transcript |
| Generic audio chat | LlamaAudioContent in engine.create |
Audio-capable GGUF projectors; .litertlm bundles with audio |
Audio as a chat content part | Ordinary chat output, no transcript contract |
LiteRT-LM Web supports none of these. A prompt adapter reports
SpeechToTextImplementation.multimodalPromptAdapter: it runs whole-audio
generation on a LlamaEngine. LiteRtLmAsrAdapter reports
SpeechToTextImplementation.dedicatedBackend and does not use a chat model.
The adapter is an explicit declaration: an audio-understanding model is not advertised as ASR merely because it accepts audio.
Load or attach#
SpeechToTextEngine.load downloads each remote file into the model cache,
loads the model, and returns a recognizer that owns what it loaded.
dispose() releases it. SpeechToTextEngine.attach borrows a
LlamaEngine
you loaded yourself, for example to chat with the same model, and its
dispose() leaves that engine loaded.
load(model, ...) | attach(engine, adapter: ...) | |
|---|---|---|
| Adapters | Any | SpeechToTextPromptAdapter only |
| Who loads the files | The recognizer, from ModelSources | You |
| Capability check | load throws LlamaUnsupportedException |
Read capabilities yourself |
dispose() |
Cancels the task and disposes the engine | Cancels the task; the engine stays loaded |
load takes download: (ModelLoadOptions for every remote file: cache
policy and directory, authentication, resume, retries and cancel token),
onProgress: (one combined progress for all files), store: (a
ModelFileStore with your own resolver or download manager), and, for a
prompt adapter, params: (ModelParams) and backend:
(by default
LlamaBackend()). Every file resolves, main file first, before anything
loads. A local file takes only the cancel token. The bearer token and headers
are never sent across hosts: when they are set and the remote files span more
than one origin (scheme, host and port), load throws
LlamaArgumentException before downloading from another host, so load such
files from one host or leave the credentials unset. ModelLoadOptions.sha256
throws
LlamaUnsupportedException for a model of more than one file. The load is
atomic: when it throws, nothing stays loaded, and downloaded files stay in the
cache.
load takes ownership of a backend: you pass: the recognizer's dispose(),
or a failed load, disposes it with the engine, so pass a backend that nothing
else uses.
Transcribe a file with Qwen3-ASR#
Use a matching model and multimodal projector pair. llama.cpp documents
Qwen3-ASR in its
multimodal model list,
and published GGUF pairs are available from
ggml-org/Qwen3-ASR-0.6B-GGUF.
Sources may be local paths, HTTP(S) URLs or hf://owner/repo/file references.
final recognizer = await SpeechToTextEngine.load(
SpeechToTextModel(
ModelSource.path('/models/Qwen3-ASR-0.6B-Q8_0.gguf'),
projector: ModelSource.path('/models/mmproj-Qwen3-ASR-0.6B-Q8_0.gguf'),
adapter: const Qwen3AsrAdapter(),
),
);
try {
final result = await recognizer.transcribeOnce(
const SpeechToTextRequest(
audio: SpeechAudioFileInput('/recordings/meeting.wav'),
contextPrompt: 'llamadart, Qwen3-ASR',
),
);
print(result.text);
} finally {
await recognizer.dispose();
}
load checks capabilities after the model and projector load and throws
LlamaUnsupportedException with the reason when they cannot recognize speech.
Projector load success alone does not prove audio support.
load rethrows what the projector load throws. On native llama.cpp that is
LlamaModelException for a projector the runtime rejects, such as the
Qwen3-TTS projector with the Qwen3-ASR model, and LlamaUnsupportedException
when the runtime lacks the mtmd functions. On Web, it is LlamaModelException
when the bridge cannot fetch or load the projector.
With an engine you loaded, attach the adapter and check capabilities
yourself:
final recognizer = SpeechToTextEngine.attach(
engine,
adapter: const Qwen3AsrAdapter(),
);
final capabilities = await recognizer.capabilities;
if (!capabilities.isSupported) {
throw StateError(capabilities.unsupportedReason!);
}
transcribeOnce returns the final result, and throws the task's failure, or
LlamaStateException when the task is cancelled. For events, use
transcribe:
final task = await recognizer.transcribe(
const SpeechToTextRequest(
audio: SpeechAudioFileInput('/recordings/meeting.wav'),
contextPrompt: 'llamadart, Qwen3-ASR',
),
);
task.events.listen((event) {
if (event is SpeechToTextPartialEvent) {
print('partial: ${event.text}');
}
});
try {
final result = await task.result;
print(result.text);
} on LlamaException catch (error) {
print('Recognition failed or was cancelled: $error');
}
transcribe itself throws typed input, state, or unsupported errors when
preflight fails before a task can start. After startup, events is a
single-subscription stream of progress that never emits an error: a
prompt adapter emits one SpeechToTextFinalEvent, and LiteRT-LM emits
partial events first. task.result returns the transcript or throws the
failure, or LlamaStateException when the task is cancelled; task.done
reports the same outcome as a SpeechToTextCompletion and never throws.
Add a model family#
To recognize speech with another audio chat model on LlamaEngine, extend
SpeechToTextPromptAdapter. The recognizer sends one user turn: the text from
promptFor, then the audio. Natively the turn goes through the model's chat
template with thinking disabled; on Web the bridge takes the text as the raw
prompt with the audio bytes. Generation is greedy and stops at
maxOutputTokens. parseTranscript turns the complete output into a
SpeechToTextTranscript; an empty transcript fails the task.
class MyAsrAdapter extends SpeechToTextPromptAdapter {
const MyAsrAdapter();
@override
String get name => 'My-ASR';
@override
bool get supportsLanguageHints => true;
@override
String promptFor(SpeechToTextRequest request) {
final language = request.languageHint;
return language == null
? 'Transcribe the audio.'
: 'Transcribe the audio. It is in $language.';
}
@override
SpeechToTextTranscript parseTranscript(String output) =>
SpeechToTextTranscript(output.trim());
}
final recognizer = await SpeechToTextEngine.load(
SpeechToTextModel(
ModelSource.parse('hf://owner/repo/asr-model.gguf'),
projector: ModelSource.parse('hf://owner/repo/mmproj-asr-model.gguf'),
adapter: const MyAsrAdapter(),
),
);
supportsLanguageHints defaults to false, and the recognizer then rejects a
languageHint with LlamaUnsupportedException. supportsContextPrompt
defaults to true; supportsLanguageDetection defaults to false and says
whether parseTranscript reports SpeechToTextTranscript.language.
capabilities reports these flags. The recognizer checks only that the loaded
projector supports audio; whether the prompt suits the model is the adapter's
responsibility.
Stream live audio with LiteRT-LM#
LiteRT-LM's dedicated ASR engines (added in LiteRT-LM v0.16) consume PCM
windows instead of an audio part in normal chat. Load the model with its
tokenizer and a LiteRtLmAsrAdapter for the model family's preset, start a
stream, and await every input push so bounded native backpressure can
throttle the producer.
final recognizer = await SpeechToTextEngine.load(
SpeechToTextModel(
ModelSource.path('/models/moonshine_tiny.tflite'),
tokenizer: ModelSource.path('/models/tokenizer.json'),
adapter: const LiteRtLmAsrAdapter(LiteRtLmAsrModelPreset.moonshineTiny),
),
);
final session = await recognizer.startStream();
final events = session.events.listen(
(event) {
if (event is SpeechToTextPartialEvent) {
print('stable=${event.confirmedText} pending=${event.pendingText}');
} else if (event is SpeechToTextFinalEvent) {
print('final=${event.result.text}');
}
},
// A session's events report a failure as an error too; done carries it.
onError: (Object _) {},
);
for (final chunk in mono16KhzFloatPcmChunks) {
await session.addPcm(chunk);
}
await session.finish();
final completion = await session.done;
await events.cancel();
print(completion.state);
await recognizer.dispose();
load probes the LiteRT-LM ASR runtime before it downloads anything and
throws LlamaUnsupportedException when the runtime is unavailable, including
on Web. It then resolves the model and tokenizer, each a local path, an
HTTP(S) URL or a Hugging Face file, to local files, model first, downloading
remote ones into the model cache as described under
Load or attach; each transcribe or startStream
starts its
own native session on them. The
adapter carries the runtime settings (backend, numberOfThreads,
maxBufferedAudio, overlapRatio, and libraryPath
for local validation),
so params: and backend: must be null. A LiteRtLmAsrAdapter
model takes
no projector, and a prompt adapter model no tokenizer.
The session runs synchronous inference in a worker isolate. confirmedText is
stable, while pendingText may change after the next inference window.
finish flushes a partial final window.
LiteRtLmAsrBackend.cpu is the only backend. Metadata presets cover Parakeet
TDT, Parakeet CTC, Moonshine Tiny, Whisper Tiny, and Qwen3-ASR 0.6B, but
callers must supply a matching model and tokenizer. The API does not capture a
microphone or resample audio. Advanced callers can use LiteRtLmRuntimeClient
and LiteRtLmAsrRuntimeSession from package:llamadart/backend.dart
directly
with local files, but those synchronous calls must not run on a Flutter UI
isolate.
Cancel, dispose and concurrency#
final task = await recognizer.transcribe(request);
// Later:
task.cancel();
final completion = await task.done;
assert(completion.state == SpeechToTextCompletionState.cancelled);
Cancellation is cooperative: task.cancel() for whole-input recognition, or
await session.cancel() for a LiteRT-LM session, which stops between native
windows. task.cancel() stops only that task: on a prompt adapter it cancels
the task's own generation and no other request on the same LlamaEngine. A
session's cancel() returns a future because it ends a live input
stream and releases the native recognizer; a session has no single result. Cancelling or pausing an event subscription neither cancels nor
throttles native inference; LiteRT-LM producers must await addPcm for input
backpressure.
dispose() cancels a running task or stream, waits for it to stop, and
disposes the engine that load created. Calling it again is safe, and
isDisposed reports it. After dispose(), transcribe
and startStream
throw LlamaStateException and capabilities reports unsupported.
All prompt-adapter recognizers and text-to-speech wrappers over one
LlamaEngine share a one-task lease. Direct LlamaEngine.create
calls on a
borrowed engine must not run during a task. A LiteRT-LM recognizer allows one
active task per SpeechToTextEngine instance.
Input formats and length#
Qwen3-ASR recognition is validated up to 30 seconds per input. Longer inputs can drop or repeat sentences without reaching any limit, so split longer recordings into windows of at most 30 seconds. Built-in windowing is tracked in #327.
A Qwen3-ASR prompt grows by about 13 tokens per second of audio, and the
transcript shares the same context. On native llama.cpp, a prompt-adapter task
that reaches the context size or maxOutputTokens before the transcript ends
fails with LlamaSpeechTranscriptTruncatedException. Its limit
names the
limit that stopped recognition and partialTranscript holds the text produced
before it. A task whose transcript is empty, for example from silent input,
fails with LlamaSpeechException.
Native llama.cpp accepts WAV, MP3, and FLAC file or byte inputs. Raw PCM is
unsupported for prompt adapters because projector sample rates are
model-specific. LiteRT-LM accepts SpeechAudioPcmInput for a complete mono
16 kHz normalized Float32List buffer, or the incremental session above.
SpeechAudioFormat carries optional encoding and MIME metadata. Final results
reserve segment and word timing, confidence, and speaker fields for future
backends; prompt adapters return one untimed segment.
Web#
Web runs prompt adapters, such as Qwen3-ASR, through WebGPU bridge assets
v0.1.30+. The hosted chat app derives the capability from the immutable
llama-web-bridge-assets tag; custom hosts can set
window.__llamadartBridgeSpeechToTextSupported before the backend is created.
An older bridge, no loaded projector, a projector without audio support, or a
failed runtime audio probe makes load throw LlamaUnsupportedException, and
leaves an attached recognizer's capabilities.isSupported false, with an
actionable reason. On Web, load passes the sources to the browser runtime,
which fetches each file itself: a ModelSource.path is a URL relative to the
document, or a blob: URL, and download must stay at its defaults.
WebGPU accepts encoded WAV, MP3, and FLAC bytes. Read the selected file into
memory and pass SpeechAudioBytesInput with a SpeechAudioFormat
whose
encoding is 'wav', 'mp3' or 'flac'; local filesystem paths, other
encodings, raw PCM, and bytes without that metadata are rejected. The bridge
reports truncation through onUsage when its supportsCompletionUsage
probe is true (v0.1.54+). A cut-off Web transcript then fails with
LlamaSpeechTranscriptTruncatedException, whose limit is
LlamaSpeechTranscriptLimit.runtime: the bridge combines output-token,
context and media caps, so the error cannot name the exact cause. Its
partialTranscript holds the text produced before the limit. Older bridges
without this signal retain their existing completion behavior. The browser
needs enough memory for the roughly 1.02 GB Qwen3-ASR
0.6B Q8_0 model/projector pair.
Known limits#
-
Validated with Qwen3-ASR 0.6B Q8_0 on WAV up to 33 s, and on MP3 and FLAC
copies of the 11 s
jfk.wavon native macOS and in headless Chromium. The audio prompt grows with duration (3,890 tokens for 297 s in #636). -
Qwen3-ASR may emit a leading
language English<asr_text>marker.Qwen3AsrAdapterstrips that marker, but does not expose it as reliable detected-language metadata until language behavior has a dedicated validation contract. - There are no word/segment timestamps, confidence scores, or speaker diarization. Incremental audio and partial text are LiteRT-LM-only.
- Inference backend correctness and performance remain device dependent; establish a CPU baseline before claiming GPU support for a deployment.
- The chat app shows file transcription, microphone capture, and live dictation built on this API.