On-device text to speech
Generate complete PCM and WAV audio with the experimental typed llama.cpp Qwen3-TTS API on native and WebGPU runtimes.
On this page
TextToSpeechEngine is the typed API for speech synthesis. It is separate from
LlamaEngine.create because audio sample metadata, speaker reference input,
progress, cancellation, and output buffering are not chat-token semantics.
The first implementation is experimental and deliberately narrow. It supports
Qwen3-TTS through native llama.cpp or WebGPU bridge assets v0.1.33+ with the
matching audio-generation projector. It reports prompt/frame progress, then
returns one complete 24 kHz mono float32 PCM buffer. It does not stream playable
audio chunks yet.
Current support matrix#
| Runtime | Typed TextToSpeechEngine | Output |
|---|---|---|
| Native llama.cpp / GGUF | Experimental Qwen3-TTS adapter | Complete float32 PCM; WAV helper |
| WebGPU / GGUF | Experimental Qwen3-TTS adapter with bridge assets v0.1.33+ |
Complete float32 PCM; WAV helper |
Native LiteRT-LM / .litertlm |
Unsupported by the pinned native artifact | None |
| LiteRT-LM Web | Unsupported | None |
Load Qwen3-TTS#
Use a matching model and projector pair. The chat example pins the Q4_K_M base
model and Q8_0 projector from
ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF.
A TextToSpeechModel names the model and projector, each a ModelSource,
and the adapter that drives them. Native runtimes download and cache remote
sources before loading their local files; Web passes the same sources to the
browser runtime and its cache.
const revision = 'ca27d74bc954b73dadab5b71ca265d87fc861a7c';
final synthesizer = await TextToSpeechEngine.load(
TextToSpeechModel(
ModelSource.huggingFace(
repoId: 'ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF',
revision: revision,
filePath: 'Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf',
),
projector: ModelSource.huggingFace(
repoId: 'ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF',
revision: revision,
filePath: 'mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf',
),
adapter: const Qwen3TtsAdapter(),
),
);
try {
final result = await synthesizer.synthesizeOnce(
const TextToSpeechRequest(text: 'Hello from llamadart.', language: 'en'),
);
print(result.duration);
} finally {
await synthesizer.dispose();
}
load creates a LlamaEngine, loads the model with params: and then the
projector, checks capabilities, and throws LlamaUnsupportedException
with
the reason when they cannot synthesize speech with the adapter. Loading a
projector alone does not prove that the active native or Web runtime exports
the required TTS ABI or that the projector matches the model. download:
takes ModelLoadOptions for every remote file, onProgress: reports both
files together, store: takes a ModelFileStore with your own resolver or
download manager, and backend: the LlamaBackend (by default
LlamaBackend()). Both files resolve, model first, before anything loads;
a local file takes only the cancel token. The bearer token and headers are
never sent across hosts: when they are set and the two remote files are on
different origins (scheme, host and port), load throws
LlamaArgumentException before downloading from another host. ModelLoadOptions.sha256
throws
LlamaUnsupportedException, since one checksum cannot cover both files. The
load is atomic: when it throws, the engine is disposed, and downloaded files
stay in the cache.
The synthesizer owns the engine load created and a backend: you passed,
and dispose(), or a failed load, disposes both.
To share a LlamaEngine you loaded yourself, attach the adapter instead and
check capabilities yourself; dispose() then leaves your engine loaded:
final synthesizer = TextToSpeechEngine.attach(
engine,
adapter: const Qwen3TtsAdapter(),
);
final capabilities = await synthesizer.capabilities;
if (!capabilities.isSupported) {
throw StateError(capabilities.unsupportedReason!);
}
The runtime generates the audio itself and reports which audio-generation
model it loaded, so an adapter can target only a model the runtime supports.
Qwen3TtsAdapter accepts Qwen3-TTS and maps request languages to its codes.
TextToSpeechAdapter is open for another family once a runtime reports it:
supportsModel accepts the BackendTextToSpeechModel it drives,
supportedLanguages lists its codes, and normalizeLanguage maps a request
language to one of them.
Synthesize and save WAV#
final task = await synthesizer.synthesize(
const TextToSpeechRequest(
text: 'Hello from llamadart.',
language: 'English',
),
);
task.events.listen((event) {
if (event is TextToSpeechProgressEvent) {
print('Generated ${event.framesGenerated} frames');
}
});
try {
final result = await task.result;
final wavBytes = result.toWavBytes();
print('${result.duration}: ${wavBytes.length} WAV bytes');
} on LlamaException catch (error) {
print('Synthesis failed or was cancelled: $error');
}
synthesize throws typed validation, state, or unsupported errors when
preflight fails before a task starts. After startup, the single-subscription
event stream carries progress, then one TextToSpeechFinalEvent with the PCM,
and never emits an error. task.result returns the audio or throws the
failure, or LlamaStateException when the task is cancelled; task.done
reports the same outcome as a TextToSpeechCompletion and never throws.
synthesizeOnce skips the events: it is
(await synthesizer.synthesize(request)).result.
For models that advertise speaker-reference support, encoded bytes are the
portable representation. In this example, referenceWavBytes is a
Uint8List obtained through the host application's file picker or recorder:
final task = await synthesizer.synthesize(
TextToSpeechRequest(
text: 'This utterance uses the supplied reference voice.',
language: 'English',
speakerReference: SpeechAudioBytesInput(referenceWavBytes),
),
);
Native applications may alternatively use
SpeechAudioFileInput('/recordings/reference.wav'). Browser runtimes cannot
read arbitrary local filesystem paths.
For Qwen3-TTS, the canonical language codes are zh, en, ja, ko,
de,
fr, ru, pt, es, and it. Common English names are normalized to those
codes, so language: 'English' is equivalent to language: 'en'. Other values
fail during typed preflight instead of reaching the backend model as an invalid
prompt.
Treat reference recordings as sensitive input. The typed API does not retain them after the backend request completes, but application code remains responsible for its own files, byte buffers, permissions, and disclosures.
Cancellation, concurrency, and buffering#
Call task.cancel() to request cooperative cancellation of that synthesis.
Cancelling only the event-stream subscription does not cancel synthesis. On native llama.cpp with
llamadart-native v0.4.1-1 or later, which the default pin meets, a cancel
stops a Qwen3-TTS audio decode in progress at its next chunk boundary. Older
runtimes finish the native step first, which can include the whole decode.
LlamaEngine.unloadModel() and dispose() cancel an active synthesis the same
way, and its task reports cancelled.
TextToSpeechEngine.dispose() cancels a running task, waits for it to stop,
and disposes the engine that load created; an attached engine stays loaded.
Calling it again is safe, and isDisposed reports it. After dispose(),
synthesize throws LlamaStateException and capabilities
reports
unsupported.
All typed STT and TTS wrappers over one LlamaEngine share a one-task speech
lease. Do not run chat generation, transcription, or another synthesis on the
same engine until the active speech task completes.
The current native and Web wrappers produce PCM only after all requested audio-codec
frames have been generated. supportsIncrementalAudio and
supportsOutputBackpressure are therefore false. Progress events are useful
for status and cancellation, but are not playable audio chunks.
The chat app has a Qwen3-TTS mode built on this API.
Known limits#
- The first backend returns complete 24 kHz mono output only; no incremental audio chunks, timestamps, or output backpressure are available.
- Qwen3-TTS and llama.cpp audio generation are experimental. Voice quality, latency, supported reference formats, and accelerator behavior remain model/device dependent.
-
Web requires published bridge assets
v0.1.33+, WebAssembly memory64 for the pinned roughly 1.48 GB model/projector pair, and a browser/device with enough memory. Older bridge assets fail capability discovery clearly. -
The chat example pins
v0.1.54, whose bridge retries a synthesis once on the main thread, more slowly, after eligible WebGPU errors and worker timeouts; see WebGPU bridge fallback behavior. - Web speaker references are selected-file bytes only; microphone speaker recording remains a native chat-example feature.