On-device text to speech

Generate complete PCM and WAV audio with the experimental typed llama.cpp Qwen3-TTS API on native and WebGPU runtimes.

On this page

TextToSpeechEngine is the typed API for speech synthesis. It is separate from LlamaEngine.create because audio sample metadata, speaker reference input, progress, cancellation, and output buffering are not chat-token semantics.

The first implementation is experimental and deliberately narrow. It supports Qwen3-TTS through native llama.cpp or WebGPU bridge assets v0.1.33+ with the matching audio-generation projector. It reports prompt/frame progress, then returns one complete 24 kHz mono float32 PCM buffer. It does not stream playable audio chunks yet.

Current support matrix#

RuntimeTyped TextToSpeechEngineOutput
Native llama.cpp / GGUF Experimental Qwen3-TTS adapter Complete float32 PCM; WAV helper
WebGPU / GGUF Experimental Qwen3-TTS adapter with bridge assets v0.1.33+ Complete float32 PCM; WAV helper
Native LiteRT-LM / .litertlm Unsupported by the pinned native artifact None
LiteRT-LM WebUnsupportedNone

Load Qwen3-TTS#

Use a matching model and projector pair. The chat example pins the Q4_K_M base model and Q8_0 projector from ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF.

A TextToSpeechModel names the model and projector, each a ModelSource, and the adapter that drives them. Native runtimes download and cache remote sources before loading their local files; Web passes the same sources to the browser runtime and its cache.

const revision = 'ca27d74bc954b73dadab5b71ca265d87fc861a7c';
final synthesizer = await TextToSpeechEngine.load(
  TextToSpeechModel(
    ModelSource.huggingFace(
      repoId: 'ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF',
      revision: revision,
      filePath: 'Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf',
    ),
    projector: ModelSource.huggingFace(
      repoId: 'ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF',
      revision: revision,
      filePath: 'mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf',
    ),
    adapter: const Qwen3TtsAdapter(),
  ),
);
try {
  final result = await synthesizer.synthesizeOnce(
    const TextToSpeechRequest(text: 'Hello from llamadart.', language: 'en'),
  );
  print(result.duration);
} finally {
  await synthesizer.dispose();
}

load creates a LlamaEngine, loads the model with params: and then the projector, checks capabilities, and throws LlamaUnsupportedException with the reason when they cannot synthesize speech with the adapter. Loading a projector alone does not prove that the active native or Web runtime exports the required TTS ABI or that the projector matches the model. download: takes ModelLoadOptions for every remote file, onProgress: reports both files together, store: takes a ModelFileStore with your own resolver or download manager, and backend: the LlamaBackend (by default LlamaBackend()). Both files resolve, model first, before anything loads; a local file takes only the cancel token. The bearer token and headers are never sent across hosts: when they are set and the two remote files are on different origins (scheme, host and port), load throws LlamaArgumentException before downloading from another host. ModelLoadOptions.sha256 throws LlamaUnsupportedException, since one checksum cannot cover both files. The load is atomic: when it throws, the engine is disposed, and downloaded files stay in the cache.

The synthesizer owns the engine load created and a backend: you passed, and dispose(), or a failed load, disposes both. To share a LlamaEngine you loaded yourself, attach the adapter instead and check capabilities yourself; dispose() then leaves your engine loaded:

final synthesizer = TextToSpeechEngine.attach(
  engine,
  adapter: const Qwen3TtsAdapter(),
);
final capabilities = await synthesizer.capabilities;
if (!capabilities.isSupported) {
  throw StateError(capabilities.unsupportedReason!);
}

The runtime generates the audio itself and reports which audio-generation model it loaded, so an adapter can target only a model the runtime supports. Qwen3TtsAdapter accepts Qwen3-TTS and maps request languages to its codes. TextToSpeechAdapter is open for another family once a runtime reports it: supportsModel accepts the BackendTextToSpeechModel it drives, supportedLanguages lists its codes, and normalizeLanguage maps a request language to one of them.

Synthesize and save WAV#

final task = await synthesizer.synthesize(
  const TextToSpeechRequest(
    text: 'Hello from llamadart.',
    language: 'English',
  ),
);

task.events.listen((event) {
  if (event is TextToSpeechProgressEvent) {
    print('Generated ${event.framesGenerated} frames');
  }
});

try {
  final result = await task.result;
  final wavBytes = result.toWavBytes();
  print('${result.duration}: ${wavBytes.length} WAV bytes');
} on LlamaException catch (error) {
  print('Synthesis failed or was cancelled: $error');
}

synthesize throws typed validation, state, or unsupported errors when preflight fails before a task starts. After startup, the single-subscription event stream carries progress, then one TextToSpeechFinalEvent with the PCM, and never emits an error. task.result returns the audio or throws the failure, or LlamaStateException when the task is cancelled; task.done reports the same outcome as a TextToSpeechCompletion and never throws. synthesizeOnce skips the events: it is (await synthesizer.synthesize(request)).result.

For models that advertise speaker-reference support, encoded bytes are the portable representation. In this example, referenceWavBytes is a Uint8List obtained through the host application's file picker or recorder:

final task = await synthesizer.synthesize(
  TextToSpeechRequest(
    text: 'This utterance uses the supplied reference voice.',
    language: 'English',
    speakerReference: SpeechAudioBytesInput(referenceWavBytes),
  ),
);

Native applications may alternatively use SpeechAudioFileInput('/recordings/reference.wav'). Browser runtimes cannot read arbitrary local filesystem paths.

For Qwen3-TTS, the canonical language codes are zh, en, ja, ko, de, fr, ru, pt, es, and it. Common English names are normalized to those codes, so language: 'English' is equivalent to language: 'en'. Other values fail during typed preflight instead of reaching the backend model as an invalid prompt.

Treat reference recordings as sensitive input. The typed API does not retain them after the backend request completes, but application code remains responsible for its own files, byte buffers, permissions, and disclosures.

Cancellation, concurrency, and buffering#

Call task.cancel() to request cooperative cancellation of that synthesis. Cancelling only the event-stream subscription does not cancel synthesis. On native llama.cpp with llamadart-native v0.4.1-1 or later, which the default pin meets, a cancel stops a Qwen3-TTS audio decode in progress at its next chunk boundary. Older runtimes finish the native step first, which can include the whole decode. LlamaEngine.unloadModel() and dispose() cancel an active synthesis the same way, and its task reports cancelled.

TextToSpeechEngine.dispose() cancels a running task, waits for it to stop, and disposes the engine that load created; an attached engine stays loaded. Calling it again is safe, and isDisposed reports it. After dispose(), synthesize throws LlamaStateException and capabilities reports unsupported.

All typed STT and TTS wrappers over one LlamaEngine share a one-task speech lease. Do not run chat generation, transcription, or another synthesis on the same engine until the active speech task completes.

The current native and Web wrappers produce PCM only after all requested audio-codec frames have been generated. supportsIncrementalAudio and supportsOutputBackpressure are therefore false. Progress events are useful for status and cancellation, but are not playable audio chunks.

The chat app has a Qwen3-TTS mode built on this API.

Known limits#

  • The first backend returns complete 24 kHz mono output only; no incremental audio chunks, timestamps, or output backpressure are available.
  • Qwen3-TTS and llama.cpp audio generation are experimental. Voice quality, latency, supported reference formats, and accelerator behavior remain model/device dependent.
  • Web requires published bridge assets v0.1.33+, WebAssembly memory64 for the pinned roughly 1.48 GB model/projector pair, and a browser/device with enough memory. Older bridge assets fail capability discovery clearly.
  • The chat example pins v0.1.54, whose bridge retries a synthesis once on the main thread, more slowly, after eligible WebGPU errors and worker timeouts; see WebGPU bridge fallback behavior.
  • Web speaker references are selected-file bytes only; microphone speaker recording remains a native chat-example feature.

Searches the latest release. Esc to close.