Multimodal input: vision and audio
Send images and audio to multimodal models: GGUF model plus projector pairs, LiteRT-LM bundles, capability checks and web notes.
On this page
Multimodal inference requires a model/runtime path that supports vision or
audio behavior. GGUF models use a model plus projector pair. Native
.litertlm bundles use LiteRT-LM's bundle-native media processors and do not
load a separate projector.
GGUF projector flow#
await engine.loadModel('/path/to/model.gguf');
await engine.loadMultimodalProjector('/path/to/mmproj.gguf');
Use source-based loading when the projector should be resolved, downloaded, and cached like a remote model source:
await engine.loadModelSource(
ModelSource.parse('hf://owner/repo/model-Q4_K_M.gguf'),
);
await engine.loadMultimodalProjectorSource(
ModelSource.parse('hf://owner/repo/mmproj.gguf'),
);
Native/file-backed backends download remote projectors through the configured
ModelDownloadManager before loading the cached local path. URL-loading web
backends support remote unauthenticated projector URLs directly and reject local
filesystem paths or options that require native cache IO such as auth headers,
checksum verification, explicit cache policy changes, custom cache directories,
disabled resume, and custom retry counts.
Projector offload follows effective model-load configuration. If model loading
is CPU-only (preferredBackend: GpuBackend.cpu or gpuLayers: 0), projector
initialization also runs CPU-only.
unloadModel() and dispose() release the projector with the model.
unloadMultimodalProjector() releases only the projector and keeps the model
loaded. Loading another projector replaces the active one.
LiteRT-LM bundle flow#
await engine.loadModel('/path/to/model.litertlm');
Native LiteRT-LM accepts LlamaImageContent and LlamaAudioContent backed by
local paths or encoded media bytes. Remote image URLs and raw PCM
Float32List audio samples are rejected before native generation because the
current LiteRT-LM C message loader expects a local path or base64 blob.
Native LiteRT-LM starts audio preprocessing on the selected backend (CPU when NPU is selected). If that fails, it retries on CPU and keeps CPU audio for the loaded model; with the GPU backend, the Gemma 4 E2B bundle resolves to GPU text and vision with CPU audio.
Build multimodal message#
final message = LlamaChatMessage.withContent(
role: LlamaChatRole.user,
content: const [
LlamaImageContent(path: '/path/to/image.jpg'),
LlamaTextContent('Describe this image in one sentence.'),
],
);
await for (final chunk in engine.create([message])) {
final text = chunk.choices.first.delta.content;
if (text != null) {
print(text);
}
}
Capability checks#
final supportsVision = await engine.supportsVision;
final supportsAudio = await engine.supportsAudio;
final supportsVideo = await engine.supportsVideo; // false in current releases
Always prefer these runtime checks over model-card assumptions. A loaded
projector can expose only a subset of the family-level modalities. The current
Gemma 4 E2B GGUF projector path in native llama.cpp mtmd reports both vision
and audio support; audio remains experimental upstream. Web continues to rely
on the loaded bridge's runtime capability report.
Native .litertlm bundles process media themselves. loadMultimodalProjector*,
supportsVision and supportsAudio apply only to GGUF projectors.
Video isn't supported; send extracted frames as LlamaImageContent.
LlamaVideoContent fails with LlamaUnsupportedException.
LlamaAudioContent is generic audio input routed through normal generation; it
does not by itself provide a transcript contract. For typed transcription, see
Speech to Text.
Web notes#
- Web uses bridge runtime paths.
- Multimodal projector loading on web is URL-based.
-
loadMultimodalProjectorSource(...)accepts remote unauthenticated projector URLs on URL-loading web backends; source options that require the native download/cache manager are unsupported there. - Local file path media inputs are native-first; web flows use browser file bytes/URLs.
- LiteRT-LM web through
@litert-lm/coreremains text-only inllamadart.
Tuning notes#
- Start with smaller images or audio inputs before changing backend settings.
-
The example chat app caps picked image inputs to a
384pxmax edge before staging them, but directLlamaImageContent(...)usage does not resize media for you. -
Projector load success does not imply every modality is available. Re-check
engine.supportsVision/engine.supportsAudioafter loadingmmproj. - Keep context and generation budgets tighter than your text-only defaults.
- Follow-up turns after an image can still overflow the active context window if conversation history grows too large.
- If multimodal is unstable on GPU, establish a working CPU baseline first.
- For broader tuning workflow and diagnostics guidance, see Performance Tuning.