Multimodal input: vision and audio
Send images and audio to multimodal models: GGUF model plus projector pairs, LiteRT-LM bundles, capability checks and web notes.
On this page
Multimodal inference requires a model/runtime path that supports vision or
audio behavior. GGUF models use a model plus projector pair. Native
.litertlm bundles use LiteRT-LM's bundle-native media processors and do not
load a separate projector.
GGUF projector flow#
await engine.loadModel('/path/to/model.gguf');
await engine.loadMultimodalProjector('/path/to/mmproj.gguf');
Use source-based loading when the projector should be resolved, downloaded, and cached like a remote model source:
await engine.loadModelSource(
ModelSource.parse('hf://owner/repo/model-Q4_K_M.gguf'),
);
await engine.loadMultimodalProjectorSource(
ModelSource.parse('hf://owner/repo/mmproj.gguf'),
);
Native/file-backed backends download remote projectors through the configured
ModelDownloadManager before loading the cached local path. URL-loading web
backends support remote unauthenticated projector URLs directly and reject local
filesystem paths or options that require native cache IO such as auth headers,
a cancelToken, checksum verification, explicit cache policy changes, custom
cache directories, disabled resume, and custom retry counts.
Projector offload follows effective model-load configuration. If model loading
is CPU-only (preferredBackend: GpuBackend.cpu or gpuLayers: 0), projector
initialization also runs CPU-only.
unloadModel() and dispose() release the projector with the model.
unloadMultimodalProjector() releases only the projector and keeps the model
loaded. Loading another projector replaces the active one.
LiteRT-LM bundle flow#
await engine.loadModel('/path/to/model.litertlm');
Native LiteRT-LM accepts LlamaImageContent and LlamaAudioContent backed by
local paths or encoded media bytes. Remote image URLs and raw PCM
Float32List audio samples are rejected with LlamaUnsupportedException
before native generation because the current LiteRT-LM C message loader
expects a local path or base64 blob.
Native LiteRT-LM starts audio preprocessing on the selected backend (CPU when NPU is selected). If that fails, it retries on CPU and keeps CPU audio for the loaded model; with the GPU backend, the Gemma 4 E2B bundle resolves to GPU text and vision with CPU audio.
Build multimodal message#
final message = LlamaChatMessage.withContent(
role: LlamaChatRole.user,
content: const [
LlamaImageContent(path: '/path/to/image.jpg'),
LlamaTextContent('Describe this image in one sentence.'),
],
);
await for (final chunk in engine.create([message])) {
final text = chunk.choices.first.delta.content;
if (text != null) {
print(text);
}
}
On native llama.cpp, a request that carries image or audio parts and sets
GenerationParams.thinkingBudget or speculative decoding
(speculativeDecoding or speculativeDecodingConfig) throws
LlamaUnsupportedException; both are text-only there. Leave them unset for
media turns.
Capability checks#
final supportsVision = await engine.supportsVision;
final supportsAudio = await engine.supportsAudio;
final supportsVideo = await engine.supportsVideo; // false in current releases
Always prefer these runtime checks over model-card assumptions. With no
projector loaded, a GGUF model (native llama.cpp or WebGPU) rejects image
and audio parts with LlamaUnsupportedException instead of answering from the
text alone. A loaded projector can expose only a subset of the family-level
modalities. The current Gemma 4 E2B GGUF projector path in native llama.cpp
mtmd reports both vision and audio support; audio remains experimental
upstream. Web continues to rely
on the loaded bridge's runtime capability report.
Native .litertlm bundles process media themselves. loadMultimodalProjector*,
supportsVision and supportsAudio apply only to GGUF projectors.
Video isn't supported; send extracted frames as LlamaImageContent.
LlamaVideoContent fails with LlamaUnsupportedException.
LlamaAudioContent is generic audio input routed through normal generation; it
does not by itself provide a transcript contract. For typed transcription, see
Speech to Text.
Web notes#
- Web uses bridge runtime paths.
- Multimodal projector loading on web is URL-based.
-
loadMultimodalProjectorSource(...)accepts remote unauthenticated projector URLs on URL-loading web backends; source options that require the native download/cache manager are unsupported there. -
Local file path media inputs are native-first; web flows use browser file
bytes/URLs.
LlamaImageContent.urlis read only by the web bridge: nativellama.cppthrowsLlamaUnsupportedExceptionfor it, so download the image and pass its bytes. - LiteRT-LM web through
@litert-lm/coreremains text-only inllamadart.
Tuning notes#
- Start with smaller images or audio inputs before changing backend settings.
-
The example chat app caps picked image inputs to a
384pxmax edge before staging them, but directLlamaImageContent(...)usage does not resize media for you. -
Projector load success does not imply every modality is available. Re-check
engine.supportsVision/engine.supportsAudioafter loadingmmproj. - Keep context and generation budgets tighter than your text-only defaults.
- Follow-up turns after an image can still overflow the active context window if conversation history grows too large.
- If multimodal is unstable on GPU, establish a working CPU baseline first.
- For broader tuning workflow and diagnostics guidance, see Performance Tuning.