Platform & Backend Matrix
Check which native and web runtimes are supported by llamadart and how backend selection works per platform.
On this page
- Platform/architecture coverage
- Model format routing
- Configuring native runtime families
- Video input boundary
- LiteRT-LM runtime coverage (v0.16.0-native.2)
- Runtime capability notes
- Current llama.cpp module availability by bundle (v0.3.0)
- Selector names and aliases
- Configuring native backend modules
- Selection and fallback behavior
- Vulkan cooperative matrix driver crashes
- Related docs
This page combines platform support, runtime-family selection, and
backend-module configuration for
llamadart.
The native-assets hook currently pins llamadart-native tag
v0.3.0 and
litert-lm-native release v0.16.0-native.2 (hook/build.dart). Apps can
override the llama.cpp native GitHub source with
hooks.user_defines.llamadart.llamadart_native_tag and
hooks.user_defines.llamadart.llamadart_native_repository, or use a local
bundle source with hooks.user_defines.llamadart.llamadart_native_path. Module
availability below is for the pinned/default artifacts.
Speech support is narrower than general runtime availability. llama.cpp/GGUF
has an experimental whole-file SpeechToTextEngine adapter when the caller
explicitly selects the Qwen3-ASR profile and the loaded projector reports audio
capability. Native accepts WAV, MP3, and FLAC file or byte inputs. WebGPU bridge
assets v0.1.30+ opt into the validated Qwen3-ASR path with WAV bytes only;
older or custom runtimes stay unsupported unless the host explicitly declares
the capability. Native llama.cpp and WebGPU bridge assets v0.1.33+ also expose
experimental Qwen3-TTS synthesis through the separate typed
TextToSpeechEngine. Native LiteRT-LM v0.16 also
supports experimental CPU-only streaming ASR through
SpeechToTextEngine.liteRtLm, with the native session owned by a worker
isolate. LiteRT-LM Web does not expose typed speech. See the
speech recognition support matrix.
Available override tags are published on the
leehack/llamadart-native releases page
or via gh release list --repo leehack/llamadart-native --limit 20.
Accepted tags are stable vMAJOR.MINOR.PATCH, historical bNNNN, and existing
nightly artifacts. New nightly wrapper rebuilds use bNNNN-N; existing
bNNNN-llamadart.N artifacts remain consumption-only overrides. A stable
wrapper-only rebuild of upstream vM.m.p uses vM.m.p-N and preserves that
upstream prefix in the release manifest. New wrapper and nightly releases are
GitHub prereleases and require an explicit tag. Historical bNNNN and
bNNNN-llamadart.N artifacts may retain older prerelease=false
metadata, but
remain explicit compatibility inputs. Build-hook overrides must always name an
explicit tag; latest is limited to maintainer synchronization and
header/binding regeneration, where it accepts only an unsuffixed stable tag
regardless of GitHub metadata. Nightly cores and positive rebuild counters
reject leading zeros.
The selected release must include a bundle asset named
llamadart-native-<bundle>-<tag>.tar.gz for the target being built.
Native source overrides do not regenerate Dart FFI bindings or symbol lookups,
so the selected binary must remain ABI- and symbol-compatible with the default
runtime revision.
Platform/architecture coverage#
| Platform target | Hook bundle key | llamadart_native_backends configurable? |
Backend behavior | Status |
|---|---|---|---|---|
| Android arm64 | android-arm64 |
Yes | Defaults: cpu, vulkan (when present) |
Supported |
| Android x64 | android-x64 |
Yes | Defaults: cpu, vulkan (when present) |
Supported |
| Linux arm64 | linux-arm64 |
Yes | Defaults: cpu, vulkan (when present) |
Supported |
| Linux x64 | linux-x64 |
Yes | Defaults: cpu, vulkan (when present) |
Supported |
| Windows arm64 | windows-arm64 |
Yes | Defaults: cpu, vulkan (when present) |
Supported |
| Windows x64 | windows-x64 |
Yes | Defaults: cpu, vulkan (when present) |
Supported |
| iOS arm64 (device) | ios-arm64 |
No (fixed in hook) | Consolidated runtime: cpu, metal |
Supported |
| iOS arm64 (simulator) | ios-arm64-sim |
No (fixed in hook) | Consolidated runtime: cpu, metal |
Supported |
| iOS x86_64 (simulator) | ios-x86_64-sim |
No (fixed in hook) | Consolidated runtime: cpu, metal |
Supported |
| macOS arm64 | macos-arm64 |
No (fixed in hook) | Consolidated runtime: cpu, metal |
Supported |
| macOS x86_64 | macos-x86_64 |
No (fixed in hook) | Consolidated runtime: cpu, metal |
Supported |
| Web (browser) | N/A (JS bridge path) | N/A | Router: llama.cpp WebGPU/CPU for .gguf; LiteRT-LM JS for .litertlm URLs |
Experimental; see WebGPU Bridge and LiteRT-LM web notes below |
All iOS targets above require the consuming Flutter/Xcode project to use a
minimum deployment target of 16.4 or newer. Flutter macOS targets require
macOS 14.0 or newer. If an iOS app still uses CocoaPods, set the Podfile
platform to 16.4 or newer too. LiteRT-LM iOS Simulator artifacts are
arm64-only, so apps that include LiteRT-LM must exclude the x86_64 Simulator
architecture; llama.cpp remains available for x86_64 Simulator builds.
Model format routing#
LlamaBackend() routes by model file format:
-
.ggufand unknown extensions use llama.cpp. Native targets load the bundled native runtime; web targets use the WebGPU bridge router. -
Native
.litertlmpaths use LiteRT-LM and the companion runtime bundles fromlitert-lm-native. -
Web
.litertlmURLs use the browser LiteRT-LM backend, which wraps the official@litert-lm/coreJavaScript API. Apps can preloadwindow.LiteRtLmEngine = module.Engineor setwindow.__llamadartLiteRtLmModuleUrlto an@litert-lm/coreESM URL before loading the model.
Use the same high-level LlamaEngine, ModelSource, and download/cache APIs
for both formats. Native/file-backed targets cache remote .litertlm sources
before local load and can use ChatSession for multi-turn chat. LiteRT-LM web
currently forwards only single-turn text prompts through @litert-lm/core, so
it does not preserve ChatSession history, system prompts, or tool
declarations with native LiteRT-LM semantics yet.
For GGUF multimodal projectors, loadMultimodalProjectorSource(...) uses the
same ModelSource and native download/cache APIs as loadModelSource(...).
Native/file-backed targets load the cached local projector path. URL-loading web
targets accept remote unauthenticated projector URLs and reject local filesystem
sources plus options that require native cache IO, including auth headers,
checksum verification, explicit cache policy changes, custom cache directories,
disabled resume, and custom retry counts.
Select LiteRT-LM CPU/GPU/NPU with ModelParams.liteRtLmBackend.
LiteRtLmBackendPreference.auto currently maps to GPU on Android, iOS, macOS,
and web, and CPU on other LiteRT-LM targets. NPU selection is Android native
only; web rejects it explicitly.
Configuring native runtime families#
Use llamadart_native_runtimes to choose which native runtime families are
bundled:
llama_cpp: GGUF model support through llama.cpp.litert_lm:.litertlmmodel support through LiteRT-LM.
The v0.3.0 native llama.cpp pin (llama.cpp v0.3.0) includes BailingMoE3 and
GraniteSWA/GraniteMoeSWA model loading and LFM2 target/draft support for
DSpark speculative decoding. These architectures use the existing GGUF APIs;
LFM2 DSpark uses SpeculativeDecodingConfig.draftDspark(...). No
model-specific Dart preset is required, but representative target/draft and
backend validation remains necessary before enabling DSpark in production.
Video input boundary#
| Runtime path | Public video input | Current evidence |
|---|---|---|
| Native llama.cpp / GGUF | Not consumable |
The published
v0.3.0
archive exports upstream video helper symbols, but this release has not been qualified for end-to-end video input. The companion build does not opt into
LLAMA_SUBPROCESS
/
MTMD_VIDEO
or package FFmpeg/ffprobe; the public Dart path remains unsupported until matching native, packaging, and frame-lifecycle validation exists.
|
| Native LiteRT-LM | Not consumable | The public direct-media path accepts image/audio content only. |
| WebGPU / Web LiteRT-LM | Not consumable | No validated public Dart video transport or frame-lifetime contract exists. |
| Android / iOS | Not consumable | Explicitly unsupported pending native packaging and device validation. |
LlamaEngine.supportsVideo therefore returns false. Passing
LlamaVideoContent throws LlamaUnsupportedException with either the native
compile/dependency blocker or, for a custom video-enabled native build, the
remaining Dart frame-ingestion/lifetime blocker. Do not infer support from
exported mtmd_helper_video_* symbols.
Unset or empty config means all runtime families available for the target. Apps that only ship one model format can trim package size:
hooks:
user_defines:
llamadart:
llamadart_native_runtimes: [llama_cpp]
Per-platform overrides can use OS keys or the exact bundle keys from the tables on this page. Exact target keys override OS keys:
hooks:
user_defines:
llamadart:
llamadart_native_runtimes:
runtimes: [llama_cpp, litert_lm]
platforms:
ios: [llama_cpp]
macos: [llama_cpp, litert_lm]
android-arm64: [litert_lm]
linux-x64: [llama_cpp]
Accepted aliases include llama.cpp, gguf, litert, and litert-lm.
Use all or both to include every available runtime family for a target.
Explicitly selecting litert_lm for a target without a pinned LiteRT-LM
runtime fails during the build hook instead of producing an app that cannot
load .litertlm models.
LiteRT-LM runtime coverage (v0.16.0-native.2)#
| Platform target | LiteRT-LM bundle key | Selectable backends | Status |
|---|---|---|---|
| Android arm64 | android-arm64 |
cpu, gpu, npu |
CPU/GPU supported. NPU is selectable only for compatible device/model/runtime deployments and may require a packaged LiteRT dispatch directory. |
| Android x64 | android-x64 |
cpu, gpu, npu |
CPU/GPU supported for emulator/test targets; NPU is deployment-specific and not validated for generic emulators. |
| iOS arm64 (device) | ios-arm64 |
cpu, gpu |
Supported |
| iOS arm64 (simulator) | ios-arm64-sim |
cpu, gpu |
Supported |
| iOS x86_64 (simulator) | Not published | N/A | Unsupported; exclude litert_lm for this target |
| macOS arm64 | macos-arm64 |
cpu, gpu |
Supported |
| macOS x86_64 | macos-x64 |
cpu |
Supported; the published x64 bundle does not include the WebGPU companion libraries |
| Linux arm64 | linux-arm64 | cpu | Supported |
| Linux x64 | linux-x64 | cpu | Supported |
| Windows x64 | windows-x64 | cpu | Supported |
| Web (browser) | N/A (@litert-lm/core) |
cpu, gpu |
Experimental; web-compatible .litertlm URLs only |
LiteRT-LM does not currently expose embeddings, state persistence, or external
multimodal projector APIs through llamadart. On native LiteRT-LM targets,
LlamaImageContent / LlamaAudioContent path/blob inputs are routed through
the normal generation path for .litertlm bundles whose native
template/runtime supports media; remote URLs and raw PCM samples are rejected.
Native .litertlm loads can also pass one default-scale text LoRA adapter from
ModelParams.loras; runtime setLora / removeLora, adapter stacking,
adapter scaling, and LiteRT-LM web LoRA remain unsupported.
High-level thinking and tool-call parsing still run through LlamaEngine
for
compatible templates, but llama.cpp-style GBNF grammar constraints are not
supported for .litertlm generation. Native LiteRT-LM can opt into runtime
speculative decoding through GenerationParams.speculativeDecoding; Web
LiteRT-LM rejects that option until the browser runtime exposes an equivalent
control. Web LiteRT-LM also does not expose tokenizer operations or multimodal
inputs and is limited to single-turn text prompts, so it should not be treated
as a multi-turn ChatSession or tool-calling backend yet.
llamadart rejects unsupported operations explicitly for .litertlm
loads
instead of silently ignoring llama.cpp-only settings.
Native LiteRT-LM exposes these load-time runtime controls through
ModelParams. Nullable fields keep the pinned v0.16.0-native.2
runtime
default.
| Native C API | Dart field | Support decision |
|---|---|---|
litert_lm_engine_settings_set_num_threads |
numberOfThreads |
Exposed for native
.litertlm
;
0
keeps LiteRT-LM automatic thread selection.
|
litert_lm_session_config_set_lora_path |
loras |
Exposed for one default-scale text LoRA adapter at model load; multiple adapters, custom scales, and runtime LoRA updates are rejected. |
litert_lm_engine_settings_set_activation_data_type |
liteRtLmActivationDataType |
Exposed for native
.litertlm
; typed as
float32
,
float16
,
int16
, or
int8
.
|
litert_lm_engine_settings_set_prefill_chunk_size |
liteRtLmPrefillChunkSize |
Exposed for CPU dynamic models; positive values only. |
litert_lm_engine_settings_set_parallel_file_section_loading |
liteRtLmParallelFileSectionLoading |
Exposed as a nullable boolean; native default remains parallel loading. |
litert_lm_engine_settings_set_litert_dispatch_lib_dir |
liteRtLmDispatchLibDir |
Exposed for Android NPU deployments that need a packaged LiteRT dispatch directory. |
LiteRT-LM web rejects these native-only fields because @litert-lm/core does
not expose equivalent runtime controls.
Android NPU support is not implied by the backend selector alone. A target app
still needs a .litertlm model bundle and LiteRT dispatch library set that
support the device SoC. If native LiteRT-LM cannot create an NPU engine for the
device/model bundle, use cpu or gpu for that artifact.
Runtime capability notes#
-
LoRA adapters are supported on native llama.cpp/GGUF runtimes with the
complete metadata inspection ABI below. Activated LoRA (aLoRA) is not yet
supported: llamadart detects invocation-token metadata and rejects the
adapter before activation. Native overrides must export both
llama_adapter_get_alora_n_invocation_tokensandllama_adapter_get_alora_invocation_tokens; runtimes without the complete, compatible metadata inspection ABI fail closed withLlamaUnsupportedExceptionrather than risk applying aLoRA eagerly. -
Thinking budgets (
GenerationParams.thinkingBudget) use llama.cpp's reasoning-budget sampler on native text-only GGUF generation.engine.createresolves known template delimiters automatically; raw generation requires explicit delimiters. LiteRT-LM and WebGPU reject the setting, as does the llama.cpp speculative-decoding path. -
Experimental DSpark speculative decoding is available as an explicit
opt-in on native llama.cpp/GGUF through
SpeculativeDecodingConfig.draftDspark(...); it is never enabled by default. It maps to upstreamdraft-dspark, requires a compatible external draft GGUF, and remains subject to target/draft/backend parity, acceptance, and throughput validation. It is an external non-MTP draft-context strategy and requires at least theb10356-llamadart.1wrapper fix; the package-pinned default runtime satisfies that ABI. WebGPU and LiteRT-LM reject this llama.cpp-specific strategy explicitly. -
State persistence (
LlamaEngine.stateSaveFile(...)/stateLoadFile(...)) is available on native backends and on WebGPU bridge assetsv0.1.15+that exposestateSaveFile/stateLoadFilebridge APIs. On web, state paths refer to the bridge WASMFS virtual filesystem and are not durable across page reloads. Durable browser storage currently requires app-level export/import outside the DartstateSaveFile/stateLoadFilehelpers. LiteRT-LM currently reports state persistence as unsupported. -
WebGPU readiness is browser/device/runtime dependent. Check secure
context,
navigator.gpu, adapter/features,window.crossOriginIsolated, loaded bridge asset source/version, and model memory pressure before treating a web load failure as a package bug. The WebGPU Bridge page has the browser-console probe and Flutter Web smoke-test path.
Current llama.cpp module availability by bundle (v0.3.0)#
| Bundle key | Available backend modules in bundle |
|---|---|
android-arm64 |
cpu, vulkan, opencl |
android-x64 | cpu, vulkan, opencl |
linux-arm64 | cpu, vulkan, blas |
linux-x64 |
cpu, vulkan, blas, cuda, hip |
windows-arm64 | cpu, vulkan, blas |
windows-x64 |
cpu, vulkan, blas, cuda |
ios-*, macos-* |
Consolidated Apple runtime (
cpu
+
metal
path; no split
ggml-*
module selection in hook)
|
Selector names and aliases#
llamadart_native_backends values are matched against modules discovered in
the selected bundle. Current configurable-bundle module names are:
cpuvulkanopenclcudablaship
Aliases:
vk->vulkanocl->openclopen-cl->opencl
GpuBackend.metal remains valid as a runtime backend preference on Apple
targets, but Apple targets are non-configurable in
llamadart_native_backends.
Configuring native backend modules#
Use hooks.user_defines.llamadart.llamadart_native_tag and
hooks.user_defines.llamadart.llamadart_native_repository to test another
GitHub release source,
hooks.user_defines.llamadart.llamadart_native_path to use a local source, and
hooks.user_defines.llamadart.llamadart_native_backends to select split
llama.cpp backend modules:
hooks:
user_defines:
llamadart:
# Optional. Defaults to llamadart's tested native runtime pin.
llamadart_native_tag: v0.3.0
# Optional. GitHub repository slug or github.com URL.
llamadart_native_repository: leehack/llamadart-native
# Optional. Takes precedence over GitHub downloads when set.
# Relative paths are resolved from the pubspec defining this config.
# llamadart_native_path: ./native-bundles
llamadart_native_backends:
platforms:
android-arm64:
backends: [vulkan]
cpu_profile: full # default; use compact for baseline-only
linux-x64: [vulkan, cuda]
windows-x64:
backends: [vulkan, cuda, blas]
Android arm64 CPU policy keys (platforms.android-arm64):
cpu_profile: full(default): include all Android ARM CPU variant modules.cpu_profile: compact: include baseline CPU variant module only.-
cpu_variants: [...](advanced): explicit CPU variant list, overridescpu_profile.
Supported canonical cpu_variants values:
android_armv8.0_1(baseline)android_armv8.2_1android_armv8.2_2android_armv8.6_1android_armv9.0_1android_armv9.2_1android_armv9.2_2
Variant feature differences:
| Variant | Optional feature set used by that module |
|---|---|
android_armv8.0_1 | baseline |
android_armv8.2_1 | DOTPROD |
android_armv8.2_2 |
DOTPROD + FP16_VECTOR_ARITHMETIC |
android_armv8.6_1 |
DOTPROD + FP16_VECTOR_ARITHMETIC + MATMUL_INT8 |
android_armv9.0_1 |
DOTPROD
+
FP16_VECTOR_ARITHMETIC
+
MATMUL_INT8
+
SVE2
|
android_armv9.2_1 |
DOTPROD
+
FP16_VECTOR_ARITHMETIC
+
MATMUL_INT8
+
SVE
+
SME
|
android_armv9.2_2 |
DOTPROD
+
FP16_VECTOR_ARITHMETIC
+
MATMUL_INT8
+
SVE
+
SVE2
+
SME
|
Accepted cpu_variants input forms are normalized, for example:
baselinearmv8_6_1v9_0_1android-armv9.2_2libggml-cpu-android_armv8.2_2.so
If cpu_variants contains unknown entries, they are ignored with warnings. If
no valid entries remain, selection falls back to cpu_profile (or default
full).
Selection and fallback behavior#
- Configurable targets start from defaults (
cpu,vulkan) if available. -
llamadart_native_runtimescontrols whole native runtime families:llama_cpp,litert_lm, or both. Unset or empty means all runtime families available for the target. -
llamadart_native_backendscontrols only llama.cpp module files insidellama_cpp; it does not affect LiteRT-LM assets. cpuis auto-added as fallback when present in the bundle.- Android arm64 defaults to
cpu_profile: full. -
cpu_variants(if provided) takes precedence overcpu_profilefor Android arm64. - If requested modules are unavailable for a bundle, the hook warns and falls back to defaults.
- If defaults are also unavailable, all available modules in that bundle are used as fallback.
-
Backend-owned runtime dependencies follow the selected backend module. CUDA
runtime DLLs (
cudart64_*,cublas64_*,cublaslt64_*) are bundled only whencudais selected, and OpenBLAS runtime libraries are bundled only whenblasis selected. Unknown runtime libraries are kept for compatibility with future native bundle layouts. -
Apple targets (
ios-*,macos-*) supportcpu+metal, but ignore per-backend module config in this hook path because runtime libraries are consolidated. -
Flutter Apple apps use Swift Package Manager only through runtime companion
packages:
llamadart_llama_cpp_flutterfor GGUF/llama.cpp andllamadart_litert_lm_flutterfor iOS.litertlm/LiteRT-LM. -
For Flutter iOS apps, installed companion packages choose the Apple SPM
runtime families and win over
llamadart_native_runtimes. Flutter macOS LiteRT-LM builds currently use the core native-assets fallback while the hook path remains responsible for the complete runtime. If neither companion package is installed, the core native-assets fallback is used. -
For non-Flutter projects and non-Apple targets,
llamadart_native_runtimesremains the selector even if a Flutter companion package is accidentally present in dependencies. -
llamadart_native_tag,llamadart_native_repository, andllamadart_native_pathcustomize hook-managed native assets. Apple SPM binary target URL/checksum pins are owned by the companion packages underpackages/and can be customized with path/git overrides or forks of those packages. - Standalone Dart macOS runs keep the native-assets path for compatibility.
-
Custom standalone Dart macOS launchers can point
LLAMADART_LITERT_LM_LIB_DIRat the extracted LiteRT-LM cache directory when the default cache search is not suitable. -
windows-x64performs extra runtime dependency validation:cudarequirescudartandcublasDLLs.blasrequires OpenBLAS DLL.
-
If
llamadart_native_tagpoints at a release without a matching bundle asset, the native-assets hook fails while downloading that asset. -
Available override values are
leehack/llamadart-nativerelease tags, notllamadartpackage versions. -
llamadart_native_repositoryaccepts a GitHubowner/reposlug orhttps://github.com/owner/repoURL. -
llamadart_native_pathtakes precedence over GitHub downloads and can point directly at an archive, at an extracted bundle directory, or at a directory containing<tag>/<bundle>/,<bundle>/, or the expected archive file. - Native source overrides do not regenerate Dart FFI bindings or symbol lookups, so they are only safe with compatible native binaries.
-
If you change
llamadart_native_tag,llamadart_native_repository,llamadart_native_path,llamadart_native_runtimes, orllamadart_native_backends, runflutter cleanonce to clear stale native-asset outputs. -
If a native release tag is republished with refreshed assets, also run
flutter cleanbefore rebuilding so an older same-tag extracted bundle does not stay in use.
Vulkan cooperative matrix driver crashes#
Some Vulkan drivers advertise cooperative matrix support but crash inside the
Vulkan property query calls used by upstream ggml-vulkan. This is a driver
failure, not a llamadart loader failure. Use upstream ggml-vulkan's opt-out
environment variables before starting the Dart/Flutter process:
GGML_VK_DISABLE_COOPMAT=1
GGML_VK_DISABLE_COOPMAT2=1
On Windows PowerShell:
$env:GGML_VK_DISABLE_COOPMAT = "1"
$env:GGML_VK_DISABLE_COOPMAT2 = "1"
flutter run -d windows
These variables disable the cooperative matrix optimized Vulkan paths for that process. They can reduce Vulkan performance, so use them only when the Vulkan driver crashes or reports device loss in the cooperative matrix path.