Native runtime configuration
Configure which prebuilt llama.cpp and LiteRT-LM runtimes and backend modules the llamadart build hook bundles, and how to override the native release.
On this page
The llamadart build hook (hook/build.dart) downloads prebuilt llama.cpp and
LiteRT-LM runtimes for the target platform, so apps never compile C++. The hook
reports the downloaded .so, .dylib and .dll files as
package:code_assets
code assets, and the Flutter or Dart build bundles them into the APK, IPA or
desktop app. On macOS, LiteRT-LM libraries stay in the hook cache and load from
there.
Configure the hook under hooks.user_defines.llamadart in the app's
pubspec.yaml. Every key is optional:
| Key | Selects | Default |
|---|---|---|
llamadart_native_runtimes |
Runtime families to bundle:
llama_cpp
,
litert_lm
, and the opt-in
stable_diffusion
|
llama_cpp and litert_lm where published |
llamadart_native_backends |
llama.cpp backend modules, and Android arm64 CPU variants | cpu and vulkan where present |
llamadart_stable_diffusion_backends |
The
stable_diffusion
build on Linux and Windows:
cpu
or
vulkan
|
Follows llamadart_native_backends |
llamadart_native_tag |
leehack/llamadart-native release to download |
The pinned release |
llamadart_native_repository |
GitHub repository to download llama.cpp bundles from | leehack/llamadart-native |
llamadart_native_path |
Local archive or bundle directory used instead of a download | Unset |
hooks:
user_defines:
llamadart:
llamadart_native_runtimes: [llama_cpp]
llamadart_native_backends:
platforms:
android-arm64:
backends: [vulkan]
cpu_profile: compact
linux-x64: [vulkan, cuda]
windows-x64: [vulkan, cuda, blas]
After changing any of these keys, or after a native release tag is republished
with new assets, run flutter clean once so stale native assets are not
reused.
Choose runtime families#
llamadart_native_runtimes keeps whole runtimes out of the app when it ships
only one model format. The value is a list, or a map with a top-level
runtimes list and per-platform overrides under platforms:
hooks:
user_defines:
llamadart:
llamadart_native_runtimes:
runtimes: [llama_cpp, litert_lm]
platforms:
ios: [llama_cpp]
android-arm64: [litert_lm]
-
Platform keys are OS names (
android,ios,linux,macos,windows) or bundle keys:android-arm64,android-x64,ios-arm64,ios-arm64-sim,ios-x86_64-sim,linux-arm64,linux-x64,macos-arm64,macos-x86_64,windows-arm64andwindows-x64. For each bundle the exact bundle key wins, then its OS key, thenruntimes, then every family. -
Aliases:
gguf,llama,llama.cppforllama_cpp;litert,litert-lm,litertlm,.litertlmforlitert_lm;stable-diffusionforstable_diffusion.allandbothselectllama_cppandlitert_lm. Unknown names are dropped with a warning. -
Selecting
litert_lmby name for a target without a LiteRT-LM runtime, such as the iOS x86_64 simulator or Windows arm64, fails the build. When it is only implied by the default orall, the hook drops it with a warning. -
An empty or all-unknown selection falls back to
llama_cppandlitert_lm; one that leaves no runtime, such asnone, fails the build.
Opt-in stable_diffusion runtime (experimental)#
stable_diffusion bundles the
stable-diffusion.cpp runtime
from leehack/stable-diffusion-native for the experimental
ImageGenerationEngine. It adds about 40 to
70 MB per target, so it is never bundled by default or by all/both; name it
explicitly:
hooks:
user_defines:
llamadart:
llamadart_native_runtimes:
runtimes: [llama_cpp, stable_diffusion]
-
Published for
android-arm64,ios-arm64,ios-arm64-sim,ios-x86_64-sim,macos-arm64,macos-x86_64,linux-arm64,linux-x64andwindows-x64, built for iOS 16.4 and macOS 13.3 or newer. At run time the engine checks the CPU before loading the library: Android needs an Armv8.2 CPU with dot-product and fp16 (asimddp,fphp,asimdhp), and Linux and Windows x64 need AVX2, FMA, F16C and BMI2. OtherwiseImageGenerationEngine.loadthrowsLlamaUnsupportedExceptioninstead of crashing on an illegal instruction. On Windows the runtime also needs the latest Microsoft Visual C++ v14 Redistributable (x64); stock Windows Server lacks it, and the error names the DLL that failed to load. -
Apple builds use the Metal build and Android the CPU build. Linux and Windows publish a CPU build (about 38 MB) and a Vulkan build (about 72 MB) that needs the system Vulkan loader (
libvulkan.so.1orvulkan-1.dll) at run time.llamadart_stable_diffusion_backendspicks one, as a list for every platform or aplatformsmap likellamadart_native_backends:hooks: user_defines: llamadart: llamadart_stable_diffusion_backends: [cpu] # or per platform: # llamadart_stable_diffusion_backends: # platforms: # linux: [vulkan] # windows: [cpu]It accepts
cpuandvulkan(aliasvk) and leaves the llama.cpp backends unchanged. A list namingvulkanselects the Vulkan build even when it also namescpu. Without an entry for the platform, or with an empty one, the build followsllamadart_native_backendsfor the same platform: Vulkan when Vulkan is selected there, which is the default, and CPU otherwise. A value naming neithercpunorvulkanis ignored with a build warning and the same fallback applies. -
Other targets, such as
android-x64or Windows arm64, are skipped with a warning. Naming it for that exact bundle key, for exampleandroid-x64: [stable_diffusion], fails the build instead. -
Flutter iOS and macOS apps should add the
llamadart_stable_diffusion_fluttercompanion instead; see Flutter Apple apps. Without it the hook bundles the runtime, and App Store Connect rejects that iOS framework: Flutter writesMinimumOSVersion13.0 into it, while the library needs iOS 16.4. The hook reports this as an Xcode build warning, which Xcode andxcodebuildshow but plainflutter buildandflutter runoutput does not.
Choose llama.cpp backend modules#
llamadart_native_backends filters the split llama.cpp modules inside the
llama_cpp family; it does not affect LiteRT-LM. It is set per platform,
under platforms or as a map keyed by platform. A platform value is a list, a
comma-separated string, or a map with a backends list. A bare top-level list
applies to no platform.
Modules in the pinned bundles:
| Bundle | Modules |
|---|---|
android-arm64, android-x64 |
cpu, vulkan, opencl |
linux-arm64 | cpu, vulkan, blas |
linux-x64 |
cpu, vulkan, blas, cuda, hip |
windows-arm64 | cpu, vulkan, blas |
windows-x64 |
cpu, vulkan, blas, cuda |
ios-*, macos-* |
One consolidated CPU and Metal runtime; not configurable |
Aliases: vk for vulkan; ocl and open-cl for opencl.
GpuBackend.metal still selects Metal at runtime on Apple targets.
Selection rules:
- With no request, the hook bundles
cpuandvulkan, where present. cpuis always added when the bundle has it.- A request that names any module the bundle lacks is rejected whole: the hook warns and uses the defaults. If the defaults are missing too, it bundles every module.
-
CUDA runtime DLLs (
cudart64_*,cublas64_*,cublaslt64_*) ship only withcuda, and OpenBLAS libraries only withblas. -
On
windows-x64, the hook rejects a bundle whosecudamodule lacks the cudart and cuBLAS DLLs, or whoseblasmodule lacks OpenBLAS.
Linux modules need system libraries; see Linux prerequisites.
On Windows every module needs the latest Microsoft Visual C++ v14
Redistributable (x64 or arm64), at least as new as the build tools of the
bundled DLLs. The bundles ship only the OpenMP runtime
(vcomp140.dll on x64, libomp140.aarch64.dll on arm64), not
msvcp140.dll or vcruntime140.dll. vulkan also needs a GPU driver that
provides the Vulkan loader, vulkan-1.dll, and cuda an NVIDIA driver,
which provides nvcuda.dll. When one of those is missing, the module's
startup diagnostic, which a failed model load reports, names it.
When a requested backend is not bundled#
ModelParams.preferredBackend selects only a module the app has. With the
default cpu and vulkan, GpuBackend.cuda has no module to load, so on
Linux and Windows the model loads on CPU with 0 GPU layers rather than on
another GPU backend. llamadart logs a LlamaLogLevel.warn record through the
Dart logger that names the backend and this user-define, and
getBackendName() reports CPU. To use CUDA, add it for the platform:
hooks:
user_defines:
llamadart:
llamadart_native_backends:
platforms:
linux-x64: [cpu, vulkan, cuda]
windows-x64: [cpu, vulkan, cuda]
-
On Windows,
dart runanddart testload only the modules the app bundles, as does a Linux app built withdart build cliand run from another directory. -
On Linux, a process whose working directory is the package root, as with
dart runanddart test, copies each backend module missing from.dart_tool/libout of the hook's download cache,.dart_tool/llamadart/native_bundles/<tag>/linux-<arch>/extracted, when the backend starts. The cache holds every module in the release,cudaandhipincluded on x64, so there CUDA loads without the user-define when its system libraries are installed.
Android arm64 CPU variants#
The android-arm64 map form also takes cpu_profile and cpu_variants:
cpu_profile: full(default) bundles all seven CPU variant modules.cpu_profile: compactbundles only the baselineandroid_armv8.0_1.-
cpu_variants: [...]lists variants explicitly and overridescpu_profile. Unknown entries are dropped with a warning; if none remain,cpu_profileapplies.
| Variant | Optional CPU features |
|---|---|
android_armv8.0_1 | Baseline |
android_armv8.2_1 | DOTPROD |
android_armv8.2_2 |
DOTPROD, FP16_VECTOR_ARITHMETIC |
android_armv8.6_1 |
DOTPROD, FP16_VECTOR_ARITHMETIC, MATMUL_INT8 |
android_armv9.0_1 |
DOTPROD
,
FP16_VECTOR_ARITHMETIC
,
MATMUL_INT8
,
SVE2
|
android_armv9.2_1 |
DOTPROD
,
FP16_VECTOR_ARITHMETIC
,
MATMUL_INT8
,
SVE
,
SME
|
android_armv9.2_2 |
DOTPROD
,
FP16_VECTOR_ARITHMETIC
,
MATMUL_INT8
,
SVE
,
SVE2
,
SME
|
Variant names are normalized, so baseline, armv8_6_1, v9_0_1,
android-armv9.2_2 and libggml-cpu-android_armv8.2_2.so are all accepted.
Override the llama.cpp release#
hooks:
user_defines:
llamadart:
llamadart_native_tag: vX.Y.Z
llamadart_native_repository: leehack/llamadart-native
# llamadart_native_path: ./native-bundles
-
llamadart_native_tagnames allamadart-nativerelease tag, not allamadartpackage version. The hook acceptsvMAJOR.MINOR.PATCH,vMAJOR.MINOR.PATCH-N,bNNNN,bNNNN-NandbNNNN-llamadart.N, neverlatest. List releases withgh release list --repo leehack/llamadart-native --limit 20. -
The release must contain
llamadart-native-<bundle>-<tag>.tar.gzfor the target; otherwise the download fails the build. -
llamadart_native_repositorytakes anowner/reposlug or ahttps://github.com/owner/repoURL. -
llamadart_native_pathwins over downloads. It can point at an archive, an extracted bundle directory, or a directory containing<tag>/<bundle>/,<bundle>/, or the expected archive. Relative paths resolve from thepubspec.yamlthat sets them.
Overrides do not regenerate the Dart FFI bindings, so the binary must stay ABI- and symbol-compatible with the pinned release; the hook logs a warning when an override is active. Two checks fail closed:
-
LoRA adapters need both
llama_adapter_get_alora_n_invocation_tokensandllama_adapter_get_alora_invocation_tokens. Without a compatible pair,setLoraand loads withModelParams.lorasthrowLlamaUnsupportedExceptionrather than activate an adapter whose type it cannot check. -
DSpark speculative decoding (
SpeculativeDecodingConfig.draftDspark) needs at least theb10356-llamadart.1wrapper fix; the pinned release has it.
LiteRT-LM has no override keys; the hook always downloads the pinned
litert-lm-native release and verifies its checksum. Tag grammar for
maintainers: Native and web sync.
Flutter Apple apps#
Flutter iOS and macOS apps link runtimes through Swift Package Manager when a companion package is a dependency:
llamadart_llama_cpp_flutterlinks the llama.cpp XCFrameworks.llamadart_litert_lm_flutterlinks the LiteRT-LM iOS XCFrameworks.-
llamadart_stable_diffusion_flutterlinks the stable-diffusion.cpp XCFramework for image generation.
When the llama.cpp or LiteRT-LM companion is present, the installed companions
choose the Apple llama_cpp and litert_lm families and the rest of
llamadart_native_runtimes is ignored with a warning. stable_diffusion
is
decided on its own: its companion selects it, and otherwise the hook bundles it
when llamadart_native_runtimes names it. Adding only the stable_diffusion
companion leaves llama.cpp and LiteRT-LM on the hook. The build checks the
resolved llama.cpp and stable_diffusion companions' pins against the core
package and rejects local Artifacts overrides. The tag,
repository, path and backend keys do not change SwiftPM binaries; their pins
live in each companion's Package.swift, so use a path or git override or a
fork of the companion. Flutter macOS LiteRT-LM still uses the hook-managed
runtime. Without a companion, and for non-Flutter or non-Apple builds, the hook
path above applies.
Standalone Dart on macOS keeps LiteRT-LM libraries in the hook cache. A custom
launcher can set LLAMADART_LITERT_LM_LIB_DIR to the extracted LiteRT-LM
directory.
How the hook resolves a build#
sequenceDiagram
autonumber
participant Build as flutter build / dart run
participant Hook as hook/build.dart
participant Cache as Local bundle cache
participant Release as GitHub release asset
participant Assets as code_assets
Build->>Hook: invoke native-assets hook
Hook->>Hook: resolve bundle key and runtime families
Hook->>Cache: check cached bundle
alt cache miss or stale
Hook->>Release: download bundle archive
Hook->>Hook: extract and validate libraries
end
Hook->>Hook: select llama.cpp modules
Hook->>Assets: report code assets
Runtime sources: llama.cpp bundles come from
leehack/llamadart-native
and
LiteRT-LM bundles from
leehack/litert-lm-native.
The hook checks each LiteRT-LM archive for its required libraries and
downloads it again when a cached copy is corrupt or incomplete. Which
repository owns which change: Runtime ownership.