Recent releases
Review recent llamadart release highlights and jump to the canonical changelog for full release notes.
On this page
- Unreleased
- 0.9.0
- 0.8.24
- 0.8.23
- 0.8.22
- 0.8.21
- 0.8.20
- 0.8.19
- 0.8.18
- 0.8.17
- 0.8.16
- 0.8.15
- 0.8.14
- 0.8.13
- 0.8.12
- 0.8.11
- 0.8.10
- 0.8.9
- 0.8.8
- 0.8.7
- 0.8.6
- 0.8.5
- 0.8.4
- 0.8.3
- 0.8.2
- 0.8.1
- 0.8.0
- 0.7.2
- 0.7.1
- 0.7.0
- 0.6.17
- 0.6.16
- 0.6.15
- 0.6.14
- 0.6.13
- 0.6.12
- 0.6.11
- 0.6.10
- 0.6.9
- 0.6.8
- 0.6.7
- 0.6.6
- 0.6.5
- 0.6.4
- 0.6.3
- 0.6.2
- 0.6.1
- 0.6.x line highlights
- 0.5.x line highlights
- Release usage guidance
For canonical full release notes, use:
Unreleased#
- Add an observability guide and tested optional OpenTelemetry example with Langfuse and Grafana recipes.
0.9.0#
- Document generic JSON tool calling as an intentional fallback, including its prompt and model-reliability limits; runtime behavior is unchanged (#755).
-
Load Qwen3.5-0.8B on the Web CPU (WebAssembly) backend at the default
contextSizeusing smaller processing batches, while preserving full-context defaults for unknown models, including embedding models (#752). -
Report a failed Web model load on a page without cross-origin isolation as
LlamaModelExceptionwith its real cause, not as a COOP/COEP worker-thread error; only a real worker-thread failure still names COOP/COEP (#753). -
Log a Dart warning when an explicit
preferredBackendGPU module is not bundled and the model loads on CPU instead, as happens forcudawith the default Windows bundle; the native runtime docs now say when CUDA is bundled (#756). -
Fix the
llamadart_serverexample exiting at startup on Windows; it stops on Ctrl+C there, and on SIGINT or SIGTERM elsewhere (#757). - Accept MP3 and FLAC bytes, as well as WAV, for Qwen3-ASR speech to text on Web (#723).
-
Apply
presencePenalty,minPandthinkingBudget, and runtime LoRA adapters (setLora,removeLora,clearLoras), on WebGPU with bridge assets whose capability probes report them; other assets still reject them (#722). -
Run speculative decoding on WebGPU with bridge assets whose capability
probe reports the strategy, and report each runtime's strategies in
backendGenerationCapabilities.speculativeDecodingStrategies; other assets still reject it (#722). -
Reject a non-zero
GenerationParams.minPon WebGPU when the bridge lacks Min-P, withLlamaUnsupportedExceptioninstead of ignoring it, and ignore a stop sequence equal to apreservedTokensentry there, as native llama.cpp does (#661). -
Add
LlamaEngine.backendGenerationCapabilities, which reports whether the loaded runtime appliespresencePenalty,minPandthinkingBudget; the example chat app uses it to send Min-P and enable its slider only where supported (#661). - Return a DeepSeek V3 forced-open thought that never closes as reasoning, as llama.cpp does with the DeepSeek V3.1 template, instead of as content with its tool calls (#743).
-
Keep escaped
\nand\rin Qwen3-Coder XML reasoning, as llama.cpp does with the Qwen3.5 template; before, they became line breaks unless a tool call ended the thought (#743). -
Stream content and reasoning with the whitespace the non-streamed parse
keeps, with or without tools, so streamed answers and
ChatSessionhistory no longer start with the blank lines after</think>. Only whitespace at either end, and text that may be a tag or tool-call opening, waits for more output, so reasoning still streams token by token. Without tools, Hermes, DeepSeek R1, Qwen3-Coder XML and the other formats the template engine guide lists also drop a start tag repeated at the start of a forced-open thought, as the parse does. The guide lists the exceptions (#754). -
Stream content that equals the non-streamed parse for Qwen3-Coder XML,
Mistral Nemo and 15 more tool-call formats, and for output parsed with a PEG
parser, so text before a tool call no longer carries the tool-call envelope
into streamed content or
ChatSessionhistory. Content and reasoning are trimmed as for Hermes, and text after a call arrives at the end of the stream. See the template engine guide for the formats and exceptions (#732). -
Stream Hermes-format content that equals the non-streamed parse, so text
before a tool call no longer carries the
<tool_call>envelope into streamed content orChatSessionhistory; only a possible envelope opening and trailing whitespace wait for more output. With tools, streamed content is now trimmed as the parse trims it, and text the parse keeps after a tool call, including a malformed envelope, arrives at the end of the stream. Streamed reasoning, and soChatSessionthinking, is trimmed per thought as the parse trims it. The exception is a forced-open thought that never closes and starts with whitespace: the parse keeps it untrimmed, but it streams without its leading and trailing whitespace, so" \n Hello there. \n\n"streams as"Hello there.". Before, it streamed as the parse gives it, except for some thoughts containing a backslash, depending on chunking. After a forced-open thought, text after a tool call arrives at the end (#701). -
Throw
LlamaModelExceptionwhen a WebGPU model load fails with a bridge error that has no specific mapping, andLlamaInferenceExceptionorLlamaStateExceptionfor such Web embedding, next-token scoring and state errors, with URL credentials and signed query values redacted from the details and the load-failure console log (#704). -
Keep URL credentials, signed query values and fragments out of
LlamaEnginemodel and projector load errors and logs and themodelfield of completion chunks for every URL form, including scheme-relative//user:pass@host/...URLs and relative paths with a query, and out of native model download errors. A projector load error that is not aLlamaExceptionnow throwsLlamaModelException. Thedetailsof a model or projector load failure is now a{type, message}map instead of the original error, and a native download that fails with a network error carries the error text as aStringindetails(#704). - Keep the text after a U+0000 in native llama.cpp tokenization, embeddings and generation prompts instead of dropping it (#608).
-
Make
DecisionEngine.loadthrowLlamaStateExceptionwhen another model is loaded while it runs, even under the same backend handle (#626). -
Bound speech validation pack memory by a footprint counter instead of the
resident set:
phys_footprinton macOS and iOS,RssAnonplusRssShmemplusVmSwapon Linux and Android, read aftermalloc_trim(0)where the C library provides it (glibc, not Android), andPrivateUsageplusSharedCommitUsageon Windows (PrivateUsagealone on builds without it). Evicting file-backed pages, such as the mmapped weights, compressing memory under pressure, or glibc keeping freed memory across reloads no longer failspeak_memory_boundwithout memory growth, and each report names its counter (#633, #762). -
Fail speech validation
leak_slope_boundwhen the least-squares footprint slope over cleanup cycles 1-8 exceeds 7 MiB per cycle; it failed only when every cycle grew by more than 7 MiB, and passed leaks of 16 MiB per reload (#762). - Detect chat template capabilities with llama.cpp's probes, and give templates that read only typed content text parts, as llama.cpp does: SmolVLM prompts keep the message text, Ministral 3 renders an image followed by a reasoning-only turn as llama-server does, TranslateGemma 2B keeps the text next to an image, and Kimi-K2 tool results after an image are plain text (#720).
-
Use the LFM2 format for LFM2.5 templates that list tools without
<|tool_list_start|>, as llama.cpp does, so LFM2.5-1.2B-Instruct and LFM2.5-1.2B-Thinking tool prompts drop the stray "Respond in JSON format" instruction (#716). -
Report per-request usage on WebGPU with the newly pinned bridge assets
v0.1.54, on the finalcreatechunk and to observers (#696). -
Give assistant turns that hold only tool calls or only reasoning empty
content instead of
nullin chat templates, as llama.cpp does: QwQ-32B renders them instead of throwing, and LFM2 and Devstral prompts drop a straynullor<function text>(#715). - Pass Map and List tool results as compact JSON text to LFM2, gpt-oss, Solar Open, Ministral, DeepSeek V3 and TranslateGemma templates too, instead of Python-style or spaced text (#717).
- Pass earlier tool-call arguments as JSON objects to templates that read them as objects, as llama.cpp does, so Qwen3, Ministral, Devstral, gpt-oss and similar prompts format them with the template's own JSON spacing (#702).
-
Require
dinja1.2.0, so more chat prompts match llama.cpp:tojsonoutput such as tool declarations uses llama.cpp's spacing, number format and non-ASCII text; Qwen3-Coder, GLM-4.6, GLM-4.7-Flash, MiniMax-M2, Nemotron-3-Nano, Command R7B, Cohere2 MoE and Ling 3.0 prompts lose stray indentation; Functionary v3.1 adds no tool instructions without tools; Granite 3.3 spells out the month in its date; Hunyuan Hy3 keeps the system prompt first instead of merging it into the user turn; and Bielik 11B v3 tool-call turns without text render instead of throwing. -
Add
LlamaEngine.scoreNextToken(...)for next-token log-probabilities on native llama.cpp and WebGPU bridge assetsv0.1.52+, matching llama-servern_probs; checksupportsNextTokenScoringfirst (#694). -
Add
example/laya_command_bar, a Flutter text field that reshapes into a reminder, message, calculation or other command as you type, read by a Laya decision model, by EmbeddingGemma and labelled examples, or by small LLMs' next-token scores, including the decision model decider-2b. -
Count generated tokens with an empty text piece in llama.cpp
getPerformanceContext()evalTokensandsampleCountwithout speculative decoding, as the speculative path already did (#706). -
Report per-request token usage and timings on the final
createchunk asLlamaCompletionChunk.usageon native llama.cpp (#696). -
Add
LlamaEngine(observers: ...), which reports chat and text completions, embeddings and model loads, with their usage, to tracing and metrics code (#696). -
Throw
LlamaModelExceptionwhen native llama.cpp cannot find or load a multimodal projector, andLlamaUnsupportedExceptionwhen the runtime lacks the mtmd functions;LlamaEngine.supportsAudioalso throws the latter. Speech-to-text capabilities now say when no projector is loaded, using the newLlamaEngine.hasMultimodalProjector(#325). -
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@v0.5.0(llama.cppv0.5.0) with Apple companion0.0.20, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and aligned current README/website native override docs. -
Honour
LlamaEngine.cancelGeneration()issued right after listening to acreate,generateorChatSession.createstream, before it reaches the backend, instead of running the whole generation (#602). -
Cancel an active text-to-speech synthesis on
LlamaEngine.unloadModel()anddispose()instead of waiting for it to finish (#628). -
Cancel an active Qwen3-ASR transcription on
LlamaEngine.unloadModel()anddispose()instead of completing it with the transcript cut at the unload (#670). - Send LiteRT-LM tool calls and tool results in the runtime's own message format, so Gemma 4 reads tool output and Qwen3 tool histories no longer fail (#681).
-
Start a native llama.cpp generation requested right after a cancel once the
cancelled run stops, instead of failing with
generation is already in progress. An overlap with a running generation that was not cancelled now throwsLlamaStateException(#655). -
Render Qwen3 prompts as llama.cpp does: an earlier assistant tool-call
turn without reasoning no longer gets an empty
<think>block (#691). -
Require
dinja1.1.0. Its Jinja string comparison makes three more chat templates render as llama.cpp does: MiniMax-M1 adds no empty system block for an empty or whitespace-only system message; NVIDIA Nemotron Nano v2 drops the blank line before a tool call, the blank lines before its tool instructions when tools come with an empty or whitespace-only system message, and an empty final assistant turn without a generation prompt; and Functionary v3.2 tool declarations drop stray// Format=<|NONE|>lines and spell out nested object parameters (#351). -
Cancel a generation's backend run as soon as its stream subscription is
cancelled, instead of at its next token, which during prompt evaluation
meant after the whole prompt. A native llama.cpp generation requested right
after such a cancel now waits for it instead of throwing
LlamaStateException, and native llama.cpp sees a cancel between text prompt micro-batches (ModelParams.microBatchSize, 512 tokens by default) or, with speculative decoding, between batches (ModelParams.batchSize) (#663, #660). -
Render every result of a tool message holding several
LlamaToolResultContentparts, as onetoolmessage per result like llama.cpp, instead of only the first;LlamaChatMessage.toJsonlists them all (#683). - Render Gemma 4 tool calls and tool results as llama.cpp does, so Gemma 4 GGUF models can read tool output (#669).
- Stop a Qwen3-TTS audio decode at its next chunk boundary when native text-to-speech is cancelled, instead of finishing the native step in progress first. This needs llamadart-native v0.4.1-1 or later; older runtimes keep the previous behaviour (llamadart-native#86, #322).
-
Add an experimental
DecisionEnginefor Laya-style decision models (a ModernBERT encoder GGUF plus a safetensors head) on native llama.cpp, with typedChoiceKey,ScoreKeyandNoulKeyquestions (#604). -
Add
example/basic_app/bin/llamadart_decision_example.dart, a console demo that triages a support ticket withDecisionEngine(#604). -
Add
example/laya_tetris, a Flutter app in which a Laya decision model plays real-time Tetris throughDecisionEngine(#604). -
Run
example/laya_tetrison Web through the WebGPU bridge, with a live demo at https://leehack-flutter-laya-tetris.static.hf.space. -
Add a notebook in
example/laya_tetris/training/that fine-tunes a Laya decision head for the Tetris example and exports it forDecisionEngine(#604). -
Run
DecisionEngineon WebGPU through the bridge decision API (apiVersion 1), which bridge assetsv0.1.47+include (#604). - Stop native image and audio requests from seeding the repeat penalty with leftover memory, which made output depend on the previous request (#603).
-
After a failed native prompt decode, the next
reusePromptPrefixrequest no longer runs on the wrong KV cache or keeps failing (#601). -
Reject
embed()andembedBatch()on rank-pooled reranker GGUFs withLlamaUnsupportedExceptionon native, instead of returning memory read past llama.cpp's classifier-score buffer (#583). -
Throw
LlamaInferenceExceptionfrom nativeembed()andembedBatch()when input to an encoder-only model or a model without a KV cache (such as BERT-family and ModernBERT GGUFs) does not fit onemicroBatchSizepass, instead of aborting the process or embedding only the last chunk (#607). -
Aligned default WebGPU bridge assets to
v0.1.54for the decision API, next-token scoring, presence penalty, Min-P, thinking budgets, runtime LoRA adapters, speculative decoding and the Web runtime fixes below; the bridge also adds itssupportsCompletionUsageflag, which llamadart does not use yet (#729). The assets embed llama.cppv0.5.0, are qualified against nativev0.5.0, and keep Web/native llama.cppv0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03bparity and Web@litert-lm/core@0.15.0. Immutable Web asset manifest:8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176. -
On Web, an invalid GBNF grammar now fails generation with a
LlamaInferenceExceptionwhose details contain(invalid grammar), and the loaded model stays usable, instead of aborting the WebGPU bridge runtime (llama-web-bridge#125). -
On Web, when a bridge worker fails and the main-thread reload fails too, or
a replacement worker cannot start, the bridge now forgets the model instead
of being left broken with a
TypeError. Later calls fail withNo model loaded. Call loadModelFromUrl first.; callunloadModel(), thenloadModel()and any projector again to recover (llama-web-bridge#123, llama-web-bridge#127). - On Web, a replacement bridge worker reloads the current model before its next request (llama-web-bridge#126).
-
On Web, grammar-constrained generation no longer aborts the bridge runtime
when top-k or top-p keeps only tokens the grammar rejects; it resamples as
llama.cpp does, and fails with
Grammar rejected every candidate tokenonly when the grammar cannot continue (llama-web-bridge#118). - On Web, an ordinary error from a healthy bridge worker, such as a prompt that overflows the context or empty embedding input, is now rethrown with the worker kept, instead of moving the session to the main thread for good and re-running the request (llama-web-bridge#119, llama-web-bridge#120).
-
Extend the GGUF speech-to-text validation pack with four synthetic edge
fixtures built in-process, so no extra audio is stored: generated digital
silence, plus a truncated RIFF, a stereo 44.1 kHz re-encode and a 33-second
concatenation, the last three derived from the locked
jfk.wav. Every GGUF STT pack run executes the four checks and each one gatesfunctional_pass; the LiteRT-ASR and TTS packs pass no edge fixtures (#325). -
Raise every
stt,ttsandlitert-asrspeech validation pack run from 8 to 15 lifecycle checks: an immediate cancel, three cancel/dispose/load/generate cycles, and bounds that fail the run when a cancellation takes over 500 ms to end its task or the peak resident set exceeds 1.10x the one sampled after the first generation (#594). -
Add cross-platform validation cases for a cancel issued right after
listening, a generation requested right after a cancel, an overlapping
generation, an invalid GBNF grammar and
ToolChoice.autoon a prompt that needs no tool, and run the tool cases on the GGUF chat profiles (#602, #655, #654). -
Add
decision-gguf-{cpu,metal,vulkan,cuda,webgpu}validation profiles that checkDecisionEnginetoken ids, raw logits and answers against the Laya 0.3.5 reference, plus batching, reload and typed rejections, on desktop, mobile, Web WebGPU and GCE CUDA (#604). -
Add a Web-only
chat-gguf-webgpuvalidation profile, hash Web validation models while they stream so GGUFs over 2 GiB pass preparation, list every bundled profile in the validation app, and verify iOS GGUF GPU placement from the XCTest console log. - Run eight cleanup cycles instead of three in every speech validation pack, and fail a run whose resident set grows by more than 7 MiB in each of seven warm cycles. The 1.10x peak ratio no longer applies on Linux CUDA, where reload overhead that levels off failed it without a leak (#686).
-
Add GGUF speech validation pack checks:
ttsunloads and disposes the engine during a synthesis, cancels one during its audio decode and bounds the resident set those checks add, andsttmust fail withLlamaSpeechTranscriptTruncatedExceptionatmaxOutputTokensand at the context size.sttruns now execute 28 checks andttsruns 25 (#628, #636, #322). -
Force greedy
topK: 1for zero-temperature LiteRT-LM Web generation, matching the native clamp (#548). -
Log
Model … loaded from …; native engine creation is deferred until the first generation or tokenizer callinstead ofloaded successfullywhen the native LiteRT-LM backend finishesloadModel, since it creates the engine lazily (#569). -
Require
dxcompiler.dllanddxil.dllin the Windows x64 LiteRT-LM runtime cache and desktop validation bundle checks, matching the hook's v0.17.0-6 inventory. The runtime does not preload them: Dawn's D3D12 backend loads the pair at GPU engine creation (#570). -
Document that Linux llama.cpp loads need the OpenMP runtime
(
libgomp.so.1;libgomp1on Ubuntu/Debian,libgompon Fedora and Arch) and that Linux LiteRT-LM GPU needs a hardware Vulkan ICD: with only Mesa llvmpipe the runtime segfaults after model load instead of failing cleanly (llamadart-native#82, #572). -
Record one startup diagnostic when every candidate of a native backend
module family fails to load, or when a
ggml/wrapper symbol is missing from both the primary FFI asset and every fallback library. Candidates are named by asset URI or file name only and loader errors are classified, never quoted, so no directory or loader search path reaches the diagnostic (#416). -
Forward llama.cpp and LiteRT-LM worker-isolate log records to the
LlamaEngine.configureLogginghandler. A worker takes the Dart logger level when it starts andLlamaEngine.setDartLogLevel/setLogLevelupdate a running worker; the defaultnonesends nothing anddebugrecords are capped at 1000 per worker. The LiteRT-LM program-cache pruning warnings are now ordinarywarnrecords gated by that level instead of the native log level. AddsLlamaLogger.leveland theBackendDartLogLevelcapability (#567). -
Replace the token in
Bearer <token>and the value intoken=,key=,secret=,password=,api_key=andapikey=<value>outside HTTP URLs in native startup diagnostics with<redacted-secret>; URL and control-character handling is unchanged (#551). -
Add
ModelParams.liteRtLmCacheDirto choose the native LiteRT-LM runtime cache directory and opt-inModelParams.liteRtLmMaxProgramCacheBytes, which deletes*_mldrift_program_cache.binfiles above the cap before each engine create and logs a warning per deleted file. Defaults are unchanged: the same per-platform directory and no pruning (#552). -
Force greedy
topK: 1for zero-temperatureLiteRtLmRuntimeClient.createConversationcalls, which returned incoherent text on the LiteRT WebGPU sampler with the default top-k. -
Keep root-cause native startup diagnostics when the buffer or the rendered
startupDiagnostics=[...]suffix overflows: teardown entries, now prefixedteardown:, are dropped first, duplicates are recorded once, entries are capped at 2048 characters, and each omitted run renders as...(#415). -
Skip the Windows altered-search-path preload for wrapper library candidates
whose absolute path does not exist, so lazy wrapper API lookups no longer
record a
Failed to preload Windows backend modulestartup diagnostic per missing candidate (#550). -
Accept
promptTemplateon the non-nativeLiteRtLmRuntimeClient.createConversationplaceholder, so callers passing it compile for Web as they do on native (#549). -
Cache
TemplateCaps.detectresults in a per-isolate LRU keyed by exact template source and bounded at 16 entries, so repeated chat-template renders skip both Jinja parses and all four capability probes. Detections in which any analysis step failed are not cached and keep logging on every call (#448). -
Detect
supportsToolsandsupportsToolCallsfor chat templates that reject two tool calls in one assistant message (Llama 3.2) or a user turn directly after a tool call (Ministral 3). The tools capability probe now renders a single tool call, and a separate parallel probe clears onlysupportsParallelToolCallswhen it throws (#557). -
Limit the Ministral tool-call grammar to a single
[TOOL_CALLS]block unless parallel tool calls are enabled; it previously always allowed repeats while the parser kept only the first call (#559). - Limit the Nemotron v3 tool-call grammar (Qwen3-Coder XML format) to a single tool call unless parallel tool calls are enabled; it previously always allowed repeats while the parser kept only one call (#562).
- Pin the WebGPU model-load retry ladder with browser tests for the ladder advance, the wasm64 BigInt restart on wasm32, the restart without the remote fetch backend, and forced remote-fetch chunk halving stopping on both its ten-restart cap and its 4 KiB minimum chunk, then collapse the duplicated attempt thread-count switch into one helper (#361).
-
Apply the JSON Schema
patternkeyword when generating GBNF, for anchored patterns built from literals, positive character classes,(...)groups nested at most 32 deep, grouped alternation and*/+/?/{m,n}repetition; any other pattern, including deeper nesting, falls back to the rule the schema would have produced without it, sominLengthandmaxLengthstill apply there. An applied pattern replacesminLength/maxLengthas it does in llama.cpp, so a schema carrying bothpatternandmaxLengthis no longer length-bounded. A schema carryingpatternbut no explicittypenow yields a string rule instead of throwingUnrecognized schema. Mistral Nemo and Magistral tool-call ids are now grammar-constrained to exactly nine alphanumerics (#582). - Stop a cancelled llama.cpp image or audio prompt at the next prompt chunk, or between a media chunk's encode and its embedding decode, instead of after the whole prompt is ingested. The native call already running still finishes (#599).
-
SpeechToTextEnginenow fails a native Qwen3-ASR transcript that reaches the context size ormaxOutputTokenswithLlamaSpeechTranscriptTruncatedExceptioninstead of completing with truncated text, andcreate()reportsfinishReason: 'length'when native llama.cpp stops at either limit. The chat app caps Qwen3-ASR recordings at the validated 30 seconds (#636). -
Stop
ToolChoice.autoon WebGPU from forcing a tool call: it now skips the lazy tool-call grammar and parses tool calls best-effort. WebGPU rejectsGenerationParams.grammarLazyand a non-rootgrammarRootwithLlamaUnsupportedException; backends report this through the newBackendLazyGrammarSupport(#654). -
Name the CUDA 12 runtime libraries (
libcudart.so.12,libcublas.so.12) that the Linuxcudabackend needs on the default loader path; llamadart does not ship them, and validation bundles refuseLD_LIBRARY_PATH(#587). -
Remote validation runs resolve
packages/llamadart_validationbefore building the report, instead of reportingFAILEDwith a null error on a fresh checkout. A failed report step is now the run's error, with its exit code and a redacted stderr tail (#688). -
Select the devices of an explicit
GpuBackend.metalorGpuBackend.hipon llama.cpp: they looked up ggml registries namedMetalandHIP, but ggml names themMTLandROCm, so loading fell back to automatic device selection. A HIP load on a ROCm build now reports its backend asHIPinstead ofCPU(#611). -
Report a WebGPU model load that fails with
error 138as the documented cross-origin isolation (COOP/COEP)UnsupportedError, asthread constructor failedalready was, instead of rethrowing the raw bridge error (#598). -
Throw
LlamaModelExceptionwhen WebGPU cannot fetch or load a multimodal projector, instead of the raw JavaScript error. Its details drop these parts of the projector URL the app passed, as written, JSON-escaped, percent-encoded or percent-decoded: the userinfo and password, as whole tokens of any length; the?queryand#fragment, where they directly follow a non-space character; the query and each&-separated part that contain=, as whole tokens; and bare query values and the fragment of 10 or more characters, as whole tokens. A whole token has no ASCII letter or digit directly before or after it. Shorter bare values printed on their own, such as the1of?v=1, stay. Other URLs in the details lose userinfo, query and fragment on a best-effort basis (#642). -
Leave no envelope text in the parsed
contentwhen Qwen2.5 wraps a Hermes tool call in double braces (<tool_call>{{"name": ...}}</tool_call>, with any number of extra closing braces) without a grammar. Calls are extracted as before; a double-brace call with other malformed envelope text keeps that text. This deliberately differs from upstream llama.cpp, which rejects that output and extracts no call (#662).
0.8.24#
-
Align native
leehack/llamadart-native@v0.4.1on upstreamb29c606e28a01b1bc8c1351026a0fa6e616bf6c4, with matching Dart bindings and Apple companion0.0.19. This resolves the 0.8.23 grammar limitation: native{2000}repetitions are accepted again (llamadart-native#76). -
Aligned default WebGPU bridge assets to
v0.1.44for matching Web/nativev0.4.1@b29c606e28a01b1bc8c1351026a0fa6e616bf6c4parity. Immutable Web asset manifest:8d61f453753ac7a7d839ac12318b70986a814748d86029993118c19454293aa9. -
Update native LiteRT-LM to
v0.17.0-6with Apple companion0.0.11, Qwen3 tokenizer compatibility, corrected Linux loading, and explicit Linux and Windows GPU selection while retaining CPU defaults. The Windows x64 runtime bundlesdxil.dllanddxcompiler.dll, which D3D12 GPU engine creation requires (litert-lm-native#47). Web LiteRT-LM stays at@litert-lm/core@0.15.0. - Preserve required iOS LiteRT-LM provider and Metal plugins, handle dependency ordering in companion libraries, and exclude metadata/import archives from runtime library inventories.
- Respect greedy sampling for zero-temperature native LiteRT-LM generation.
- Restore native Qwen3 chat text when thinking is disabled and preserve plain system instructions when seeding LiteRT-LM conversation history.
- Settle pending LiteRT-LM requests when a worker stops, close response ports, and report unverified native cleanup as an error.
- Preserve Unicode when detokenizing native GGUF tokens and suppress caller stop markers across chunk boundaries and speculative decoding.
- Restore native Qwen3-ASR file/encoded-byte transcription parity by keeping encoded audio out of string chat-template prompts.
- Render typed tool results as JSON text while preserving string results, validate Qwen XML argument types against schemas, and reject malformed or undeclared tool calls without exposing executable tool deltas.
- Prevent split MiniMax M3 thinking delimiters from leaking into reasoning.
- Discover Windows backend libraries in compiled CLI bundles and provide bounded diagnostics for unavailable native thinking-budget helpers.
- Fix fresh macOS Flutter dependency scanning while retaining Apple ABI and local-override guards.
- Add locked Gemma 4/Qwen3.5 validation profiles, explicit NPU coverage, and opt-in speech and voice-pipeline diagnostics.
- Preserve physical iOS speech-test failure-phase diagnostics and keep repository writer checks independent of generated website output.
- Retain open LiteRT-LM qualification gaps: macOS Qwen3.5 GPU reload latency (#521), Gemma exact-history behavior (#513), and Qwen3-0.6B arithmetic on Android CPU and iOS CPU/GPU (#509). The affected cases remain unqualified; these changes do not resolve the failures or establish their remaining owning layer.
0.8.23#
-
Fail Apple builds before native symbol lookup when the resolved llama.cpp companion does not match the core native runtime, with actionable upgrade guidance.
-
Adopted native llama.cpp v0.4.0 with matching bindings and multimodal calls. Saved native sessions from older runtimes must be regenerated. Apple companion
0.0.18supplies the matching native runtime. -
Aligned default WebGPU bridge assets to
v0.1.43for Web/native llama.cppv0.4.0@5266f24da75dc449bd56cbed7addb9c8e4a6a73eparity. Web@litert-lm/core@0.15.0and native LiteRT pins are unchanged. Immutable manifest:111eefc3588842cebfe665b363378edca34924764610263e1eda5280dfcfaa27. -
Known upstream limitation: llama.cpp v0.4.0 can reject large grammar repetitions, such as
root ::= "a"{2000}. The post-v0.4.0 correction is tracked in native #76 and is not included in this release.
0.8.22#
-
Updated
llamadart_llama_cpp_flutterto0.0.17with the Apple SwiftPMv0.3.0runtime pin. -
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@v0.3.0, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and aligned current README/website native override docs. -
Aligned default WebGPU bridge assets to
v0.1.41for corrected TypeScript declarations and TTS recovery guidance, retaining Web/native llama.cppv0.3.0@c1d0e7a004015f23bc0233470b747b596f29b264parity and Web@litert-lm/core@0.15.0. Immutable manifest:fe97604daabaad6aefa223a8637d5fd9dcac09dd4a61b2ef19cd6aabb39392b9. -
Consolidated native release tag grammar across Dart, Python, Bash, workflows, and documentation via a machine-readable fixture contract (
#404).
0.8.21#
-
Aligned default WebGPU bridge assets to
v0.1.39(immutable manifestb355d01040604f6ae2c5c5fe5bb42b858101a96f03f67e4b27b32fe41ce3b2bf), restoring Web and native llama.cpp upstreamv0.2.0@bb4caa7540188872173c44d161602d9271386413parity with native anchorllamadart-native@v0.2.0-1while preserving Web@litert-lm/core@0.15.0packaging. -
Fixed native Qwen3-ASR transcription by applying the model chat template to audio turns, while preserving the validated raw-prompt Web bridge contract. Empty ASR output now fails explicitly instead of reporting an empty result.
-
Fixed Qwen 2.5/3 LiteRT-LM
ToolChoice.requiredrequests silently running without their required-call grammar and finishing with no call. They now fail with an actionableLlamaUnsupportedExceptionbefore generation when the backend cannot enforce the declared tool schema; Gemma 4 compatibility andauto/nonetool routing are unchanged. -
Fixed Gemma 4 thinking-budget output so split channel controls and tool-call envelopes stay out of visible assistant content while preserving ordinary whitespace. The chat example now also validates custom tool declarations, executes declared host handlers exactly once, appends tool results, and performs a bounded continuation for both llama.cpp and LiteRT-LM backends.
-
Patched the website's vulnerable
nanoidanduuiddependency paths. Until Docusaurus replaces its unpatched image parser, automatic local Markdown images are rejected; website contributors should use static pathname URLs. -
Fixed
llamadart_native_runtimesvaluesnone,off, and the stringfalseselecting every runtime family instead of none; the build hook fails with itsNo native runtimes selectederror again, as it did before 0.8.0. A YAML booleanfalseclears the selection too. Unset, empty, and all-unrecognised config still select every family. -
The published package no longer ships the
doc/directory; that contributor and maintainer documentation is maintained on GitHub, and the packaged files that link to it now use absolute URLs. -
Fixed unanchored
docs/andwebsite/publish-exclusions that matched those directory names at any depth and droppedtool/docs/plus thellamadart_serverexample's OpenAPI spec and Swagger UI sources from the package, leaving the published example unable to analyze. Both patterns are now root-anchored. -
Narrowed the
dinjadependency constraint to>=1.0.0 <1.1.0so chat template capability detection cannot silently resolve against an unverified Jinja parser minor. A 1.0.x patch can still reorganise the private sources the analyzer imports; a new coupling test turns that into a named failure. -
Fixed Command R7B, Hermes, and Hunyuan V3 tool grammars so distinct tool or parameter names cannot collide after conversion to internal GBNF rule names.
-
Fixed DeepSeek V3.2 DSML tool calls using their upstream
<|DSML|function_calls>envelope while preserving DeepSeek V4's distinct<|DSML|tool_calls>grammar and parser behavior. -
Fixed partial GLM 4.5, Poolside Laguna, and Muse Glimmer tool envelopes leaking into streamed assistant content, while preserving completed calls, malformed final output, and ordinary text surrounding Muse recipient channels.
-
Fixed schema-constrained tool calls for Kimi K3, MiniMax M1/M3, DeepSeek V3.2/V4, and Muse Glimmer, including exact escaped names, required fields, declared value types, matching MiniMax M3 element tags, zero-argument calls, and strings containing delimiter characters. Required-tool mode now accepts each format's reasoning/content prefix while still requiring a call. MiniMax M3, DeepSeek DSML, Muse Glimmer, Poolside Laguna, and GLM 4.5 now reconstruct argument values from the declared tool schema instead of guessing from text. Added
ToolParam.nullTypefor null-only JSON Schema properties. -
Made native video-input capability truthful without claiming end-to-end support. Explicit video content now receives a typed actionable rejection, public capability remains false, and the native probe calls
mtmd_helper_support_videobecause helper symbols are exported even when video is compiled out. Full support still requires companion FFmpeg/ffprobe packaging and Dart frame-lifecycle wiring. -
Native release synchronization and build-hook overrides now accept stable
vMAJOR.MINOR.PATCHartifacts and orderedvMAJOR.MINOR.PATCH-Nwrapper rebuilds plus nightlybNNNN-Nrebuilds, while preserving historicalbNNNNandbNNNN-llamadart.Nartifacts. Leading-zero nightly tags, rollback, wrapper/nightlylatestresults, missing bundles, and manifest/checksum/version skew fail closed; the default native pin is unchanged. -
Fixed Web/native backend API parity.
WebAutoBackendnow forwards grammar constraint support from its active runtime, so strict structured output fails early with an actionable error on unsupported Web backends, and the Web-safeLiteRtLmRuntimeClientstub now exposes the native client's thinking-tag configuration method. -
A failed llama.cpp model load now reports the startup diagnostics collected during native library discovery, so a missing or unloadable runtime library explains itself instead of surfacing as a bare load failure. Platforms that record no diagnostics keep their previous message unchanged.
-
llama.cpp worker errors now keep their type. Every backend method routes an
ErrorResponsethrough the file's own error mapper instead of rebuilding a bareException,tokenizeanddetokenizeno longer discard the worker's error entirely, and a coreUnsupportedErrorraised for an unavailable native capability is classified rather than flattened. State-file failures now throwLlamaStateException. -
Fixed multimodal media placeholders being normalized inconsistently.
<img>,<|img|>,<start_of_image>and indexed markers such as<|image_1|>are now rewritten to the mtmd marker on every path, rather than depending on which layer rendered the prompt. MiniMax-M2 and MiniCPM-5 also now detect a forced-open thinking block using the same rule as every other handler. -
aLoRA adapters are now rejected with
LlamaUnsupportedExceptioninstead of being applied like ordinary LoRA adapters. An aLoRA adapter must activate only after its invocation tokens appear in the prompt, so applying it from the start of generation silently changed output. Missing metadata-inspection symbols in custom native runtimes also fail closed with the same typed error, and rejected adapters are released when the cleanup ABI is available. LoRA errors from the worker keep their typed exception instead of arriving as a bareException, and a failed adapter load now throwsLlamaModelException. -
Deprecated
LiteRtLmRuntimeClient.conversationTokenCount()andreplaceConversationWithClone(). Both are unused and are scheduled for removal in the next major release; open an issue if you depend on either. -
NativeLlamaBackend.modelLoadFromUrlnow throwsLlamaUnsupportedExceptioninstead ofUnimplementedError, bringing it into theLlamaExceptionhierarchy. It is the same exception typeLlamaEngine.loadModelFromUrlalready throws for this condition; each keeps its own message. -
Updated the default native llama.cpp runtime to the immutable
leehack/llamadart-native@v0.2.0-1release (llama.cppv0.2.0), adding LFM2 DSpark support plus current upstream correctness and backend performance fixes. Matching Dart FFI bindings, including the new multimodal projector-device field, and the Apple SwiftPM artifact checksum were refreshed. Linuxlibmtmd.so.0now loads without the oldlibmtmd.so.SOVERSIONcompatibility alias. -
Removed the abandoned Dart-side MTP/n-gram speculative-decoding scaffolding from the llama.cpp backend; speculative decoding behavior is unchanged.
-
Bumped
llamadart_llama_cpp_flutterto0.0.16so thev0.2.0-1Apple SwiftPM pin actually publishes;0.0.14was already on pub.dev, so release automation skipped it and Apple builds would have kept theb10514runtime. -
Corrected the WebGPU bridge docs, which claimed the pinned
v0.1.37bridge assets match the default native llama.cpp runtime. They embedb10514and now trail the nativev0.2.0-1pin. -
Chat-template capability detection now logs a debug message naming the probe (
string-content,typed-content,system-role,tools) when its render throws, so a template that fails to render is distinguishable from one that genuinely lacks the capability.
0.8.20#
-
Updated WebGPU bridge assets to
v0.1.37(llama.cppb10514), restoring native/Web parity and provisioning an explicit 1 MiB Wasm stack for wasm32 and memory64 so Qwen3-ASR memory64 context construction does not overflow the default stack. -
Improved Web microphone transcription with browser-capture warmup trimming and early short, silent, and unsupported PCM WAV diagnostics.
-
Added logical and micro-batch controls for llama.cpp/WebGPU models to the Flutter chat example.
-
Made Android Auto probe the packaged Vulkan device before choosing GPU offload, avoiding unnecessary CPU fallback on capable models.
-
Updated the native LiteRT-LM runtime to
v0.16.0-native.2; the Apple companion packages the iOS Gemma constraint provider and Metal plugins required by the published runtime. -
Added an experimental
SpeechToTextEngine.liteRtLmpath with bounded mono 16 kHz float PCM, partial/final transcript events, worker-isolated CPU inference, backpressure, and cancellation. -
Added experimental live English dictation to native Flutter chat models, including generic audio-chat models, using selectable checksum-pinned Moonshine Tiny (recommended, 54 MB) and Parakeet TDT 0.6B (optional, 615 MB) LiteRT sidecars. Live dictation is CPU-only, English-only, capped at five minutes, and unavailable on Linux and Web. Audio-chat models retain Ask with voice as a separate action.
-
Improved Flutter chat example model downloads with bounded retries for transient network failures, safe resume after truncated responses, and a distinct integrity-verification state after transfer reaches 100%. The redesigned onboarding and Lab surfaces preserve model-card position while downloads reorder and keep streaming responses from pulling users away from chat history.
-
Added experimental typed Qwen3-TTS synthesis on native llama.cpp and WebGPU with capability discovery, cancellation, speaker references, complete PCM/WAV output, and synthesis/playback/export controls in the Flutter chat example. Apple apps discover the TTS ABI in the embedded llama framework; current LiteRT-LM artifacts remain unsupported.
-
Updated the native llama.cpp runtime to
b10514, adding BailingMoE3, GraniteSWA/GraniteMoeSWA, speculators-format DSpark checkpoints, and current upstream multimodal/backend fixes and performance improvements. Matching Dart FFI bindings and the Apple SwiftPM artifact checksum were refreshed. -
Added an experimental typed Qwen3-ASR whole-file transcription workflow, a checksum-pinned Qwen3-ASR 0.6B native-and-Web chat-app preset, and file and microphone transcription. Web accepts WAV bytes only; native LiteRT-LM live dictation remains a separate implementation.
-
Added Ask with voice to the native Flutter chat example for Gemma 4 E2B LiteRT-LM and audio-capable GGUF models. It sends a short microphone recording through normal multimodal chat so the model answers the spoken request, while remaining separate from typed speech-to-text.
0.8.19#
-
Updated the native llama.cpp runtime to
b10333and WebGPU bridge assets tov0.1.27(llama.cppb10333), including matching Dart FFI bindings and refreshed Apple SwiftPM artifacts. -
Fixed corrupt Qwen3.5 output on Android Vulkan by preserving the KQV offload required for correct hybrid model inference while retaining the remaining conservative Android context settings.
0.8.18#
-
Updated native llama.cpp to
b10276, including Qwen3-TTS model-loading primitives, explicit bundled-MTP loading, automatic model-specific token suppression, recent model/runtime improvements, the matching load-mode and penalty-sampler ABI migrations, and refreshed Apple artifacts. Speech generation is not yet exposed through the public Dart API. -
Updated the default LiteRT-LM runtimes to native
v0.15.0-native.3and Web@litert-lm/core@0.15.0. The native artifact includes a corrected v0.15 streaming callback bridge and an Android Dawn rollback for Mali-G715 GPU device loss; incompatible callback runtimes now fail safely before generation, and concrete macOS app, framework, and cache libraries take precedence over process-linked assets. -
Updated WebGPU bridge assets to
v0.1.26(llama.cppb10276), refreshing both WebAssembly runtimes while preserving the existing bridge API. -
Disabled automatic WebGPU fetch-backed model loading by default. Streamed loading remains the safe default, with explicit opt-in available for controlled range-capable deployments.
0.8.17#
-
Updated the default llama.cpp native runtime to
b10075with matching bindings and Apple artifacts. - Added Tencent Hunyuan V3 chat-template, reasoning, and tool-call support.
- Fixed Gemma 4 LiteRT-LM text generation in the Web chat app after model loading completed successfully.
- Restored GGUF loading in deployed Web chat apps by packaging the pinned WebGPU runtime assets with Flutter Web builds.
0.8.16#
-
Updated the default llama.cpp native runtime to
b9982, including safer multimodal UTF-8 prompt handling and refreshed Dart/SwiftPM bindings. -
Improved llama.cpp batching defaults and
ChatSessioncontext management for more predictable generation under constrained contexts. - Added llama.cpp presence-penalty sampling and thinking-budget controls; unsupported WebGPU and LiteRT-LM paths now fail explicitly.
- Reworked the runnable TUI coding agent with a focused Pi-style workflow, Unsloth Qwen3.6 defaults, shared model-source loading, and clearer output.
-
Improved the OpenAI-compatible server with standard client-managed tool-call
transcripts, named
tool_choice, configurable thinking behavior, and shared model-source loading.
0.8.15#
-
Added clipboard media attachments to the runnable chat app, including
Cmd/Ctrl+Vscreenshots and copied image/audio files on desktop and web plus a Paste attachment action on mobile, while preserving normal text paste. - Added an app-owned FIFO model-download queue to the runnable chat app, with per-card queue positions/cancellation and a responsive shell progress pill that remains visible after settings closes.
- Replaced the runnable chat app's broad built-in catalog with a focused Unsloth-first set. Added cross-platform Gemma 4 E4B plus native-desktop Gemma 4 12B/26B-A4B/31B and Qwen3.6 35B-A3B presets. Added downloaded-first ordering, name/capability search, Mobile & Web/Desktop filters, clearer incompatible-model states, quieter model cards, and independent remove-from-library and downloaded-file actions for custom entries.
-
Enabled Gemma 4 audio attachments in the runnable chat app for the current
native GGUF projector and LiteRT-LM bundle while keeping LiteRT-LM Web
text-only. Persisted capability settings now distinguish direct model media
input from external
mmprojinput. - Made native GGUF Auto tuning model- and memory-aware. Max now requests full llama.cpp offload, while Auto preserves the requested context when the model fits and reduces context before selecting partial offload under memory pressure. Auto intent persists separately from resolved values so each model load, including after an app restart, recalculates current device headroom.
- Fixed native LiteRT-LM response limits being misapplied as forced benchmark decode counts. Short answers now stream without waiting for the full token allowance, and the chat app flushes the first LiteRT-LM token immediately.
-
Updated the default native LiteRT-LM runtime to
v0.14.0-native.2, fixing Android GPU plugin symbol resolution and using the checksum-pinned official Apple XCFrameworks for Metal-capable iOS and macOS packaging.
0.8.14#
-
Improved the runnable chat app's web model-cache and loading experience by reusing cached GGUF and LiteRT-LM bundles, preserving browser model caches during app cache cleanup, allowing text-only downloads without a multimodal projector, and polishing download/load progress states.
-
Updated the chat app's web runtimes to pinned WebGPU bridge assets
v0.1.18(llama.cppb9915) and@litert-lm/core@0.14.0for reproducible hosted and local inference. -
Updated the default native llama.cpp runtime to
leehack/llamadart-native@b9935, regenerated matching Dart FFI bindings, refreshed the Apple SwiftPM companion checksum, and aligned current runtime documentation.
0.8.13#
-
Fixed load lifecycle guards so repeated model loads preserve the active model state, URL load unsupported-runtime diagnostics stay typed, and unload cancels active generation before freeing llama.cpp handles.
-
Tightened LiteRT-LM runtime validation and local smoke coverage by requiring complete macOS arm64 runtime caches, adding the missing iOS-compatible SwiftPM Gemma provider target, and keeping Flutter macOS LiteRT-LM companion-package builds on hook-managed native assets while the current SwiftPM artifact set is incomplete.
-
Reworked the README into a shorter entry point, fixed stale docs/examples found during the documentation review, and aligned release, Android smoke, WebGPU mem64, native sync, and capability-support wording with the current workflows and runtime behavior. WebGPU runtime LoRA calls now throw an unsupported-operation error instead of reporting no-op success.
-
Fixed llama.cpp n-gram speculative configuration mapping so
draftTokenMaxno longer implicitly overrides upstreamngramSizeM, and documented upstream comparison commands plus measured n-gram benchmark results. -
Added llama.cpp upstream speculative decoding parity through
SpeculativeDecodingConfigconstructors for draft-simple, EAGLE3, MTP, DFlash, ngram-simple, ngram-map-k, ngram-map-k4v, ngram-mod, ngram-cache, and mixed n-gram plus one draft-model strategy, including generic native wrapper bindings, docs, and local benchmark matrix coverage. The benchmark tooling can generate a llama.cpp-compatible static n-gram cache file forngram-cacheE2E validation. Speculative benchmark prompts now render with configured or loaded GGUF chat templates before the generic fallback instead of silently falling back to a hard-coded Gemma prompt. -
Documented compatible DFlash GGUF metadata, a known-good public target/draft model pair, and troubleshooting guidance for incompatible
dflash-draftor missingdflash.target_layersartifacts. -
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9873-llamadart.2, keeping theb9873llama.cpp ABI/bindings while picking up wrapper fixes for native release provenance and backend-selected speculative sampler acceptance. Refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum and aligned current README/website native override docs. -
Hardened LiteRT-LM generation validation so llama.cpp-only speculative decoding knobs fail loudly instead of silently degrading to LiteRT-LM's boolean speculative toggle.
-
Added
LlamaStructuredOutputandLlamaEngine.createStructuredJson(...)helpers for strict JSON-object / JSON-schema generation with final-output validation and typed decoding. -
Added
LlamaEngine.loadMultimodalProjectorSource(...)so GGUF multimodal projector files can use the sameModelSourceresolver and native download/cache options asloadModelSource(...). -
Improved the runnable chat app's Manage Models cache UX so model and mmproj asset cache states are shown separately, missing multimodal projectors can be re-cached without re-fetching already cached model assets, and runtime media capability mismatches surface as user-readable warnings. Custom signed or tokenized Hugging Face URLs now require confirmation before they are saved.
0.8.12#
-
Updated the default LiteRT-LM native runtime pin to
leehack/litert-lm-native@v0.14.0-native.1, refreshed native-assets and Apple SwiftPM checksums, and exposed the new native LiteRT-LM 0.14 load/generation controls. -
Hardened LiteRT-LM 0.14 runtime packaging across Linux, Windows, Android, Apple SwiftPM, and macOS local runtime-prep paths.
-
Hardened release automation by adding CODEOWNERS coverage for publication-sensitive files and making pub.dev/GitHub Release propagation waits configurable with longer defaults.
-
Added post-merge release automation so a merged release-prep PR can publish missing companion package versions, push the core release tag, wait for pub.dev, and confirm the GitHub Release without a separate manual tag step.
Added
LlamaEngine.getModelFileType()for llama.cpp/GGUF models.-
Refreshed llama.cpp
b9860native runtime pins and template parity for DeepSeek V4 and MiniCPM5, including MiniCPM5 XML tool-call handling.
0.8.11#
-
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9829, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and aligned current README/website native override docs for the release.
0.8.10#
-
Potentially breaking behavior change: native model cache defaults changed without breaking Dart source compatibility.
DefaultModelDownloadManager()now prefers the platform shared cache on desktop/server instead of the process temp directory, and mobileDefaultModelDownloadManager.auto()without an explicit app-private directory now uses a best-effort temporary/cache fallback instead of throwing. Apps or tests that asserted the old temp path or mobile exception should pass an explicit cache directory or followMIGRATION.md. -
Added optional
androidAppPrivateCacheDirectoryandiosAppPrivateCacheDirectoryarguments toDefaultModelDownloadManager.auto(...)so apps can provide platform-specific mobile cache roots without constructor-levelPlatform.isAndroid/Platform.isIOSbranching. -
Updated the default native
DefaultModelDownloadManager()constructor to use the per-user shared model cache on desktop/server platforms and the mobile app-private cache fallback, so plainLlamaEngine(...)remote source loads use a platform-appropriate default while preserving a temporary fallback for hosts that cannot expose a desktop cache environment.
0.8.9#
-
Broadened the
hooksdependency constraint to support both the existing build-hooks package family and the latest stable release, restoring the pub.dev dependency freshness score without breaking downstream packages that still resolvehooks1.x. -
Made web-safe backend stubs the default conditional import/export targets, preserving native
dart:ioselection while avoiding false WASM compatibility deductions in pub.dev analysis.
0.8.8#
-
Added a CI release-doc version consistency check so current README/website install snippets and companion package READMEs stay aligned with package
pubspec.yamlversions, and documented that companion/core package publishing happens only after release-prep merge with explicit maintainer approval for each tag. -
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9803, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and aligned current README/website native override docs. -
Added
DefaultModelDownloadManager.auto(...)plus explicit model cache root constructors for shared desktop caches, app-private mobile caches, user-selected model libraries, and App Group containers. Implicit shared cache resolution now fails loudly on mobile and web where the OS cannot provide a hidden cross-developer model folder.
0.8.7#
-
Fixed multimodal chat-template rendering so templates that force-open
reasoning, such as Qwen3.5 VLM prompts ending with
<think>, preserveenable_thinkingand stream generated reasoning throughdelta.thinkinginstead ofdelta.content.
0.8.6#
-
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9776, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and aligned README/website native override docs.
0.8.5#
-
Fixed the split-library mtmd fallback ABI for image and byte-buffer
multimodal inputs so Windows
mtmd.dlland other split mtmd native bundles use the same bitmap helper signature as the generated native binding path. This avoids corrupting the first mtmd bitmap-helper call for Gemma 4/MMProj style multimodal loads and adds native symbol regression coverage for the fallback ABI.
0.8.4#
-
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9744, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and aligned current README/website native override docs. -
Expanded llama.cpp chat-template parity coverage for current upstream
fixtures, including Cohere2 MoE, LFM2.5 tool-call, and Granite 4.1 templates.
LFM2.5 prompts that use plain
List of tools: [...]now route through the LFM2 handler, andToolChoice.requireduses grammar-constrained LFM2 tool-call generation. - Fixed streaming tool-call parsing so partial North/Cohere bare action arrays are not emitted as content before the complete tool call is parsed, and expanded the local GGUF feature smoke coverage for thinking, tool-call, and optional multimodal turns.
0.8.3#
-
Fixed Windows CUDA backend discovery when the native asset bundle directory is
not on the app
PATH. Apps using the CUDA llama.cpp backend can now resolve bundled CUDA redistributables besideggml-cuda.dllwithout manually adding.dart_tool/libor the native bundle path toPATH.
0.8.2#
-
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9694, regenerated matching Dart FFI bindings, refreshed thellamadart_llama_cpp_flutterApple SwiftPM checksum, and updated the default WebGPU bridge asset pin toleehack/llama-web-bridge-assets@v0.1.17(llama.cppb9699). The WebGPU backend now caps unset large-model browser batches so Gemma 4 mem64 loads do not fall back to context-sized compute buffers. -
Added
BackendGpuEnumeration.listGpuDevices({probeBackends})andLlamaEngine.listGpuDevicesso apps can enumerate GPU-class devices and select llama.cpp offload targets by backend-specificmainGpuindex. -
Added Cohere2 MoE / North Code chat-template detection and parsing so
<|START_TEXT|>responses and<|START_ACTION|>tool-call arrays are handled separately from older Command-R templates.
0.8.1#
-
Fixed docs references that still pointed at
llamadart_litert_lm_flutter0.0.1and the pre-native.1LiteRT-LM release after the 0.8.0 native pin sync moved LiteRT-LM Apple/runtime artifacts tov0.13.1-native.1. -
Routed native
.litertlmimage/audio chat parts through LiteRT-LM Conversation message JSON so bundles with native media processors can acceptLlamaImageContent/LlamaAudioContentpath and encoded-byte inputs without a separatemmprojprojector.
0.8.0#
-
Split Flutter Apple SwiftPM runtime linking into companion packages:
llamadart_llama_cpp_flutterfor GGUF/llama.cpp andllamadart_litert_lm_flutterfor.litertlm/LiteRT-LM. The core package remains a native-assets package without Flutter plugin metadata; the companion package sources live underpackages/in this repository. -
Changed unset or empty
llamadart_native_runtimesto mean all available runtime families. Flutter iOS/macOS companion packages decide Apple SPM runtimes when present; other builds continue to usellamadart_native_runtimes. -
Added opt-in native
.litertlmModelParamsfor activation data type, prefill chunk size, parallel file-section loading, and Android NPU LiteRT dispatch library directory, forwarding the pinned LiteRT-LMv0.13.1-native.1engine-settings C APIs while keeping defaults unchanged. - Extended the LiteRT-LM engine smoke tool with matching environment variables and documented the support decision for each candidate runtime knob.
- Kept LiteRT-LM web rejecting these native-only settings explicitly.
- Added llama.cpp MTP benchmark diagnostics and local smoke/benchmark tools so baseline-vs-MTP runs can report decode timing, draft/accepted token counts, draft verification timing, and acceptance rate.
-
Added
SpeculativeDecodingConfig.mtp(draftModelPath: ...)for llama.cpp external draft-model MTP sessions. - Removed the Android Vulkan MTP allow-list dart define and the model-name based Android Vulkan acceleration shortcut. Vulkan MTP now runs only when callers explicitly request Vulkan plus MTP in runtime parameters.
0.7.2#
- Added explicit pub.dev platform metadata for Android, iOS, Linux, macOS, web, and Windows. This keeps the package listing aligned with the actual cross-platform runtime support even though Flutter plugin registration is only needed for Darwin app integration.
0.7.1#
-
Added Flutter iOS/macOS Swift Package Manager integration so Apple apps link
pinned
leehack/llamadart-nativeandleehack/litert-lm-nativeXCFramework artifacts throughdarwin/llamadart/Package.swift. -
Disabled the legacy hook-managed Apple bundle path for Flutter iOS/macOS
builds, avoiding wrapper/framework
MinimumOSVersionmismatches in App Store uploads. - Raised the Flutter Apple runtime floors to iOS 16.4 and macOS 14.0 to match the published XCFramework artifacts.
-
Kept Android native builds on both
llama_cppandlitert_lmby default; iOS, macOS, Linux, and Windows now default tollama_cpponly. Non-Android.litertlmapps should opt in withllamadart_native_runtimes. - Added native release pin automation for Apple SPM checksums, excluded local SwiftPM artifact caches from pub archives, and hardened main-branch CI against Hugging Face tiny-model download rate limits.
- Compatibility note: no Dart API breaking changes. Flutter Apple apps must target iOS 16.4/macOS 14.0 or newer.
0.7.0#
-
Added LiteRT-LM as a first-class backend for native
.litertlmbundles and single-turn web-compatible.litertlmURLs, alongside the existing llama.cpp/GGUF path. -
Added
ModelParams.liteRtLmBackendso callers can select LiteRT-LM CPU, GPU, or Android NPU execution where the pinned runtime supports it. - Added native LiteRT-LM tokenization, detokenization, log-level control, runtime metrics, cached Hugging Face loading, and package hook overrides for testing compatible native runtime sources.
-
Added
GenerationParams.speculativeDecodingfor native LiteRT-LM and wired the benchmark app so speculative runs are reflected in metrics. -
Fixed Gemma 4
.litertlmthinking and tool calling with canonical templates, thought-channel parsing, reasoning suppression, and a filename-keyed template registry for Gemma and Qwen LiteRT-LM bundles. -
Fixed iOS
.litertlmloading by resolving embeddedLiteRtLmandStreamProxyframeworks from the app bundle. -
Added WebGPU mem64 selection through
ModelParams.preferMemory64andModelParams.modelBytesHintso large GGUF models such as Gemma 4 E2B can choose the 64-bit bridge core. - Fixed chat-app web downloads, LiteRT-LM web loading/generation, unsupported token-count refreshes, and misleading LiteRT-LM load progress.
- Hardened native and LiteRT-LM cancellation/disposal, multimodal cleanup, parser correctness, grammar generation, model download timeouts, and partial download resume behavior.
- Added Gemma 4 benchmark tooling, GGUF chat-feature smoke coverage, and the WebGPU Gemma 4 mem64 E2E scenario.
- Updated README and website docs for backend choice, capability limits, platform support, package-size controls, benchmark results, model templates, and pinned runtime artifacts.
-
Compatibility note: no public API breaking changes for existing GGUF /
llama.cpp callers. LiteRT-LM support is additive, with deprecated benchmark
wrappers retained for compatibility; unsupported llama.cpp-only parameters are
rejected for
.litertlmloads instead of being silently ignored.
0.6.17#
-
Synced native hook pinning and regenerated bindings through
leehack/llamadart-native@b9371, picking up llama.cppb9371. -
Picked up the Apple mobile Metal stability fix that disables Metal residency
sets on iOS/tvOS/visionOS native bundles, avoiding affected device
context-creation failures such as
MTLLibraryErrorDomain Code=3. -
Compatibility note: no public API breaking changes in
0.6.17; existing0.6.16callers remain compatible.
0.6.16#
-
Fixed native
getVramInfo()so llama.cpp GPU-class backend devices can report free/total VRAM when available, with Windows split-bundle registry fallback handling for backend-device symbols. - Improved browser recovery for large remote WebGPU model/projector loads by retrying wasm32 model-staging aborts with the wasm64 core before surfacing memory-pressure failures.
-
Improved the runnable chat app's web remote-model startup path so model assets
are prefetched into browser cache when available, browser
CacheStoragefailures fall back to direct network loading, and credentialed/signed model URLs skip persistent browser cache storage. - Improved the runnable chat app's mobile download behavior so lifecycle pauses no longer deliberately cancel active foreground downloads; the app now lets short screen-lock/background interruptions continue when the OS permits and still keeps explicit pause/dispose cancellation paths.
- Added in-app and docs guidance for mobile large-model downloads, including resumable partial files, foreground Dart lifecycle limits, and the need for opt-in native background download/model-store integrations for robust cross-app GGUF management.
-
Compatibility note: no public API breaking changes in
0.6.16; existing0.6.15callers remain compatible.
0.6.15#
- Fixed GLM-OCR and other multimodal chat-template workarounds so image and audio content parts are preserved when tool-call normalization runs, system prompts are merged before leading media parts, and invalid tool-call serialization fails loudly instead of silently falling back to the wrong template shape.
-
Added
tool/testing/run_local_e2e.dartas a discovery and orchestration entry point for heavyweight local-only Dart E2E, Flutter device, and Web/Playwright smoke scenarios. -
Hardened the upstream llama.cpp chat/template E2E runner against current
llama.cpp target renames, dynamic backend library lookup, and full
test-chatserver/mtmd build requirements. -
Documented that real-model/device/WebGPU scenarios remain skipped from
default CI and should be opted into explicitly with
--listand--dry-runfirst. -
Compatibility note: no public API breaking changes in
0.6.15; existing0.6.14callers remain compatible. The chat-template changes fix multimodal serialization behavior for affected templates, and the local E2E runner is additive.
0.6.14#
-
Updated the default WebGPU bridge asset pin to
leehack/llama-web-bridge-assets@v0.1.16(llama.cppb9165), picking up the published JS bridge build, TypeScript declaration asset, and refreshed bridge docs. - Added WebGPU readiness guidance covering browser capability checks, cross-origin isolation, bridge asset/version diagnostics, fallback behavior, model/configuration pressure, and the Flutter Web real-model smoke path.
-
Added
ModelDownloadController, a dependency-free helper that turnsModelDownloadManagercache/download work into app-facing lifecycle states for resolving, cache checks, downloads, verification, ready, failed, cancelled, and retry flows. -
Wired the runnable chat app example through a
ModelDownloadManageradapter so its model-management UI demonstrates the controller while preserving the example's multi-asset and web-cache service behavior. -
Compatibility note: no public API breaking changes in
0.6.14; the WebGPU bridge asset update andModelDownloadControllerare additive, and existing0.6.13callers remain compatible.
0.6.13#
-
Added package-managed model source downloads and cache management:
ModelSource,ModelLoadOptions,ModelCachePolicy, resolver targets, download/cache metadata, progress callbacks, cache inspection, removal, clearing, and age/size pruning. -
Added native/file-backed
DefaultModelDownloadManagersupport for streaming HTTP downloads,.partfiles with atomic promotion, authenticated bearer and custom headers, cooperative cancellation, retry, HTTP Range resume, cache hit/refresh/cache-only/no-cache policies, SHA-256 verification, and persisted redacted metadata for signed URLs. -
Improved Hugging Face
hf://ergonomics with?revision=...parsing for branch/ref names containing slashes, plus docs for private/gated bearer-token usage, separatemmprojassets, sharded-GGUF limitations, and redaction guarantees. -
Hardened download/cache correctness by serializing concurrent same-entry
downloads, recovering missing or malformed cache metadata sidecars, treating
mismatched byte-count/SHA-256 metadata as cache misses, and rejecting
remote-only options for local
ModelSource.path(...)inputs. -
Added
LlamaEngine.loadModelSource(...)so local path sources keep using the existing native loader, remote HTTP(S)/Hugging Face sources download through the package-managed native cache before local loading, and URL-capable web backends keep using direct URL loading for simple unauthenticated requests. -
Added KV-cache state persistence APIs:
LlamaEngine.supportsStatePersistence,stateSaveFile(...),stateLoadFile(...), backend support diagnostics, and WebGPU bridge forwarding for bridge assetsv0.1.15+. -
Compatibility note: no public API breaking changes in
0.6.13; existingloadModel(...)callers are unchanged.
0.6.12#
-
Synced default WebGPU bridge asset pinning to
leehack/llama-web-bridge-assets@v0.1.14(llama.cppb9016) to match the native runtime pin. - Picked up bridge-side Qwen UTF-8 streaming stabilization and multimodal fallback narrowing while preserving control-token output for parser consumers.
- Picked up the bridge-side BERT embedding thread-pool sizing fix so automatic thread selection does not exceed the compiled WebAssembly pthread pool.
-
Forwarded native-compatible
ModelParamsload tuning knobs through the WebGPU bridge path, including sequence slots, flash attention, KV cache type, RoPE overrides, split mode, and main GPU. -
Matched native batch defaults on WebGPU so unset
batchSizeandmicroBatchSizeusen_batch = n_ctxandn_ubatch = n_batch, avoiding first-embedding aborts for BERT-class/non-causal encoder models while preserving model-specific Qwen3.5-0.8B WebGPU safety tuning. - Filtered backend-owned runtime dependencies during native asset bundling so CUDA runtime DLLs and OpenBLAS runtime libraries are emitted only when their owning backend module is selected, while unknown runtime libraries stay bundled for forward compatibility.
- Compatibility note: no public API breaking changes in
0.6.12.
0.6.11#
-
Synced native hook pinning and regenerated bindings through
leehack/llamadart-native@b8955. -
Fixed Gemma 4 streaming so
<|channel>thought ... <channel|>output is emitted as thinking deltas instead of content text, including when channel markers are split across streamed chunks. - Tracked the chat app lockfile for stable generated Flutter plugin metadata in CI and release validation.
- Compatibility note: no public API breaking changes in
0.6.11.
0.6.10#
-
Synced native hook pinning and regenerated bindings through
leehack/llamadart-native@b8638. -
Hardened multimodal prompt overflow handling so native failures surface as
Dart exceptions, and reduced staged chat-app image size to a
384pxmax edge to lower multimodal context pressure. - Added built-in Gemma 4 template detection/render/parse support, including thinking and tool-call handling.
-
Added runtime projector capability gating so multimodal flows and the chat app
respect actual
supportsVision/supportsAudioresults instead of model-family assumptions. - Compatibility note: no public API breaking changes in
0.6.10.
0.6.9#
-
Documented that iOS builds require a minimum deployment target of
16.4or newer across the README, docs site, and example docs. -
Updated
example/chat_appiOS Podfile and Runner project settings to use deployment target16.4. -
Honored
ggml_backend_scoreduring Android asset-based backend fallback so unsupported CPU variant libraries are skipped before initialization. -
Changed Android
autobackend resolution to prefer CPU by default while keeping Vulkan available for explicit opt-in. -
Clarified that changing
hooks.user_definesrequiresflutter clean && flutter pub getbefore rebuilding. - Compatibility note: no public API breaking changes in
0.6.9.
0.6.8#
-
Synced native hook pinning and regenerated bindings to
leehack/llamadart-native@b8480. - Refreshed generated low-level FFI bindings to match the synced upstream headers.
- Compatibility note: no public API breaking changes in
0.6.8.
0.6.7#
-
Synced native hook pinning and regenerated bindings to
leehack/llamadart-native@b8373. -
Hardened Linux bundle loading for packaged apps and improved versioned
libllamadartdependency resolution. -
Fixed Hermes tool-call parsing when whitespace appears between
<tool_call>and the JSON payload. - Compatibility note: no public API breaking changes in
0.6.7.
0.6.6#
- Synced native hook pin to
leehack/llamadart-native@b8216. -
Updated default web bridge asset pinning to
leehack/llama-web-bridge-assets@v0.1.10(llama.cppb8216). - Switched bundled Qwen3.5 example presets to Unsloth
Q4_K_MGGUFs. -
Added native perf diagnostics chips in the chat app (
p_eval,eval,sample,reuse) and Android-specific Qwen tuning guidance. -
Restored a targeted Android Vulkan fast path for local Qwen3.5
0.8B/2B/4Bmodels while keeping CPU as the recommended Android preset for0.8B/2B. - Fixed local web chat app bridge/runtime handling for Qwen prompt streaming and multimodal fallback behavior.
- Compatibility note: no public API breaking changes in
0.6.6.
0.6.5#
-
Added embedding APIs:
LlamaEngine.embed(...)andLlamaEngine.embedBatch(...). - Added backend embedding capability interfaces for custom backend implementations.
-
Added multi-sequence embedding batching support via
ModelParams.maxParallelSequences(n_seq_max). -
Added native embedding benchmark tooling:
tool/testing/native_embedding_benchmark.dartandtool/testing/native_embedding_sweep.dart. - Added website docs for embeddings and updated basic-app docs with embedding examples.
-
Added a Basic App SQLite vector retrieval example using
bin/llamadart_sqlite_vector_example.dart. -
Updated default WebGPU bridge asset pinning to
leehack/llama-web-bridge-assets@v0.1.8. - Improved WebGPU runtime stability/tuning in chat app flows (backend switching, streaming smoothness, and multimodal regression gating).
- Added GPU-path multimodal image-size capping to reduce memory/runtime pressure on larger image inputs.
- Compatibility note: no public API breaking changes in
0.6.5.
0.6.4#
- Aligned multimodal projector offload with effective model-load settings, including CPU-only configurations.
- Added safer backend selection/discovery APIs and improved runtime backend status plus GPU-layer diagnostics accuracy.
- Improved web large-model handling with cache-prefetch download UX, bridge worker fallback paths, memory-pressure retries, and wasm64-core fallback wiring.
-
Synced native hook tag to
b8157and added Android arm64 CPU-profile and variant policy support with loader hardening.
0.6.3#
-
Synced native runtime to llama.cpp
b8138and picked up Android arm64 crash/compatibility hardening. - Example app performance/UX polish and web model handling improvements.
-
Added
example/tui_coding_agent, a terminal coding agent example with default stable text-protocol tool mode. - Added persisted settings log-level fallback handling with regression tests.
0.6.2#
- Native inference performance improvements (request overhead, stream batching, and prompt-prefix reuse with parity-safe fallback).
- Added native benchmark and prompt-reuse parity tooling, plus CI parity coverage.
0.6.1#
- Publishing compatibility fix for hook backend-config code paths.
- Continued parity hardening around template/parser behavior.
0.6.x line highlights#
- Expanded llama.cpp template and parser parity.
- Stronger handling for tool payload fidelity.
- More deterministic behavior around template routing and fallback removal.
0.5.x line highlights#
- Public API tightening and migration cleanup.
- Split Dart/native log controls.
- Example/runtime reliability improvements.
Release usage guidance#
- For upgrade planning, combine this page with Upgrade Checklist.
- For breaking changes, always validate against the exact release tag notes.