Changes confirmed medium confidence

Liquid AI Opens d1 Decision Models for Edge Text, Vision and Audio

The open-weight d1-3B and experimental d1-omni-600M turn text, images or audio into structured probabilities without generating a token sequence.

Edited by Tyronne Panaino

Liquid AI released two open-weight d1 decision models on October 7 for developers who need structured answers from text, images or audio without asking a generative model to compose a response. The release adds downloadable d1-3B and the experimental d1-omni-600M to the hosted d1 service the company introduced two days earlier, giving edge developers a new deployment path while leaving several performance questions open.

What changed from the hosted d1 release

The earlier Liquid AI d1 announcement introduced a hosted model that accepts text and images, reads a state and one or more questions in a single forward pass, then returns probabilities rather than generated tokens. It was available through Liquid AI's API.

The newer open d1 release makes two model variants available for download. Liquid AI says d1-3B handles text and images. The smaller experimental d1-omni-600M accepts either text plus images or text plus audio. In both cases, the intended output is a structured decision such as a yes-or-no probability, a choice among labels or a score on a defined scale.

That distinction matters for applications that need a bounded answer rather than free-form prose. A support router can choose a destination, a safety filter can classify a request, and a visual inspection system can score a defined condition without paying for a generated explanation on every call. The release does not show that a decision model can replace a general model across open-ended reasoning or generation tasks.

What the evidence supports

Liquid AI reports evaluation across seven public datasets covering reading comprehension, toxicity detection, intent classification, medical question answering and cross-language understanding. It also reports edge-device timing for d1-3B and says that model answered a single question in less than 50 milliseconds on each measured device. Those are company-run results, not independent replication.

The limits are unusually important here. Liquid AI does not report vision or audio benchmark results in the release. It says the available Decision Index vision split is private and that audio decision benchmarks remain an open problem. It also provides no speed measurements for d1-omni-600M because that model is still an early research release.

The practical conclusion is narrower than the headline benchmark claims. Developers can now inspect and run two d1 weights locally, and the models expose a different interface from token generation. The evidence does not yet establish how reliably the multimodal variant performs across real cameras, microphones, languages or operational edge cases.

Who is affected

The immediate audience is teams building on-device classification, routing and retrieval systems with tight latency or connectivity constraints. Open weights make it possible to test the models against private data and hardware without routing every decision through Liquid AI's hosted endpoint. That can change where a prototype runs, but privacy and reliability still depend on the surrounding application, local storage, logging and access controls.

Teams evaluating d1 should separate three questions: whether the model produces the required answer type, whether it stays accurate on their own distribution, and whether the complete device pipeline meets its latency and memory budget. The release provides a starting point for those tests, not a production guarantee.

Status

Confirmed. Liquid AI and its Hugging Face publication provide first-party evidence that the two open-weight models were released. Confidence is medium because the performance claims and limitations come from the vendor, without independent evaluation in the fetched evidence.

Sources

Update note: Last reviewed October 11, 2026. We will revise this post if independent evaluations or fuller multimodal benchmarks become available.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Changes coverage