Liquid AI’s New Models Make Multimodal Decisions Without Generating Text

Liquid AI has released two open-weight multimodal decision models that answer questions without writing a single token. The new Open d1 family, made up of d1-3B and d1-omni-600M, turns text, images, and audio into calibrated, typed answers through one forward pass.
That changes the role of an AI model. Instead of producing a paragraph and asking another system to interpret it, each d1 model returns a structured decision directly, using zero output tokens. For applications that need fast classifications, rankings, or other typed results, that design targets speed and lower inference costs.
Two Models, Three Input Paths
d1-3B reads text and images, while d1-omni-600M reads text with an image or text with audio. Neither model writes text, and a request carries images or audio, never both. Audio clips can run for up to 30 seconds.
The models also take different architectural paths. d1-3B has 3.12 billion parameters and starts from LFM2.5-VL-3B, a decoder-only vision-language model. It uses a 400 million parameter SigLIP2 NaFlex vision encoder and supports a 32,768-token context.
d1-omni-600M has 587 million parameters and starts from LFM2.5-Encoder-350M, a bidirectional encoder. Its context reaches 16,384 tokens, and its audio system uses a 17-layer FastConformer encoder.
Both checkpoints are available on Hugging Face, load through Transformers, and have day-one llama.cpp support. The LFM Open License v1.0 allows free commercial use below $10 million in annual revenue, giving developers and smaller companies a clear path to test the models in products.
Fast Decisions at the Edge
Liquid AI’s performance figures focus on the time needed to process questions and visual inputs. d1-3B reads a 384-pixel image in 35 milliseconds on Jetson AGX Thor, then answers one question in 16 milliseconds on that platform. The same question takes 8 milliseconds on an RTX 4090.
The model also gains efficiency when several questions use the same state. On Jetson AGX Thor, three questions over one state take 20 milliseconds, compared with 16 milliseconds for one question. That gap points toward workloads where an application needs several decisions from a shared image or situation.
On Decision Index v0.2.1, d1-3B scores 48.57, beating models under 10 billion parameters and edging Decider 35B-A3B at 47.11. Only Winnow-12B scores higher, reaching 50.02.
The model leads the Tools category at 74.5 and the Arts category at 36.3, but trails on Knowledge with 23.8. Across seven public text benchmarks, d1-3B averages 82.9, ahead of Decider 4B at 81.1.
d1-omni-600M averages 78.4 across those seven public text benchmarks, while its strongest individual results include 95.8 on Civil Comments and 79.5 on PAWS-X. On 11 image benchmarks, d1-3B averages 74.1, compared with 73.9 for its base model.
A Bigger Open-Model Race Is Taking Shape
Liquid AI’s release arrives as other companies push open-weight systems toward new combinations of efficiency, scale, and multimodal capability. Reflection AI unveiled its first model, Beam, with scores similar to GLM-5.2; Beam is cheaper to run than many competing models and three to four times as efficient as rival open models from Western companies.
Reflection AI was founded in March 2024 by Misha Laskin, its CEO, and Ioannis Antonoglou, its president and CTO. The company has a $25 billion valuation and entered a partnership in May 2024 to supply its models to the U.S. Department of Energy and the U.S. Department of War.
Reflection AI described Beam this way: “Through high-compute reinforcement learning, we were able to make Beam extremely efficient at reasoning—delivering competitive performance on coding and agentic tasks at a fraction of the token cost and inference time compute.”
Mistral is also launching Mistral Large 4, a one-trillion-parameter multimodal foundation model nicknamed “Le Chonk.” ML4 has 49 billion active parameters during inference and was trained over roughly two months on 4,000 Nvidia Grace Blackwell GPUs, compared with 3,000 Nvidia H200 GPUs used for a previous model.
ML4 was trained across more than 160 languages, including every official language of the European Union. It accepts multimodal inputs but produces text output, while its weights are scheduled for publication on October 27, 2026, after a testing period. Guillaume Lample, Mistral co-founder and chief scientist, said, “ML4 is at the frontier of open weight models.”
Its benchmark results include 62% on DeepSWE v1.1, 15% on Harvey’s Legal Agent Benchmark, 67% on Finch, 42% on Dense200, and 73% on DIOR-RSVG. ML4 does not yet appear in Artificial Analysis’ public evaluations or the DeepSWE leaderboard, so its claim as the strongest open-weight model outside China remains provisional.
Mistral was founded in 2023 by Arthur Mensch, Guillaume Lample, and Timothée Lacroix, and its first model, Mistral 7B, arrived in September 2023. The company now spans developer and enterprise products, customization services, inference infrastructure, and Mistral Compute.
The contrast is clear: ML4 expands multimodal models toward enormous scale and text generation, while Open d1 targets direct decisions with no output tokens. With open checkpoints, edge-ready runtimes, and measured response times, Liquid AI is betting that the next useful AI answer will not always be a sentence.
Based on
- Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens — marktechpost.com
- Fortune Tech: The U.S. has a new horse in the race to build the best open-source AI model | Fortune — fortune.com
- Mistral debuts Large 4 ‘Le Chonk’, a 1-trillion parameter text output model with high benchmarks planned for open weights release | VentureBeat — venturebeat.com




