Multimodal AI
-
Artificial Intelligence
Voice AI Has a Conversation Problem: Speed Versus Substance
Voice AI is getting more natural—and more complicated. ChatGPT Voice and Gemini Live now compete on the same basic promise:…
Read More » -
AI in Business & Enterprise
Gemini Live Agents Gain a New Visual Presence
Google is giving its live AI agents a face. Gemini 3.8 Live now adds Live Avatar, a feature that combines…
Read More » -
Consumer Technology
How Cal AI Turned a Meal Photo Into a $30 Million Business
A meal photo became the starting point for a technology business with more than 15 million downloads and over $30…
Read More » -
Generative AI
Alibaba’s Qwen3.8-LiveTranslate Targets Faster Real-Time Speech Translation
Alibaba’s Qwen team released Qwen3.8-LiveTranslate on September 19, 2026, introducing a next-generation model for real-time simultaneous interpretation. It listens to…
Read More » -
AI Agents & Automation
Qwen’s New Omni-Modal Agent Turns Audio and Video Into Action
Alibaba’s Qwen team has released Qwen3.8-Omni-Flash, a model that brings audio-video understanding, reasoning, and tool use into one system. The…
Read More » -
Large Language Models
DeepSeek’s New Flash Model Shrinks the Cost of Long Context
DeepSeek released V4.1-Flash on September 14, 2026, introducing a model built around one central goal: making long-context artificial intelligence cheaper…
Read More » -
Artificial Intelligence
Gemini’s New Video Agent Cuts Tokens Without Cutting Corners
Google is making video analysis less wasteful. This week, the company launched agentic video understanding across its Gemini Flash models,…
Read More » -
Machine Learning & Research
NeoMME Builds Multimodal Encoders From Scratch
NeoMME starts from zero. H Company researchers released the open-source family on September 3, 2026, with 260M- and 800M-parameter multimodal…
Read More » -
Machine Learning & Research
One Transformer Takes On Text, Images, and Document Search
NeoMME puts text and images inside one encoder. The system introduces multilingual multimodal models in 260M and 800M sizes, without…
Read More » -
Generative AI
Gemini Omni 1.1 Turns Video Generation Into Shot-Level Direction
Google is making video generation more controllable. On August 29, 2026, the company released Gemini Omni 1.1 Flash, a production…
Read More »