DeepMind’s New Embedding Model Unifies Five Data Types

DeepMind has widened the embedding map. On October 6, 2026, the company launched EmbeddingGemma 2, an open model that maps text, code, images, video, and audio into one 768-dimensional embedding space. It uses the same technology as Google’s Gemini Embedding models, but is designed to run on consumer hardware — a useful detail in a field that often treats a server rack as a minimum system requirement.
EmbeddingGemma 2 contains 740 million parameters and ships under the Apache 2.0 license. Its 24-layer architecture includes a 270-million-parameter text model, a 170-million-parameter vision encoder, and a 300-million-parameter audio encoder, with a 262,144-token vocabulary and an 8,192-token context window.
The model converts each input type into a shared representation while managing different input sizes: 280 tokens per image, 140 tokens per video frame, and 25 tokens per second of audio. It supports up to 29 images, 58 video frames, or 5.5 minutes of audio per input.
Better code retrieval without losing multilingual performance
Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera are associated with the model’s development. EmbeddingGemma 2 has passed 20 million downloads, and its multilingual text performance matches its predecessor.
The clearest improvement arrives in code. The model raises its MTEB Code score by 9.92 points, from 68.76 to 78.68, while its full-precision results include 61.36 on MTEB multilingual (v2) and 64.64 on MIEB lite.
Its other full-precision scores cover the model’s wider reach: 57.28 on MMEB v2 image retrieval, 67.84 on visual-document retrieval, 50.67 on video retrieval, 69.54 on MSEB sound retrieval, and 49.39 on MAEB audio tasks. One embedding space for five modalities is the headline; the scores suggest DeepMind wants developers to use it for retrieval across mixed media, not just admire the architecture diagram.
AI agents are turning CPUs into strategic hardware
The launch arrives as AI agents reshape the hardware argument. Meta debuted Muse in early September 2026, and the system topped the Apple App Store in less than two weeks after its release. Sensor Tower says Muse has surpassed 5 million downloads since last month’s launch, while Morgan Stanley estimates serving costs between $3 and $130 per month, with an average of $37 per user.
Meta’s system is designed to be CPU-agnostic and runs on an AMD-powered computer. OpenAI’s Dots runs on a virtual computer powered by an AMD EPYC-branded CPU, while most hyperscaler servers use CPUs made by Intel or AMD.
“As more agents are created and developed, and more people start to use them for more tasks, it’s going to start to shift the workload away from GPUs and onto CPUs,” said Ryan Shrout, president of Signal65. Daniel Newman put the division more plainly: “CPUs are actually performing the workflows while GPUs are doing the thinking.”
AMD’s numbers show why that shift matters. Its data center revenue more than doubled to $6.7 billion in the quarter ended June, and that business accounted for almost 60% of total sales. AMD’s market capitalization has entered the trillion-dollar market cap club, and the company currently commands about 46% of x86 CPU units.
Nvidia released Vera, a fully redesigned central processor, earlier this year, with Vera CPUs built specifically for agents. Nvidia expects CPUs to become a $200 billion market by 2030; Arm announced its own CPU for agents in March 2026, with Meta as the debut customer.
Forecasts vary, because apparently even silicon needs competing spreadsheets. Futurum Group estimates total CPU sales will reach $118 billion in 2027, while AMD predicts the entire CPU market will hit $220 billion in 2030 and expects to take more than half of it.
AMD CEO Lisa Su delivered a speech on May 22, 2026, as the company positioned itself across CPUs and GPUs. Meta’s spokesperson said, “We take a diverse approach to our hardware and are largely CPU-agnostic by design, which gives us the most flexibility in acquiring capacity.”
EmbeddingGemma 2 and the agent hardware race point to the same change: AI systems are becoming multimodal, distributed, and more dependent on the full computing stack. GPUs still handle the thinking, but the CPUs are performing the workflows — and companies are now building products, forecasts, and entire market strategies around that distinction.
Based on



