Skip to content

Google Launches EmbeddingGemma 2 for Multimodal Embeddings

Google announces EmbeddingGemma 2, a lightweight multimodal embedding model for on-device processing of text, images, audio, and video.

Sources in the coveragePerformance and quality statements from a vendor are claims unless the coverage cites independent measurements.

Evidence & verification 4 statements

Specifications, source statements and editorial judgement have different scopes. A citation is not an independent performance test.

Model ParametersSource statement

740 million

deepmind.google ↗Source observation date not recorded.
ArchitectureSource statement

Gemma 4

deepmind.google ↗Source observation date not recorded.
LicenseSource statement

Apache 2.0

deepmind.google ↗Source observation date not recorded.
ModularitySource statement

270M parameters for text-only, 170M for vision, 300M for audio

deepmind.google ↗Source observation date not recorded.

Google has released EmbeddingGemma 2, an open-source multimodal embedding model designed for on-device processing. It supports combinations of text, images, audio, and video, and is optimized for efficiency and performance on consumer hardware.

Key facts

Fact Detail The source says
Model Parameters 740 million “EmbeddingGemma 2 has 740 million parameters, making it optimal for on-device inference.”
Architecture Gemma 4 “Built on the Gemma 4 architecture”
License Apache 2.0 “released under a commercially permissive Apache 2.0 license”
Modularity 270M parameters for text-only, 170M for vision, 300M for audio “Requires as little as 270M parameters for text-only workloads with optional vision (170M) and audio (300M) encoders for full multimodal…”

What happened

Google has announced the release of EmbeddingGemma 2, an open-source multimodal embedding model designed for on-device processing of text, images, audio, and video. The model is built on the Gemma 4 architecture and is released under the Apache 2.0 license. It has 740 million parameters and is optimized for on-device inference, making it suitable for tasks such as finding specific video clips from voice memos or searching through audio recordings based on text queries. EmbeddingGemma 2 is modular, allowing developers to use only the necessary components for their specific use cases, and it supports dynamic vector truncation for efficient storage. The model is optimized for on-device performance, requiring as little as 191MB of active RAM for text-only weights on a Google Pixel 11 Pro.

What to weigh

  • EmbeddingGemma 2 is optimized for on-device performance.
  • The model supports dynamic vector truncation for efficient storage.

FAQ

What is the size of EmbeddingGemma 2?

EmbeddingGemma 2 has 740 million parameters.

What is the license for EmbeddingGemma 2?

EmbeddingGemma 2 is released under the Apache 2.0 license.

How much RAM does EmbeddingGemma 2 require on a Google Pixel 11 Pro?

With quantization, EmbeddingGemma 2 requires as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.


Source: Google

Keep following technology

Continue from here.

Explore more coverage or subscribe to Fathom Layer updates.

Back to the news desk ↗

Occasional verified technology updates. Unsubscribe anytime. See our privacy policy.