Transformers v5.19.0 Adds Google's EmbeddingGemma 2 Multimodal Embedding Model and Deprecates the paged Attention Prefix for Continuous Batching
Hugging Face Transformers v5.19.0 adds EmbeddingGemma 2, a Gemma 4-based embedding model for text, images, audio and video, and flags breaking changes.
Overview
The Hugging Face Transformers library has shipped v5.19.0, whose headline addition is EmbeddingGemma 2. The release page lists the model as “a multimodal embedding model from Google built on the Gemma 4 architecture” that encodes text, images, audio and video into a shared 768-dimensional vector space. The same release also lists breaking changes for continuous batching, mixture-of-experts (MoE) outputs and one object-detection model. The release page carries a publish timestamp of 2026-10-06T16:39:23Z, and the Transformers documentation says the model “was contributed to Hugging Face Transformers on 2026-10-06.”
What EmbeddingGemma 2 Is
In its developer guide, dated Oct. 6, 2026, Google describes EmbeddingGemma 2 as “a single, compact open model released under the Apache 2.0 license” and a “sub-1B model” that maps text, code, images, video and audio into a unified 768-dimensional space. According to the same guide, the model scales “from 270M parameters for text and code up to 740M parameters for all modalities.” Google lists four configurations: 270M for text and code, 440M for text plus vision, 570M for text plus audio, and 740M for the full multimodal model. The guide says all modalities share an 8,192-token context window.
The Transformers documentation uses the google/embeddinggemma-2 identifier in its examples. It says unused vision or audio towers can be disabled at load time by setting vision_config=None or audio_config=None, which reduces memory footprint for text-only or single-modality deployments. The documentation names Sentence Transformers (version 6.1.0 or later) as the recommended entry point; Google’s guide likewise gives the sentence-transformers library, v6.1.0 or later, as the way to run the model across modalities.
Smaller Vectors for Vector Databases
The release notes say the model uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. Google’s guide quantifies the trade-off, and these figures are Google’s own rather than independently measured:
- At 256 dimensions, Google says, most of the full quality on text and code is retained, and about 95% on image, video and speech retrieval.
- At 128 dimensions, Google says text and code retain around 90% quality, while image, video and speech retrieval quality drop to around 75%. The guide advises users to validate on target data before deploying 128 dimensions for multimodal queries.
- For storage, Google says that in bfloat16 precision a million 768-dimensional vectors take roughly 1.5 GB of memory, against 250 MB when truncated to 128 dimensions.
The Transformers documentation separately recommends evaluating the model both with and without task prompts on a specific task.
Breaking Changes in the Release
The release notes flag several changes that can affect existing code:
- Continuous batching attention: the
"paged|"prefix for SDPA and flash attention implementations is deprecated. Users should set the regular attention implementation, such assdpaorflash_attention_2, instead ofpaged|sdpaorpaged|flash_attention_2. The notes add that eager attention still requires the"paged|eager"prefix. - Cache updates in continuous batching: the cache update for the index-based and block-table paths is now fused into a single call, which “affects any custom code that calls the separate update paths.”
- MoE router logits: all MoE models whose routers compute logits now return them when
output_router_logits=True. - OWLv2:
Owlv2ForObjectDetection.embed_image_querynow selects the query box with the highest objectness score, so results may differ from earlier releases.
The notes also describe a token-dispatch implementation for expert parallelism, now the default for Qwen3 MoE and Mellum, which removes the requirement that expert-parallel size equal tensor-parallel size. They add per-layer cache configuration for DynamicCache and StaticCache, and say generate no longer mutates the user’s cache_config.
What We Don’t Know
The sources reviewed do not include independent benchmarks of EmbeddingGemma 2; the retention figures above are Google’s. The release notes do not say when the "paged|" prefix will be removed, only that it is deprecated and that the continuous-batching internals changed in preparation for removing it.