HomeAIGoogle DeepMind Introduces Gemma 4 12B
AI

Google DeepMind Introduces Gemma 4 12B Multimodal Model

Google DeepMind introduced Gemma 4 12B, an open-weights multimodal model designed to run locally on consumer laptops with 16GB of memory.

WHAT YOU NEED TO KNOW
  • Google DeepMind introduced Gemma 4 12B under an open Apache 2.0 license.
  • The model uses an encoder-free architecture that routes audio and visual inputs directly into the language backbone.
  • Gemma 4 12B operates locally on hardware with 16GB of RAM, VRAM, or unified memory.
  • Total downloads for Gemma series models have crossed 150 million.

Google DeepMind introduced Gemma 4 12B on June 3, 2026, launching a mid-sized multimodal model designed to run locally on consumer laptops. The release bridges the gap between the organization's edge-focused E4B model and its larger 26B Mixture of Experts architecture, packaging agentic capabilities inside a smaller memory footprint. Gemma 4 models have now accumulated over 150 million downloads across the developer community.

The model operates on an encoder-free architecture that routes visual and audio inputs directly into the primary language model backbone. Rather than using traditional separate encoders, Google DeepMind replaced the vision encoder with a lightweight embedding module using a single matrix multiplication, positional embedding, and normalizations. Audio signals project directly into the same dimensional space used for text tokens, making Gemma 4 12B the company's first mid-sized model featuring native audio inputs.

Hardware and deployment

Engineered to run locally, Gemma 4 12B requires 16GB of RAM, VRAM, or unified memory. The model comes equipped with Multi-Token Prediction drafters to lower latency during operation and achieves benchmark performance near Google DeepMind's 26B model at less than half of the memory footprint. Google DeepMind released pre-trained and instruction-tuned checkpoints on Hugging Face and Kaggle under an Apache 2.0 license.

Testing options include LM Studio, Ollama, the Google AI Edge Gallery App, the Google AI Edge Eloquent app, and the LiteRT-LM CLI. Local inference pipelines support Hugging Face Transformers, llama.cpp, MLX, SGLang, and vLLM, with Unsloth available for model fine-tuning. Google DeepMind also released an official Skills Repository containing a library of skills designed for agent development.

For production deployments, endpoints can be created using Google Cloud services including the Gemini Enterprise Agent Platform Model Garden, Cloud Run, and Google Kubernetes Engine. Community projects built on the Gemma architecture prior to this release range from wearable robotic arms for physical assistance to enterprise AI security tools.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →