about
Gemma 4 Technical Report (arxiv.org)
2 points by ferretj 88 days ago | hide | past | pdf | discuss on HN

In plain words: Gemma 4 is a family of free-to-use AI models that handle text, images, and sound, and can think through a problem step by step before answering. Across science, multimodal, and long-context tests, it beats earlier versions and matches much bigger open models on human-rated tasks.

Abstract

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and image patches. Furthermore, we integrate a thinking mode, enabling Gemma models to generate reasoning traces prior to responding. We improve inference speed, memory, and compute efficiency, as well as long-context abilities through critical design choices. Gemma 4 establishes a leap in performance across STEM, multimodal, and long-context benchmarks, and rivals larger, frontier open models in human-rated tasks.

Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Aditya Chawla, et al.
arXiv:2607.02770 · cs.CL, cs.AI · submitted Jul 2, 2026 · updated Jul 24, 2026
abstract · pdf · html · 17 pages, 2 figures, technical report, updated

add comment on HN
Also discussed: Jul 2026 (1 point, 0 comments) · Jul 2026 (3 points, 0 comments)