dArtBook a call
all news
Ars Technica · June 10, 2026

Google’s DiffusionGemma writes entire paragraphs at once—and it’s four times faster

Google’s DiffusionGemma writes entire paragraphs at once—and it’s four times faster

Google DeepMind just dropped a new AI model that works more like an artist than a typist. Called DiffusionGemma, it’s the latest addition to the company’s open-source Gemma family, but it breaks the mold in a big way. Instead of spitting out text one word at a time, left to right, DiffusionGemma generates whole blocks of text in parallel. Think of it like starting with a blurry image and slowly sharpening it into a clear picture—except the picture is a paragraph.

Most AI language models are autoregressive: they predict the next token based on the ones before it. DiffusionGemma flips that script. It begins with a field of placeholder tokens and runs over them multiple times, refining guesses until the entire block snaps into focus. This approach shifts the bottleneck from memory speed to raw computing power, allowing the model to produce up to 256 tokens at once.

At 26 billion total parameters, DiffusionGemma is a big model, but it only activates 3.8 billion during use—meaning it can run on a high-end gaming GPU with 18GB of RAM. On an RTX 5090, it cranks out about 700 tokens per second. On a single Nvidia H100 accelerator, that number jumps to over 1,000 tokens per second. That’s roughly four times the speed of similarly sized autoregressive Gemma models.

Google says this design makes DiffusionGemma especially good at non-linear tasks like inline editing, molecular sequencing, and even solving Sudoku puzzles—something standard models struggle with because each answer depends on later clues. By continuously self-correcting across large sets of tokens, DiffusionGemma handles those challenges with ease.

Source: Ars Technica

Want a self-updating feed like this on your site?

dArt Studio installs AI for local businesses in Broward & Palm Beach County, FL. We reply within 1 business hour.