NVIDIA now supports Google DeepMind's DiffusionGemma model across its RTX and DGX hardware platforms. This update brings day-one compatibility to NVIDIA systems without requiring custom integration work from users.
Google DeepMind open model runs natively on consumer and enterprise hardware without custom integration work
The software framework enables the open AI model to run on consumer graphics cards and enterprise data center accelerators. The architecture relies on Gemma 4 as a base with a specialized diffusion approach for token generation.

DiffusionGemma produces 256 tokens in parallel during each processing step. NVIDIA reports throughput of over 150 tokens per second on DGX Spark systems and more than 1,000 tokens per second on single H100 accelerators.
The model uses only 3.8 billion active parameters for every generation cycle despite the larger underlying structure. NVIDIA distributes BF16 and NVFP4 checkpoint files through standard repositories like Hugging Face Transformers, vLLM, and Unsloth.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.