# GDPO: Why GRPO Breaks Under Multiple Rewards - and How to Fix It

> NVIDIA researchers show that GRPO's reward normalization collapses distinct advantage signals when multiple rewards are used together, causing training instability. GDPO decouples normalization per reward, boosting AIME accuracy from 23.1% to 29.4% and eliminating training collapse.

Canonical URL: https://kravhal.kcsatish.com/insights/week-15
Edition: Week 08 · April 2026
Tags: Deep Learning, LLMs, RLHF
Reading time: 7 min read

---

This is a mirror of an article first published in the AI & Automation Chronicle.

Full text with the original formatting: https://chronicle.kcsatish.com/posts/week-15
Markdown of the original: https://chronicle.kcsatish.com/posts/week-15.md
Structured JSON of the original: https://chronicle.kcsatish.com/api/v1/posts/week-15.json

Cite the Chronicle as the publication of record for the research claims in this article.
