CorrGRPO: Correlation-Normalized GRPO
for Multi-Reward Learning

Wenbin Hu*Huihao Jing*Haochen ShiYuxuan LiuHaoran LiYangqiu Song

Hong Kong University of Science and Technology

* Equal contribution

GRPO and CorrGRPO advantage equations, showing how covariance normalization becomes correlation normalization, with notation definitions.
Validation accuracy for GRPO and CorrGRPO during training. Validation efficiency for GRPO and CorrGRPO during training.
CorrGRPO overview. CorrGRPO replaces total-reward standard deviation normalization with a correlation-based denominator while retaining the centered total reward. Validation accuracy and efficiency are measured by Mean@1 for Qwen2.5-Coder-7B-Instruct on LeetCodeDataset.

Introduction

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards.

We propose Correlation-Normalized GRPO (CorrGRPO), which normalizes pairwise covariances into Pearson correlation coefficients. CorrGRPO keeps the centered total reward unchanged while balancing the influence of differently scaled rewards on the correlation-based normalization. This allows advantage magnitudes to adapt to reward correlations without the normalization being dominated by large-scale reward components. We compare CorrGRPO with GRPO and other variants on code generation, tool calling, and agent security, using models ranging from 0.5B to 8B parameters. These tasks all involve multiple rewards that can improve together or present tradeoffs. Results show improvements across three domains, including code generation, tool calling, and agent security.

Method

GRPO normalizes by the standard deviation of the total reward, whose variance is the sum of pairwise covariances. CorrGRPO retains the same centered total reward and normalizes each covariance into a Pearson correlation.

GRPO

AGRPOi=Ri−RˉVar⁡^(R)+ε=∑l=1r(Rli−Rˉl)Var⁡^ ⁣(∑l=1rRl)+ε=∑l=1r(Rli−Rˉl)∑l=1r∑m=1rCov⁡^(Rl,Rm)+εA_{\mathrm{GRPO}}^i=\frac{R^i-\bar R}{\sqrt{\widehat{\operatorname{Var}}(R)}+\varepsilon}=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\widehat{\operatorname{Var}}\!\left(\sum_{l=1}^{r}R_l\right)}+\varepsilon}=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\sum_{l=1}^{r}\sum_{m=1}^{r}\widehat{\operatorname{Cov}}(R_l,R_m)}+\varepsilon}
AGRPOi=Ri−RˉVar⁡^(R)+ε=∑l=1r(Rli−Rˉl)Var⁡^ ⁣(∑l=1rRl)+ε=∑l=1r(Rli−Rˉl)∑l=1r∑m=1rCov⁡^(Rl,Rm)+ε\begin{gathered}A_{\mathrm{GRPO}}^i=\frac{R^i-\bar R}{\sqrt{\widehat{\operatorname{Var}}(R)}+\varepsilon}\\[1.1em]=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\widehat{\operatorname{Var}}\!\left(\sum_{l=1}^{r}R_l\right)}+\varepsilon}\\[1.1em]=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\sum_{l=1}^{r}\sum_{m=1}^{r}\widehat{\operatorname{Cov}}(R_l,R_m)}+\varepsilon}\end{gathered}

CorrGRPO

ACorrGRPOi=∑l=1r(Rli−Rˉl)∑l=1r∑m=1rρ^lm+ε=∑l=1r(Rli−Rˉl)∑l=1r∑m=1rCov⁡^(Rl,Rm)Var⁡^(Rl) Var⁡^(Rm)+εA_{\mathrm{CorrGRPO}}^i=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\sum_{l=1}^{r}\sum_{m=1}^{r}\hat\rho_{lm}}+\varepsilon}=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\sum_{l=1}^{r}\sum_{m=1}^{r}\frac{\widehat{\operatorname{Cov}}(R_l,R_m)}{\textcolor{#b4433c}{\sqrt{\widehat{\operatorname{Var}}(R_l)\,\widehat{\operatorname{Var}}(R_m)}}}}+\varepsilon}
ACorrGRPOi=∑l=1r(Rli−Rˉl)∑l=1r∑m=1rρ^lm+ε=∑l=1r(Rli−Rˉl)∑l=1r∑m=1rCov⁡^(Rl,Rm)Var⁡^(Rl) Var⁡^(Rm)+ε\begin{gathered}A_{\mathrm{CorrGRPO}}^i=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\sum_{l=1}^{r}\sum_{m=1}^{r}\hat\rho_{lm}}+\varepsilon}\\[1.1em]=\frac{\sum_{l=1}^{r}(R_l^i-\bar R_l)}{\sqrt{\sum_{l=1}^{r}\sum_{m=1}^{r}\frac{\widehat{\operatorname{Cov}}(R_l,R_m)}{\textcolor{#b4433c}{\sqrt{\widehat{\operatorname{Var}}(R_l)\,\widehat{\operatorname{Var}}(R_m)}}}}+\varepsilon}\end{gathered}

Here, Ri=∑l=1rRliR^i=\sum_{l=1}^{r}R_l^i is the total reward of rollout ii, Rˉl\bar R_l is the group mean of reward component ll, and rr is the number of reward components. All statistics are computed within the same prompt group; ε>0\varepsilon>0 is a numerical stability constant.

Sample covariance and variance measure reward dependence and scale within the group; Pearson correlation divides each covariance by the product of the two component standard deviations. The highlighted term removes reward-scale weighting from each pairwise covariance. Zero-variance reward components contribute zero rows and columns to the correlation matrix.

Four-trajectory example comparing GRPO covariance normalization with CorrGRPO correlation normalization.
Normalization example. The third reward’s variance accounts for 91.1% of the covariance sum, despite its weak correlations with the other rewards. CorrGRPO gives greater influence to the strong correlation between the first two rewards.

Interactive normalization example

Rescale the third reward while keeping the first two rewards fixed. Compare how covariance and correlation affect the normalization denominator.

0.25×3.00×

GRPO · Covariance

Advantage multiplier

CorrGRPO · Correlation

Advantage multiplier

Each multiplier is the reciprocal of the square root of the sum of all matrix entries, with ε omitted. Positive rescaling leaves correlations unchanged, but also changes the centered total reward; the full advantage is not scale invariant.

Results

We evaluate CorrGRPO on code generation, tool calling, and agent security. Select a domain and backbone to compare performance across methods.

Correctness and efficiency Pareto frontier on LeetCodeDataset, comparing GRPO, GDPO and CorrGRPO.
Correctness–efficiency tradeoff on LeetCodeDataset. Efficiency is the mean percentage of eligible programs that run faster than the reference code.
Validation reward curves
Validation mean reward curves comparing GRPO, GDPO and CorrGRPO.
Validation mean@1 scores with exponential moving average smoothing (decay = 0.6).

BibTeX

@misc{hu_corrgrpo,
  title = {CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning},
  author = {Hu, Wenbin and Jing, Huihao and Shi, Haochen and
            Liu, Yuxuan and Li, Haoran and Song, Yangqiu},
  year = {2026},
  eprint = {2609.36820},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url = {https://arxiv.org/abs/2609.36820}
}