ProCredit

ProCredit

From Outcome Rewards to Progress Credit
in Agentic Reinforcement Learning

Ming Ma1,2, Yi Zhu3,*, Yiran Zhong3,*, Feida Zhu3, Chonghan Liu4, Pengkun Jiao3,
Qichao Wang5, Yanhao Jia5, Tianming Yang1, Steven Hoi3

Affiliations

1Institute of Neuroscience, Chinese Academy of Sciences 2University of Chinese Academy of Sciences 3Tongyi Lab, Alibaba Group 4University of California, Los Angeles 5Nanyang Technological University

*Corresponding authors: Yi Zhu, Yiran Zhong

An attempt can fail.
Its progress can still teach.

Examples

Watch the checks turn progress into credit

An agent works on a phone task turn by turn. After every turn, the task’s acceptance checks run again on the new state. When the attempt ends, compare the credit GRPO and ProCredit give each turn.

01

User request

    Agent ⇄ EnvironmentReady
    Device state
    9:41

    Alarms

      Device state updates after each tool return.

      Acceptance checksΦ = 0/3
        0 / 0 Illustrative replay · credit computed with the paper’s equations (c = 0.5, γ = 1)

        Method

        Progress is as verifiable as the outcome

        The checks that decide whether a task succeeded can also run on every intermediate state. ProCredit uses them three ways.

        1. 1

          Check progress after every turn

          Φt = passed checks / K

          Rerun the task’s K acceptance checks on the new environment state. A query leaves progress unchanged; a wrong change can lower it.

        2. 2

          Reward the change

          rt = c (Φt − Φt−1)

          The attempt’s score adds its progress rewards to the outcome, S = R + c ΦT. A success still outscores every failure.

        3. 3

          Credit at two levels

          Ai,t = Aitraj + Ai,tturn

          Compare scores across attempts at the same task, then tilt credit within an attempt by the return still to be earned from each turn.

        ProCredit verifies progress after each turn and combines trajectory-level and turn-level credit to update the policy.

        Performance

        Results on AppWorld

        Qwen3.5 base models at three sizes, trained on AppWorld. Mean ± standard deviation over 3 training runs.

        TGC: task goal completion. SGC: scenario goal completion, where every task in a scenario must succeed. Axes start above zero so the differences are visible; bars share one scale within each metric. 35B is Qwen3.5-35B-A3B. GiGPO uses its default γ = 0.95.

        Citation

        @article{ma2026procredit,
          title   = {ProCredit: From Outcome Rewards to Progress Credit
                     in Agentic Reinforcement Learning},
          author  = {Ma, Ming and Zhu, Yi and Zhong, Yiran and Zhu, Feida and
                     Liu, Chonghan and Jiao, Pengkun and Wang, Qichao and
                     Jia, Yanhao and Yang, Tianming and Hoi, Steven},
          journal = {arXiv preprint arXiv:2609.27532},
          year    = {2026},
          url     = {https://arxiv.org/abs/2609.27532}
        }