Credit assignment methods for multi-turn agents
Which feedback and cross-action comparisons produce the policy's credit signal?
| Work | Credit unit | Feedback | Comparison mechanism | Additional requirements |
|---|---|---|---|---|
| GiGPO | Episode and action step4.3 Group-in-Group Policy Optimization | Trajectory returns and discounted future rewards4.1 Episode Relative Advantages · 4.2 Step Relative Advantages | Relative returns among actions at identical anchor states4.2 Step Relative Advantages · 4.2 Step Relative Advantages | Repeated states in collected rollouts; offline grouping4.2 Step Relative Advantages |
| Turn-level RL | Interaction turn4.2 Proposed Method: Turn-Level Credit Assignment for Multi-Turn LLM Agents | Turn-level verifiable rewards and final-answer rewards3.2 Reward Design · 3.2 Reward Design | Turn-level MDP advantage estimation using intermediate and outcome rewards4.2 Proposed Method: Turn-Level Credit Assignment for Multi-Turn LLM Agents | Task-specific tool-execution and answer-presence verifiers3.2 Reward Design |
| HGPO | Action step4.2 Hierarchy-of-Groups Policy Optimization · 4.2 Hierarchy-of-Groups Policy Optimization | Sparse terminal rewards and discounted returns3 Preliminaries · 4.2 Hierarchy-of-Groups Policy Optimization | Nested groups matching the current state and progressively longer histories4.2 Hierarchy-of-Groups Policy Optimization · 4.2 Hierarchy-of-Groups Policy Optimization | State and historical-context matching within existing rollouts4.2 Hierarchy-of-Groups Policy Optimization |
| RTMC | Action segment, shared across its tokens3.1 From Rollouts to Trees | Discounted returns from grouped rollouts3.1 From Rollouts to Trees | Monte Carlo action and state values in a rollout tree3 Rollout-Tree Monte Carlo Advantage Estimation · 3.1 From Rollouts to Trees | Task-specific state-action signatures; reuse of collected rollouts4 State-Action Signature Design · 4.1 Action Signature · 6.1 Advantages of the Approach |
| A3 | Shell-action turn3.3 Action Advantage Assignment | Episode return3.3 Action Advantage Assignment | Episode, turn and action-history tree comparisons using shell syntax3.1 AST measure for CLI agent actions · 3.3 Action Advantage Assignment | Bash parsing and structural action comparisons3.1 AST measure for CLI agent actions |
| TRACE | Tool-interaction turn; credit shared across its assistant tokens3.1 Tool-Boundary States · Joint optimization with outcome reward. | Final-answer verifier and prefix-conditioned gold-answer likelihoods3.1 Tool-Boundary States · Reference answer score. | Within-trajectory TD differences of log-ratio answer-likelihood values, with K-step propagation and outcome anchoringLog-ratio state value. · One-step TD action credit. · $K$ -step turn credit. · Joint optimization with outcome reward. | Gold answers and a frozen copy of the initial policy; evaluated on long-horizon search with compact answersReference answer score. · 6 Limitations |
| TCPO | Response turn, shared across its tokensTurn-Level Advantage and Token Optimization | Verifier score after every turnMulti-Turn Verifier Setting | Retrospective progress, future-best quality and fixed-history alternatives3 Method | Selected alternative responses and extra verifier calls; state restoration for AppWorldCounterfactual Estimation and Budgeted Refinement |
| IAPO | Trainable assistant action, shared across its tokensMethod | Trajectory advantage with observed support and error dependenciesMethod · Influence-Dependency Graph | Positive, bounded routing weights over a within-rollout dependency graphMethod | Frozen dependency annotator; experiments use Qwen3-32BRouted advantage. · Experimental Setup |
Qualitative comparisons of methods, information requirements and task design in the cited paper versions.
References
- Group-in-Group Policy Optimization for LLM Agent Training. Feng, Lang; Xue, Zhenghai; Liu, Tingcong; An, Bo. 2505.10978v1.
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Credit Assignment. Zeng, Siliang; Wei, Quan; Brown, William; Frunza, Oana; Nevmyvaka, Yuriy; Hong, Mingyi. 2505.11821v1.
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks. He, Shuo; Feng, Lang; Wei, Qi; Cheng, Xin; Feng, Lei; An, Bo. 2602.22817v1.
- RTMC: Step-Level Credit Assignment via Rollout Trees. Wang, Tao; Zheng, Suhang; Xu, Xiaoxiao. 2604.11037v1.
- Learning CLI Agents with Structured Action Credit under Selective Observation. Su, Haoyang; Wen, Ying. 2605.08013v1.
- TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents. Tao, Leitian; Peng, Baolin; Yao, Wenlin; Ge, Tao; Cheng, Hao; Wang, Mike Hang; Gao, Jianfeng; Li, Sharon. 2607.13988v1.
- TCPO: Turn-Level Credit Policy Optimization. Liao, Sicong; Chen, Zhi; Tang, Yaohua. 2608.01667v1.
- IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents. Ren, Bo; Mao, Yirong; Yang, Yi; Que, Wenhui. 2608.24588v2.