Credit assignment methods for multi-turn agents

Which feedback and cross-action comparisons produce the policy's credit signal?

JSON · CSV · BibTeX

WorkCredit unitFeedbackComparison mechanismAdditional requirements
GiGPOEpisode and action step4.3 Group-in-Group Policy OptimizationTrajectory returns and discounted future rewards4.1 Episode Relative Advantages · 4.2 Step Relative AdvantagesRelative returns among actions at identical anchor states4.2 Step Relative Advantages · 4.2 Step Relative AdvantagesRepeated states in collected rollouts; offline grouping4.2 Step Relative Advantages
Turn-level RLInteraction turn4.2 Proposed Method: Turn-Level Credit Assignment for Multi-Turn LLM AgentsTurn-level verifiable rewards and final-answer rewards3.2 Reward Design · 3.2 Reward DesignTurn-level MDP advantage estimation using intermediate and outcome rewards4.2 Proposed Method: Turn-Level Credit Assignment for Multi-Turn LLM AgentsTask-specific tool-execution and answer-presence verifiers3.2 Reward Design
HGPOAction step4.2 Hierarchy-of-Groups Policy Optimization · 4.2 Hierarchy-of-Groups Policy OptimizationSparse terminal rewards and discounted returns3 Preliminaries · 4.2 Hierarchy-of-Groups Policy OptimizationNested groups matching the current state and progressively longer histories4.2 Hierarchy-of-Groups Policy Optimization · 4.2 Hierarchy-of-Groups Policy OptimizationState and historical-context matching within existing rollouts4.2 Hierarchy-of-Groups Policy Optimization
RTMCAction segment, shared across its tokens3.1 From Rollouts to TreesDiscounted returns from grouped rollouts3.1 From Rollouts to TreesMonte Carlo action and state values in a rollout tree3 Rollout-Tree Monte Carlo Advantage Estimation · 3.1 From Rollouts to TreesTask-specific state-action signatures; reuse of collected rollouts4 State-Action Signature Design · 4.1 Action Signature · 6.1 Advantages of the Approach
A3Shell-action turn3.3 Action Advantage AssignmentEpisode return3.3 Action Advantage AssignmentEpisode, turn and action-history tree comparisons using shell syntax3.1 AST measure for CLI agent actions · 3.3 Action Advantage AssignmentBash parsing and structural action comparisons3.1 AST measure for CLI agent actions
TRACETool-interaction turn; credit shared across its assistant tokens3.1 Tool-Boundary States · Joint optimization with outcome reward.Final-answer verifier and prefix-conditioned gold-answer likelihoods3.1 Tool-Boundary States · Reference answer score.Within-trajectory TD differences of log-ratio answer-likelihood values, with K-step propagation and outcome anchoringLog-ratio state value. · One-step TD action credit. · $K$ -step turn credit. · Joint optimization with outcome reward.Gold answers and a frozen copy of the initial policy; evaluated on long-horizon search with compact answersReference answer score. · 6 Limitations
TCPOResponse turn, shared across its tokensTurn-Level Advantage and Token OptimizationVerifier score after every turnMulti-Turn Verifier SettingRetrospective progress, future-best quality and fixed-history alternatives3 MethodSelected alternative responses and extra verifier calls; state restoration for AppWorldCounterfactual Estimation and Budgeted Refinement
IAPOTrainable assistant action, shared across its tokensMethodTrajectory advantage with observed support and error dependenciesMethod · Influence-Dependency GraphPositive, bounded routing weights over a within-rollout dependency graphMethodFrozen dependency annotator; experiments use Qwen3-32BRouted advantage. · Experimental Setup

Qualitative comparisons of methods, information requirements and task design in the cited paper versions.

References

  1. Group-in-Group Policy Optimization for LLM Agent Training. Feng, Lang; Xue, Zhenghai; Liu, Tingcong; An, Bo. 2505.10978v1.
  2. Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Credit Assignment. Zeng, Siliang; Wei, Quan; Brown, William; Frunza, Oana; Nevmyvaka, Yuriy; Hong, Mingyi. 2505.11821v1.
  3. Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks. He, Shuo; Feng, Lang; Wei, Qi; Cheng, Xin; Feng, Lei; An, Bo. 2602.22817v1.
  4. RTMC: Step-Level Credit Assignment via Rollout Trees. Wang, Tao; Zheng, Suhang; Xu, Xiaoxiao. 2604.11037v1.
  5. Learning CLI Agents with Structured Action Credit under Selective Observation. Su, Haoyang; Wen, Ying. 2605.08013v1.
  6. TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents. Tao, Leitian; Peng, Baolin; Yao, Wenlin; Ge, Tao; Cheng, Hao; Wang, Mike Hang; Gao, Jianfeng; Li, Sharon. 2607.13988v1.
  7. TCPO: Turn-Level Credit Policy Optimization. Liao, Sicong; Chen, Zhi; Tang, Yaohua. 2608.01667v1.
  8. IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents. Ren, Bo; Mao, Yirong; Yang, Yi; Que, Wenhui. 2608.24588v2.