Long-horizon LLM agent training

HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents

HiPER separates high-level planning from low-level execution and trains both levels with Hierarchical Advantage Estimation.

Jiangweizhi Peng Yuanxin Liu Ruida Zhou Charles Fleming Zhaoran Wang Alfredo Garcia Mingyi Hong

Core idea

Train agents at the level where they actually make progress.

Humans do not solve long-horizon tasks by blindly taking one small action after another. We plan, pursue subgoals, decide when to stick with them, and switch when needed.

HiPER makes that hidden structure explicit, then assigns credit to both the plan and the execution that follows it.

Problem

Flat RL blurs the cause of success.

Sparse final rewards make it hard to tell whether an agent chose the wrong subgoal, executed poorly, or switched too early.

Structure

Plan-Execute exposes subgoals.

The agent decides whether to keep the current subgoal or switch, then acts under that subgoal while new observations arrive.

Training

HAE matches credit to the hierarchy.

Hierarchical Advantage Estimation propagates learning signals within subgoal segments and across subgoal boundaries.

Performance Summary

HiPER improves long-horizon agent training across ALFWorld and WebShop, with stronger final success rates and more stable training curves.

HiPER performance summary across ALFWorld and WebShop experiments

Approach

HiPER turns long-horizon interaction into an explicit Plan-Execute loop, then assigns credit separately to decisions made at the planning and execution levels.

Instead of optimizing a flat stream of actions, HiPER exposes the hierarchy already used by capable agents: a high-level policy selects or maintains a subgoal, while a low-level executor takes grounded actions until the next planning boundary.

Plan

Choose the next subgoal

Planning: the agent decides whether to continue the current objective or switch to a better one as observations arrive.

Execute

Execute grounded actions

Execution: the agent converts each subgoal into environment actions, keeping local behavior tied to the current plan.

Credit

Train with hierarchical advantages

Hierarchical Advantage Estimation propagates feedback within each subgoal segment and across segment boundaries.

HiPER hierarchical plan-execute reinforcement learning diagram

Example Trajectories

Appendix E of the paper illustrates how HiPER agents keep subgoals while they remain useful and switch when the task phase changes. Select a case to view the original trajectory grouped by contiguous subgoal segments.

BibTeX

Citation

@article{peng2026hiper,
  title={HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents},
  author={Peng, Jiangweizhi and Liu, Yuanxin and Zhou, Ruida and Fleming, Charles and Wang, Zhaoran and Garcia, Alfredo and Hong, Mingyi},
  journal={arXiv preprint arXiv:2602.16165},
  year={2026}
}