Appearance
The verification bottleneck
Reinforcement learning with verifiable rewards (RLVR) is why reasoning models got so good at math and code. The environment grades the output, the reward is trustworthy, and the optimization loop stays clean. Agentic tasks break that loop immediately. What is the reward for a good research plan, a well-negotiated contract, or the right tool call at step three of a five-step synthesis? There's no executable checker and no ground truth hiding in a test file. Human labels can't scale to thousand-step trajectories, and one scalar at the end won't tell you which action caused the outcome.
This cluster of eight papers attacks that bottleneck from different angles. PaperGym builds training environments out of research papers, using the paper's structure as a critic that paraphrase can't game. TASPO shows that fine-grained supervision isn't the same thing as fine-grained credit. A paper on on-policy distillation argues the teacher was mostly unnecessary all along. Two new benchmarks, S3Gym and E-Commerce Bench, stop evaluating fixed policies and start measuring whether agents can improve at all.
The through-line is counterintuitive enough to state up front: in this crop of results, removing supervision beats adding it more often than you'd expect.
Rubrics as critics: PaperGym turns a paper into an environment
Research planning is where the agentic reward problem gets nastiest. A plan has no verifiable answer, so RL lacks the environment it requires: tasks paired with a critic. PaperGym's fix is to extract both from the structure of a scientific paper.
The key move is separating sources. Existing rubric pipelines draw the question and the grading criteria from the same content, so a model can earn reward by paraphrasing. PaperGym synthesizes the question from the paper's research goal and background, and derives the criteria from the method and experiments sections. Separate sources mean paraphrase stops paying. Criterion leakage drops to 3.7%, where existing datasets sit at 11.90% to 34.10%. That's the difference between a critic an agent can cheat and one it has to satisfy on the merits.
Training runs the rubric twice. First, OPSD's self-teacher uses it as privileged context to generate high-quality trajectories. Then GRPO uses it as the reward signal. Across Qwen3-1.7B, 4B, and 8B, that schedule beats supervised fine-tuning, either stage alone, and the reverse ordering, improving five-benchmark averages by +5.6, +5.0, and +4.8 points. The 8B model reaches 73.48 on ResearchQA, above the far larger Kimi K2.6, and wins 58.1% of three-way comparisons against 28.2% for RubricHub Science.
The small-model results matter more than the leaderboard position. Qwen3-1.7B runs fine on a laptop GPU. Gains of that size coming from environment construction rather than model scale tell you the critic design is the product. The pipeline, the 20,000-instance PaperGym-20k corpus, and the PaperGym-Innov and PaperGym-Design benchmarks are all released, so you can verify that claim on your own hardware.
Does on-policy distillation actually distill?
On-policy distillation (OPD) was supposed to be the alternative to RLVR's coarse rewards. A teacher model provides dense token-level supervision over the student's own rollouts. The setup has an obvious flaw that got mostly ignored: those rollouts are off-policy for the teacher, so the grades are systematically off.
The paper quantifies the noise and finds it substantial, and growing with teacher scale. Bigger teacher, noisier supervision. Then comes the uncomfortable part: the student policy is insensitive to the noise. Removing it entirely converges to the same performance. Whatever OPD is doing, it isn't passing the teacher's signal through.
So where do the gains come from? They concentrate on tokens the student assigned low log-probability. A single fixed negative advantage on those tokens matches the performance of teacher-provided advantages. Translation: OPD mostly suppresses tail tokens the student already disliked. The mechanism is likelihood correction. Knowledge transfer plays little role, and the whole setup needs no teacher at all.
The follow-through, On-Policy Self-Adaptation (OPSA), acts on that finding. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens while redistributing probability mass among head tokens. No teacher, no privileged information, no distilled signal. On a Qwen3-1.7B base, OPSA improves Avg@32 on AIME24 by 35.41 points, a 263% relative gain, and more than doubles Pass@32 across all three benchmarks. It beats OPD by 16.77 points on the same AIME24 metric.
This matches a suspicion I've carried through my own distillation runs. When the teacher gets bigger, the supervision gets noisier, and the student never changes its behavior accordingly. Now I know why: the dense scores were mostly the student's own likelihoods dressed up as external knowledge.
Quick take: Across this cluster, the practical wins come from converting supervision into action-level credit or removing the supervision entirely, not from stacking more elaborate teachers.
Fine-grained supervision isn't fine-grained credit
If OPD shows dense token supervision is mostly noise, TASPO shows that even meaningful supervision isn't credit. The distinction is the paper's core contribution.
On-policy self-distillation re-evaluates sampled behavior with privileged information (PI) that exists only at training time. The assumption was that PI gives finer supervision than an outcome scalar. TASPO identifies the gap: PI-induced likelihood shifts describe how extra information changes policy preference. They don't determine how an executable action should inherit the verified outcome. A privileged signal can be irrelevant to the current state, operate at token granularity that doesn't align with tool calls or menu selections, and carry no outcome semantics at all.
TASPO's fix is disciplined. It builds decision-applicable PI from verified successful experience, aggregates PI-induced likelihood shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the original trajectory advantage. The verified outcome sets the update direction and average scale. PI only redistributes credit across actions inside that trajectory.
The empirical payoff is +10.6% over GRPO across three agentic benchmarks, with better generalization to unseen tasks, and the analysis shows action-level assignment stabilizes the optimization. I've watched this failure mode in my own GRPO runs: the trajectory advantage applies the same pressure to every decision, so a correct tool call at step two gets treated like a hallucinated quantity at step five. TASPO breaks that lumpiness.
Fine-grained supervision and fine-grained credit are different things. Conflating them is how agent RL wastes signal.
Five levels of human control
One paper in the cluster steps back and frames where this is all heading. Its structure is a five-level ladder, L0 to L4, describing how much of the learning loop remains under human control.
L0 is per-instance human judgments. L1 is verifiable rewards: programmatic critics that need no human. L2 is learned critics: rubrics, privileged supervision, teacher models. L3 is self-generated curricula and constructed environments, where agents build their own tasks. L4 is autonomous co-evolution, where reward generation and experience generation are themselves learned and shared.
Two axes move together. The reward axis runs from human judgments to reusable verifiers to autonomous rewards. The experience axis runs from human-curated tasks to self-generated curricula to constructed environments. The paper also names the failure modes you hit climbing the ladder: reward hacking, feedback drift, curriculum collapse, environment errors. It proposes evaluating three things separately, policy capability, feedback fidelity, and experience quality, and the authors keep a continuously updated repository of work on this ladder.
You can map every method in this cluster onto that ladder.
| Paper | Reward source | Credit granularity | Human in the loop | Headline result |
|---|---|---|---|---|
| PaperGym | Rubric from paper method/experiments | Trajectory (GRPO) | No labels after corpus build | +5.6 avg over SFT; 8B out-scores Kimi K2.6 on ResearchQA |
| OPSA | None | Token-level via entropy | None | +35.41 AIME24 Avg@32 over base; +16.77 over OPD |
| TASPO | Verified outcome + privileged info | Action-level weights | Verified trajectories only | +10.6% over GRPO; generalizes to unseen tasks |
| Single-policy chem agent | Programmatic reward off gold call chain | Outcome-level | None | +5.5% Tool F1 over MCTS at one inference per question |
| PRACTICE | Oracle trajectories + success/failure contrast | Skill-library edits | Oracle trajectories for curriculum | Outperforms experience-based baselines on EB-ALFRED and EB-Habitat |
Read down the "Human in the loop" column. Not one row needs preference labels at scale. Only PRACTICE depends on oracle trajectories, and only for its initial curriculum, not as a continuous dependency. That's the direction of travel the ladder paper describes.
Benchmarks that measure improvement, not just performance
Agent benchmarks mostly test frozen policies: set the weights, run the eval, record the score. S3Gym refuses that premise. It evaluates three coupled capabilities, self-testing, self-judging, and self-improvement, across seven text-based games with executable environment verifiers. The protocol keeps permissive exploration separate from strict held-out evaluation, so an agent can experiment freely and then gets judged on the real task.
The findings are more sobering than flashy. Self-improvement is neither automatic nor uniform. The best pathway for incorporating experience depends on task structure. Summary memory wins when experience compresses into reusable strategic rules, but underperforms raw history when success depends on precise, state-contingent information. Parameter training produces large gains on some tasks, then shows unstable improvement and severe negative transfer on others. This matches my eval habits: I reach for compressed rules when a game has stable strategy, and raw transcripts when every observation matters. The paper's conclusion is the thesis of this whole cluster: recognizing successful actions is insufficient. Agents also have to transform feedback into executable, transferable policies.
E-Commerce Bench pushes the horizon to its extreme: 365 days of concurrent store operation. An agent runs multiple online stores, researches the market, negotiates with suppliers, manages inventory and cash flow, and handles orders and returns, all toward maximizing year-end assets. Both sides of the market are deterministic, so runs reproduce, and a negotiation kernel, not an LLM, sets supplier behavior. The authors ran 18 frontier models across seven evaluation dimensions.
No model dominates.
GPT-5.6 Sol turns the 100,000 opening stake into 1,431,425, a 14.3x return. It also ranks 16th of 18 on fraud avoidance and trails Fable5 on operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads with 416,252, 38% above GLM 5.2 (high), and shows the strongest learning over the horizon, bargaining prices down across repeated orders.
The lesson isn't which model wins. A 365-day horizon surfaces strengths and weaknesses a 30-turn benchmark never will, and it does all of it without a human scoring a single action.
Key numbers from the cluster
- 263%: relative Avg@32 gain on AIME24 from OPSA over the Qwen3-1.7B base, which is +35.41 points.
- 16.77: OPSA's margin over OPD on AIME24 Avg@32, with no teacher in the loop.
- 3.7%: PaperGym's criterion leakage rate, versus 11.90% to 34.10% in existing datasets.
- 10.6%: TASPO's average improvement over GRPO across three agentic benchmarks.
- 14.3x: GPT-5.6 Sol's return on the 100,000 opening stake in E-Commerce Bench.
Practice: let the skill library edit itself
PRACTICE comes from the embodied-agent side and is the closest thing here to an agent that explicitly accumulates experience. Past experience-based methods extract skills from trajectories using hand-crafted prompting workflows. Those fixed procedures break when new experiences stop fitting the template. PRACTICE trains a skill learner to do the extraction instead, while keeping the task executor frozen.
The skill learner maintains a persistent skill library and produces structured batch edits: add, refine, merge, or remove skills, then hierarchically consolidates everything into a consistent updated library. Training uses a two-stage curriculum. First it learns basic skill generation and library maintenance from oracle trajectories. Then it contrasts successful and failed trajectories from heterogeneous executors on the same tasks, so it can spot invalid action patterns and learning recovery strategies. An online skill-edit distillation step aligns the learner with a stronger teacher on its current edit distribution, improving the policy further.
The result: a compact skill learner delivers consistent performance gains across successive library-update rounds for multiple frozen executors, and beats the strongest experience-based baselines on EB-ALFRED and EB-Habitat.
Note the irony. PRACTICE leans on a stronger teacher at the end, in the same cluster where OPSA argues the teacher was never the source of the gains. The two aren't in direct conflict: PRACTICE's teacher guides library edits, not token probabilities. The contrast is exactly the open question this cluster keeps circling. The field hasn't settled whether teachers are scaffolding or crutches.
One policy beats a whole search tree
The chemistry tool-use paper makes the same point from a different direction. CheMatAgent treated tool use as search: hierarchical evolutionary MCTS over tool-call trees, with separate policy and execution models and two learned critics, one regressed partly onto GPT-assigned scores. The new work shows a single policy is enough.
That policy interleaves reasoning, tool calls, and tool returns in one left-to-right generation. Training is a supervised warm-up followed by outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain. No learned critic, no judge, no GPT scores anywhere in the loop.
| Dimension | Single policy (this work) | Hierarchical MCTS (CheMatAgent) |
|---|---|---|
| Inferences per question | 1 | Grows with the tree |
| Learned critics | None | Two, one partly regressed on GPT scores |
| Reward | Programmatic, off gold call chain | Learned critic scores |
| Tool F1 on Qwen-2.5-7B | +5.5% over MCTS | Baseline |
| Return F1 on Qwen-2.5-7B | +9.6% over MCTS | Baseline |
On ChemToolBench's multiple-tool comprehensive setting, the single policy beats CheMatAgent's strongest search configuration by 5.5% Tool F1 and 9.6% Return F1 on Qwen-2.5-7B, and by 3.7% and 3.9% on Llama-3.1-8B, at one model invocation per question. The search baseline pays growing cost per question. The single policy doesn't branch at all.
The pattern is consistent: whenever a programmatic or structural signal exists, a learned critic is a liability, not a feature. This cluster produces evidence for that view from chemistry tools, research plans, and embodied navigation.
Common pitfalls
Treat dense supervision as credit assignment. You bolt a teacher onto the loop and read token scores as fine-grained credit. OPD's analysis shows those scores are mostly noise, and the gain comes from suppressing low-probability tokens, which a fixed negative advantage gets you without the teacher.
Draw the question and the criteria from the same content. That's the paraphrase trap PaperGym measures: leakage of 11.90% to 34.10% in existing datasets because the model can earn reward by restating the prompt. Derive them from separate parts of the structure.
Assume self-improvement is automatic. S3Gym's results show the winning pathway depends on task structure: summaries for compressible rules, raw history for state-contingent precision, and parameter training that can produce severe negative transfer. Measure each pathway, don't assume the loop works.
Add a learned critic when a programmatic reward exists. CheMatAgent runs two critics, one regressed partly onto GPT scores, and loses to a single policy trained directly on the gold call chain. Check whether the environment can give