Appearance
The problem with agent RL
RL post-training did something remarkable for math and code. Take a base model, run reinforcement learning with verifiable rewards (RLVR), and reasoning quality jumps. The recipe is simple: generate, check, reinforce. GRPO-style group-relative optimization became the default after the R1 wave.
Agents break that recipe. The reward is sparse, arriving once at the end of a long trajectory. The trajectory itself is full of tool calls, user clarifications, and observations that unfold over time. Which action actually resolved the task? The final reward doesn't say. And when the episode ends, everything resets. The agent starts the next episode with no memory of what worked.
Seven new papers on arXiv attack these problems from different angles. Two are about credit assignment. One is about making the learning signal denser. Two are about getting agents to accumulate knowledge across episodes. One is about making the optimization itself cheaper. And one, the outlier, is about teaching agents to write their own tools.
The through-line: everyone is trying to make the RL signal denser, more localized, and more trustworthy. The days of "one reward at the end, hope for the best" are ending.
| Paper | Problem it attacks | Core mechanism | Headline result |
|---|---|---|---|
| IAPO | Credit assignment with outcome-only rewards | Typed influence-dependency graph over actions; redistributes trajectory advantage | Beats multi-turn RL baselines on τ²-Bench, UserBench, AgentChangeBench |
| FARCA | Noisy factual supervision | Reliability-weighted token-level signals with counterfactual attribution | Better factuality without losing general reasoning |
| OPDVR | Sparse RLVR, teacher-bounded distillation | ReLU-gated reward alignment, zero new hyperparameters | Beats standard OPD on six reasoning benchmarks |
| Recuris | Long-horizon memory and skill misalignment | Working Memory + Experiential Memory + Meta-Agent update loop | 35 of 37 model-benchmark pairs improved |
| SkillForge | Append-only skill banks go stale | Evidence-based verification, multi-pathway induction | Beats SkillRL on ALFWorld, WebShop, AppWorld |
| SPO++ | Group-relative RL waits for siblings | Action-token-measure advantage standardization | Better online efficiency than SPO on ALFWorld and Math-TIR |
| SMITH | Tool creation decoupled from tool use | Joint RL over build and use tasks, three reward axes | 79.8% held-out accuracy with a 4B model |
Credit assignment: where the blame goes
IAPO starts from a simple observation. A completed rollout already records how information and errors flow between agent actions. You don't need to resample continuations or build separate step-level signals. You need to read the structure that's already there.
IAPO represents each rollout as a typed influence-dependency graph over trainable agent actions, with user and tool observations as evidence. Support-use and failed-use structure become routing weights that redistribute the same trajectory-level advantage. No new reward, no extra signals. Just a better way to spend the one advantage you have.
The results hold across Qwen3-4B and Qwen3-8B on three service-agent benchmarks: τ²-Bench, UserBench, and AgentChangeBench. Those model sizes mean you can reproduce the work on a single node. BFCL-v4 Multi-Turn shows the gains don't come at the cost of multi-turn function calling.
FARCA attacks the other side of the same coin. When you do have process-level factual supervision, it's noisy. Coarse-grained aggregation of factual signals, with no reliability assessment, creates a mismatch between fact verification and policy updates. FARCA decomposes this into two ambiguities: credit localization (which token caused the error?) and credit reliability (is this factual judgment trustworthy?).
The fix is to align the granularity of fact verification with the granularity of policy updates, then weight signals by reliability. FARCA uses counterfactual evidence attribution, the dependence of a factual judgment on key evidence, as a proxy for verification reliability. Unreliable signals get downweighted before they touch the policy.
IAPO and FARCA are two halves of the same coin. IAPO assumes the outcome reward is all you have and redistributes it intelligently. FARCA assumes you have fine-grained factual signals but they're unreliable, and weights them accordingly.
Quick Take: The agent RL papers landing right now aren't about bigger models or cleverer prompts; they're about fixing the learning signal itself, making it denser, more localized, and verified before it reaches the policy.
Denser signals: OPDVR
On-policy distillation (OPD) gives you dense token-level guidance, but it's bounded by the teacher. RLVR gives you task-level correctness, but the feedback is sparse. Combine them and you get the worst of both tuning nightmares: weighted losses and heuristic switching, each with its own hyperparameters.
OPDVR removes the knobs. It reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism. Correct trajectories get non-negative rewards. Incorrect ones get non-positive rewards. The distillation signal aligns with task success while preserving the teacher's distributional guidance.
The practical payoff: sampled-token OPD becomes a proper RLVR method, so it composes with any policy gradient algorithm, including GRPO. If you've been hand-tuning a lambda between a distillation loss and an RLVR loss, this eliminates that entire class of fiddling. Across six reasoning benchmarks, OPDVR consistently beats standard OPD. Code is on GitHub if you want to poke at the gating yourself.
Memory and skills: breaking the episodic reset
The cluster with the biggest claims is the one about memory. Most RL-trained agents are episodic: they learn within a rollout and forget everything at the episode boundary. Two papers take direct aim at this.
Recuris targets long-horizon tasks where growing histories obscure the task state and misalign skill invocation. The architecture couples Working Memory, which tracks task progress and guides skill selection, with Experiential Memory, which stores what the agent has learned. Execution becomes structured evidence that localizes failures to specific memory components. A fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory, which reshape the next round of execution.
The results are the strongest in this cluster. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of 37 completed model-benchmark pairs, a 95% hit rate that's unusual for a memory wrapper. On tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%. On SkillFlow it adds +16.6 and +13.5 points to Qwen3.6-27B and Qwen3.6-35B. The advantage widens as the horizon grows: +32.2 points on the longest tasks, exactly where agents usually fall apart. Common long-horizon failures drop by up to 80%, which means fewer restarts and retries in production agent loops.
SkillForge attacks the same problem from the skill-bank angle. SkillRL extracts skills from raw trajectories but treats the skill bank as append-only. Skills go stale as the policy changes, and nothing verifies whether they still work. SkillForge makes skill usage explicit during agent interaction, so RL optimizes both environment actions and skill invocation decisions. Evidence-based verification and multi-pathway induction keep the bank growing without letting it rot. On ALFWorld, WebShop, and AppWorld, it consistently beats SkillRL.
The shared bet: knowledge should persist across episodes, and it should be verified before it's trusted. That's a different training model from GRPO-on-trajectories.
Rollout efficiency: SPO++
Group-relative RL, the GRPO family, waits for sibling rollouts of the same prompt before computing advantages. For short math problems, that's fine. For long, variable tool-use trajectories, the wait dominates the cost.
SPO removed the sibling dependency with a persistent prompt-level value estimate, but it had a subtle bug. Its recipe whitens one advantage per trajectory before optimizing a token-mean actor loss. Trajectory centering doesn't center the token-weighted quantity the actor actually consumes. SPO++ fixes the mismatch by standardizing terminal-outcome advantages under the action-token measure.
It also organizes prompt evidence by the policy event that generated it, rather than by learner receipt order. That matters for async training, where rollouts arrive out of order. Across matched runs on ALFWorld at two model scales and on Math-TIR, SPO++ improves online learning efficiency over SPO. The paired ablation identifies action-token-measure normalization as the strongest component.
Key numbers:35 of 37 model-benchmark pairs improved by Recuris across four long-horizon benchmarks. +32.2 points is the Recuris gain on the longest tasks, where the advantage widens most. 87.9% tau-bench accuracy for Claude Opus 5 with Recuris, up 15.6 points. 79.8% held-out accuracy for a 4B model trained with SMITH, ahead of an untrained 30B-A3B tool-writer. 0 new hyperparameters added by OPDVR to combine distillation with verifiable rewards.
Tools: SMITH
The outlier is SMITH, and it might be the most practically useful paper in the cluster. Tool-augmented language models are bounded by the APIs humans bothered to write. Existing tool-creation systems prompt a frozen LLM at inference time, which leaves the model that writes a tool decoupled from the one that uses it. There's no signal that the schemas it produces are schemas it can invoke.
SMITH trains tool creation and tool use inside a single policy. Each rollout is either a build task (write a tool from a few examples) or a use task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient.
The numbers matter because of the model size. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks reaches 79.8% macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer, a model roughly 7.5x larger. 4B parameters means this runs on a single GPU, no cluster required. It also reaches 40.4 on TabMWP-Hard and 42.6 on out-of-domain GQA, up 7.6 over the best same-backbone inference-time baseline, and it does this without any visual or tabular training data. The gains come from tool learning, not modality-specific pretraining.
The transfer result is the one to remember: tools written by the 4B model lifted the performance of LFM-2.5-350M and Qwen3-30B-A3B on the same reasoning tasks. A small model writes tools, and bigger models benefit. That's a capability-amplification story that doesn't require training the big model at all.
Common pitfalls
These are the mistakes I keep seeing in agent RL setups. Each paper in this cluster is a direct response to one of them.
Treating trajectory-level reward as token-level signal. SPO++ shows that per-trajectory centering doesn't center the token-weighted loss the actor actually consumes. If you're normalizing advantages for an agent policy, standardize under the action-token measure. Otherwise your normalization is silently misaligned with the loss.
Append-only skill banks. Skills extracted from old trajectories go stale as the policy improves. SkillForge's evidence-based verification exists because append-only banks accumulate noise. If your skill bank never deletes or refines anything, it's a museum, not a memory.
Trusting every factual signal equally. FARCA's central point: not all factual judgments are equally reliable. Weight a low-reliability signal the same as a high-reliability one and you inject noise into the policy update. Counterfactual dependence is a cheap, effective reliability proxy.
Decoupling tool creation from tool use. If a frozen model writes tools and a separate policy uses them, there's no gradient connecting schema quality to invocation success. SMITH's joint training closes that loop. If your tool-writer isn't your tool-user, you're optimizing blind.
Waiting for siblings. Group-relative methods are fine for short episodes. For long tool-use trajectories, the sibling wait dominates wall-clock cost. Single-stream value estimates, as in SPO and SPO++, are the escape hatch.
One thing to remember
The through-line across all seven papers is a shift from outcome-level, episodic, reward-driven training toward token-level, memory-augmented, evidence-driven training. The next generation of agent RL won't look like GRPO on trajectories. It'll look like a system that redistributes sparse rewards through structure, verifies what it learns, and carries that knowledge across episodes.
The bottom line
If you're training a multi-turn service agent with only outcome rewards, adopt influence-graph credit assignment before adding more reward shaping. IAPO gets better learning from the same rollouts, no new signals required.
If you're constrained by rollout cost on long, variable tool-use trajectories, skip group-relative methods and use stream-aligned single-stream optimization. SPO++ gives you the efficiency without the sibling wait.
If you're building agents that should improve across episodes, stop treating the skill bank as append-only. Add verification and memory evolution, the way SkillForge and Recuris do. One thing to watch: recursive memory evolution is the most speculative direction in this cluster, but it has the clearest scaling story. Expect it to become a standard post-training component within a year.