Appearance
The Training Pipeline Is the Bottleneck
The base model is rarely the bottleneck when you're shipping an agent. The pipeline around it is. I've hit this wall on every agent project I've worked on: skills that don't transfer, environments that never push back, and RL loops that can't tell which turn went wrong.
These six papers treat that sloppiness as the bug. They span five stages of the agent lifecycle, and each one fixes a failure mode that shows up in real deployments.
| Paper | Pipeline stage | Core mechanism | Headline result |
|---|---|---|---|
| MidTool | Mid-training | Synthesized tool-use corpus from real APIs, MCP skills, and document workflows | Consistent gains on BFCL, tau2-Bench, and MCP Universe under both SFT and RL |
| Break It Down, Pass It On | Skill induction | Subtask-level, text-format skill induction | Subtask skills beat the no-memory baseline; task-level skills hurt |
| Best Prefix Selection (BPS) | Skill selection | Submodular optimization under a hard token budget with a (1-1/e, 1) guarantee | 0.73 task success vs 0.20-0.52 for baselines, on 28% fewer tokens |
| EnvHarness | Environment generation | Plug-in components wrap static environments; black-box synthesis of reshaped tasks | Up to +9.0 points on held-out instances with 9.8% fewer steps |
| SAPO | RL credit assignment | Shared policy/value backbone, single rollout, trajectory-level GAE | +15.1 points over PPO, +12.1 over GRPO, 33.2% faster per iteration |
| MileGPO | RL credit assignment | Milestone discovery, reliability-calibrated shaping, progress-contrastive calibration | State-of-the-art on ALFWorld and WebShop with a small OOD gap |
Read the table as a pipeline, not a list. Data synthesis feeds the base model. The skill memory decides what the agent carries into the context window. The environment decides what the agent practices. The RL loop decides what the agent learns from practice. Miss any stage and the others cap out.
Mid-Training: Tool Use Needs Its Own Stage
Mid-training sits between pretraining and post-training. It's already proven for math, science, and software engineering. MidTool asks whether general tool use deserves the same treatment. The evidence says yes.
MidTool is an open corpus construction pipeline. It combines large-scale web, PDF, and code data with synthesized supervision from real tool APIs, MCP skills, and document-grounded workflows. The corpus teaches four things: recognizing tool affordances, grounding arguments from context, composing tool-call workflows, and recovering from incomplete information.
The authors mid-train Qwen3-4B-Base and Qwen3-8B-Base on MidTool-Mix, then apply SFT and RL. Compared with baselines, MidTool-Mix improves downstream performance on BFCL, tau2-Bench, and MCP Universe under both training regimes.
Practical translation: 4B parameters runs on a single node of consumer GPUs, so this is reproducible without a cluster. The result matters beyond the benchmarks. General tool use has been treated as something post-training conjures out of thin air. MidTool says it's a learned capability with its own data requirements, just like math.
Skill Induction: Where Transfer Breaks
Break It Down, Pass It On runs a controlled study of how skill induction shapes transfer. Two axes: task-level vs subtask-level induction, and text vs code format.
The results are stark. Task-level skills mostly reduced the agent's performance below its no-memory baseline. Subtask-level skills raised it above baseline on average. Text skills transferred better than code skills.
The explanation comes down to two properties. Specificity measures how closely a skill matches real tasks. Abstractness measures how evenly its relevance spreads across tasks. Neither alone predicts task success, but their combined skill utility score does, and it correlates consistently with transfer success.
The practical payoff: computing skill utility only needs the skills and task descriptions, no execution. You can score a skill memory before spending a single rollout on it.
Key Numbers
- Task-level skills dropped agents below their no-memory baseline; subtask-level skills raised performance above it.
- Text skills transferred better than code skills in the controlled comparison.
- BPS hit 0.73 task success vs 0.20-0.52 for released routers and retrievers, on 28% fewer tokens.
- SAPO beat PPO by +15.1 points and GRPO by +12.1 points, cutting per-iteration runtime 33.2%.
- EnvHarness improved held-out task success by up to 9.0 points with 9.8% fewer execution steps.
Skill Selection: Top-k Is Not a Strategy
Most agents select skills by scoring each one against the task and taking the top-k. The BPS paper points out the flaw: independent scores ignore the set. Redundant skills burn context tokens, and poorly chosen skills can actively degrade performance.
They model skill selection as an optimization problem: choose a set under a hard token budget to maximize a monotone submodular benefit minus a context penalty. Best Prefix Selection (BPS) solves it in polynomial time with a bicriteria (1-1/e, 1) approximation. It's the first provable performance guarantee for skill selection.
On a contamination-controlled BigCodeBench variant, BPS reaches 0.73 measured task success. Released skill routers, text retrievers, and the executor's own selection land between 0.20 and 0.52. That's 40% better than the strongest baseline and more than 3x the weakest. And BPS does it on 28% fewer tokens than the strongest released router.
28% fewer tokens on a 32K context is roughly 9K tokens per episode freed up. That's extra room for tool outputs and reasoning, or a smaller context window and cheaper inference.
Quick Take: Agent capability is a pipeline property, not a model property. Every paper in this cluster improves the agent by fixing the training loop around it, not by scaling the model.
Environments That Push Back
Static environments are the quiet ceiling on agent RL. The agent improves, the environment doesn't. EnvHarness treats the environment as a training signal instead of a fixed constant.
EnvHarness is a programmable layer of plug-in components that wraps a static environment and reshapes its behavior without modifying the underlying logic. It operates through standard interfaces, so it applies across domains, and every reshaped environment keeps its original verifier. That last point matters: no verifier, no reliable reward.
EnvRigger automates the wrapping. It treats the target policy as a black box, observes execution trajectories, synthesizes components that target diagnosed flaws, and validates them with fresh rollouts.
Across five benchmarks in four domains, EnvHarness outperforms both the original environments and domain-specific environment generation pipelines. Up to 9.0 points on held-out instances with 9.8% fewer execution steps. The held-out gain is the one to focus on: those are the harder, unseen tasks, and the agent solves them more efficiently.
It also gives RL a better optimization signal. The policy and environment co-evolve, which is closer to how real skill acquisition works than training against a frozen world.
Credit Assignment: The RL Bottleneck
Agentic RL has settled on critic-free, group-relative methods like GRPO: sample multiple rollouts, compare them, estimate advantages. They avoid PPO's separate critic memory, but these papers identify three cracks. No explicit value generalization. Advantage collapse in long-horizon tasks. And a costly tradeoff between sampling budget and performance.
SAPO (Single-rollout Autoregressive Policy Optimization) attacks all three. Policy and value share a single autoregressive backbone, with predictions at distinct causal boundaries. The PPO objective and an auxiliary on-policy SARSA objective are optimized independently. A trajectory-level GAE combines lambda-returns with batch normalization to estimate each turn's contribution.
On ALFWorld and WebShop with Qwen2.5-1.5B/7B, SAPO beats PPO by a mean of +15.1 points and GRPO by +12.1 points. It eliminates the separate critic's memory cost and cuts per-iteration runtime by 33.2% versus PPO. A third of your training wall-clock back, with better results.
MileGPO takes a different route to the same problem. It derives process-level credit from grouped on-policy rollouts. Milestone discovery identifies candidate milestones on successful rollouts and recurring traps on failed ones. Reliability-calibrated shaping weights those candidates by outcome-based confidence. Progress-contrastive calibration checks whether a candidate reflects local progress and whether its incoming transition beats observed alternatives from the same state.
No auxiliary models, no extra environment interaction. State-of-the-art results on ALFWorld and WebShop, and a small in-distribution to out-of-distribution gap on ALFWorld. That gap matters: it means the credit assignment generalizes instead of memorizing the training tasks.
Common Pitfalls
Scoring skills independently and taking top-k. The set has redundancy and interaction effects that independent scores miss. BPS's gap, 0.73 vs 0.20-0.52, is not a tuning issue. It's a different problem formulation. If you're doing semantic-search retrieval into a context window, switch to set-level selection.
Inducing skills at task level. The controlled study is unambiguous: task-level skills landed agents below their no-memory baseline. Break tasks into subtasks before writing the skill. A skill that mirrors one full task is a memorized trace, not a transferable capability.
Storing skills as code because it feels more precise. Text skills transferred better in the same controlled comparison. Code format bakes in implementation details that don't survive the trip to a new task.
Training RL on a static environment. The agent overfits the hand-built world, and the environment never pushes back. EnvHarness's 9-point held-out gain shows the environment is a first-class training signal. If your environment doesn't adapt to the agent's weaknesses, you're leaving performance on the table.
Trusting group-relative advantages for long horizons. Advantage collapse is real in 20-turn trajectories. Either share a value head like SAPO or extract milestones like MileGPO. Final-reward-only credit assignment won't carry long-horizon tasks.
One Thing to Remember
The through-line across all six papers: agent capability is a pipeline property. The data stage, the skill memory, the environment, and the credit assignment each need deliberate design. Fix the weakest stage and you'll see bigger gains than swapping the base model for a bigger one.
The Bottom Line
- If you're building an agent that calls tools, add a mid-training stage on tool-use data before SFT or RL. MidTool's corpus is open, and the consistent gains on BFCL and MCP Universe show general tool use doesn't emerge from post-training alone.
- If you're constrained by context budget, replace top-k skill retrieval with set-level selection. BPS gives you a provable guarantee and 28% token savings, the cheapest win in this cluster.
- If you're training long-horizon agents, move off pure group-relative advantages. SAPO and MileGPO both beat GRPO by double digits, and one of them will likely become the default within two release cycles.