Appearance
The reward problem in code RL
Reinforcement learning with verifiable rewards (RLVR) is the engine behind most reasoning-model gains of the last two years. The recipe is simple: sample rollouts, check them against a known answer, convert the pass/fail into an advantage, update. It works out of the box for math, where the answer is a string you can compare. Code does not reduce to that, and the failure is structural.
Test-time RL (TTRL) methods typically derive rewards from answer-level self-voting: sample several outputs, check which ones agree. For problems with canonical answers, that's a cheap pseudo-verifier. Programs have no canonical surface form. Two correct solutions can be textually disjoint, so agreement on text says nothing about correctness, and disagreement says nothing either.
The ERPO paper solves this with behavioral agreement. It constructs probe inputs from the problem statement, executes candidate programs on those probes, and scores the outputs. Two programs that behave identically on well-chosen probes are plausibly solving the same problem. That behavioral consensus is the Probe Consensus Reward (PCR), and it turns open-vocabulary programs into a training signal without needing a reference solution.
The catch is that PCR is not a fully reliable verifier. Wrong programs can agree spuriously. ERPO pairs PCR with two guards. Rank masking converts low PCR into conservative negative updates instead of treating them as plain failures. An entropy ceiling stops the policy from drifting into whatever output pattern happens to agree on probes. The result: pass@1 and pass@k both improve, in-domain and in zero-shot transfer. pass@1 is your best single attempt; pass@k is what you get when you sample k programs and take any that passes. The pass@k gain is the one that matters for test-time scaling, because more sampling budget only helps if the coverage is there.
Here is where each piece of this cluster plugs into the RLVR loop:
Keep that map in mind. Every paper and dataset in this cluster targets one box in the loop, and none of them touch the optimizer itself.
Verifiable data at scale
RLVR is only as good as its verifier, and a verifier is only as good as the data under it. That's the gap openbmb targets with UltraData-RL-2609, the L3 slice of their L0-L4 data framework. It powers the RL stage of MiniCPM5-2B and feeds JustRL II, which scales small LLMs to 128K-reasoning RL with a critic. The release contains 85,995 verifiable tasks across four domains, all normalized into the same JSONL format.
| Domain | Samples | Share | How outcomes are verified |
|---|---|---|---|
| Math | 32,412 | 37.7% | Answer match against ground_truth |
| Code | 23,665 | 27.5% | Execute the submission against test cases |
| Long-Context | 18,046 | 21.0% | Answer match, context guaranteed to support the answer |
| Knowledge | 11,872 | 13.8% | Answer match against ground_truth |
| Total | 85,995 | 100% |
Three design goals run through every domain: verifiable outcomes, trustworthy rewards, calibrated difficulty. Verifiability means every item has an explicit reference target and a mechanical way to judge it. Math and Knowledge items are rewritten into short-answer form; Long-Context items must show an explicit context-question-answer link; Code items must execute in a sandbox against test cases. For Code, the ground_truth holds paired stdin/stdout arrays, and verification pipes each input and compares stdout after trailing-whitespace normalization.
Trustworthiness comes from consistency checks. Answers are re-solved by independent models and relabeled only by consensus; when there's no consensus, the item is dropped, not guessed. Code test cases, reference solutions, and expected outputs are cross-validated; invalid tests are removed.
Difficulty calibration is the part to steal. Repeated rollouts on the RL initialization checkpoint estimate each item's pass rate. Items the starting model already solves (pass rate 1) are removed, since they produce no advantage signal anyway. The learnable band is kept. Hard items with a confirmed label (pass rate 0) are kept and scheduled by online dynamic sampling. The filtering changes difficulty and sampling weight, never the label. A lot of RL data pipelines "fix" hard items by relabeling until the model passes. This one treats the label as sacred and difficulty as a scheduling problem.
The headline numbers justify the effort. In the JustRL II setup, AIME 2025 goes from 61 to 81 in about 300 RL steps, and the final MiniCPM5-2B checkpoint that consumes UltraData-RL-2609 hits 86. That's a 2B-parameter dense model, small enough to run on a phone. The companion UltraData-Code corpus adds per-file algo-relevance and quality scores for the code SFT side, but the RL release is where the verifiable-reward action is.
Quick Take: The pattern across this cluster is consistent: don't give the policy a reward you can't verify, and don't waste rollouts on prompts it already mastered or will never solve.
The selection problem
ThinkPrior attacks the second half of that sentence. In GRPO, the advantage is computed within a group of rollouts for the same prompt. If every rollout in the group is correct, or every one is wrong, the group-relative advantages are identically zero. No gradient. The paper calls these silent groups, and uniform sampling spends 39% of a run's rollouts on them. Nearly two in five rollouts contribute nothing to training.
Prompt selection has a cold-start problem of its own. History-based methods need target-policy rollouts to estimate which prompts are learnable, and those estimation rollouts are the same wasted resource you're trying to save. ThinkPrior instead constructs a difficulty prior in one offline pass, before the first target-policy rollout. An external anchor, the verifier-scored pass rate of each prompt on a reference model, initializes a Beta posterior. Selection ranks by expected learnability, and training outcomes update the posterior as you go. The loss and optimizer are untouched. This is a scheduling layer, nothing more.
Measured effects: early silent groups more than halve, wasted rollouts through step 30 drop by nearly a fifth, and the paper detects no difference in final accuracy. On a fixed 250-prompt pool, the budget is reallocated, not saved. You don't get a better final model; you get the same model for less compute. Composed with DAPO, generated rollouts drop 10.6% while both arms keep the same 3840-rollout update budget.
Key numbers
- 39% of a run's rollouts land on silent groups under uniform sampling
- 10.6% fewer generated rollouts with ThinkPrior+DAPO, same update budget
- Silent groups more than halved; wasted rollouts through step 30 down by nearly a fifth
The exploration problem
RLVR-trained models hit a second wall: pass@1 improves, pass@k stalls. The model gets better at one good attempt, but its coverage of possible solution paths doesn't expand, which caps test-time scaling. DATPO treats this as a rollout-structure problem and extracts three design principles from the analysis.
First, difficulty-adaptive rollouts matter for coverage, not just as an efficiency heuristic. Sample easier prompts when the policy is weak, and let the tree deepen as it strengthens. Second, tree-based rollout discovers correct answers more often than parallel sampling at the same budget. Third, branching should happen at sentence level, not token level. Token-level branching localizes: it produces near-duplicate continuations. Sentence-entropy-guided forking maximizes semantic diversity, so the siblings in the tree are genuinely different lines of reasoning.
The method wraps these into a tree-structured rollout with a sibling-diversity advantage term. On math benchmarks, DATPO beats GRPO-style baselines most clearly on pass@k, and pass@k is exactly what test-time scaling consumes. If you're spending inference budget on k samples at eval, this is the training-side change that makes that budget pay.
Learn to test, test to improve
ExecCritic is the most transferable result in this cluster, because it targets a failure mode that doesn't get enough attention: test generation. When the same trajectory writes the patch and the test, errors can agree. The patch fails, the test encodes the same misunderstanding, everything passes, and the loop converges to false confidence.
The scaffold separates the two roles. A Test agent independently generates repository-native tests. A fail-closed harness qualifies and freezes them. A Repair agent then revises the source code from execution feedback, with no power to change the tests. Both roles use Qwen-3.5-35B-A3B as the backbone: 35B parameters total, but only 3B active per token, so running two trained agents stays affordable. They are trained separately with role-specific RL.
The numbers make the point better than any argument. Holding the base Repair agent fixed, tests from the base Test agent reduce the resolved rate on SWE-bench Verified from 61.2% to 57.3%. Bad tests don't just fail to help; they actively cost you four points of resolution. Tests generated by GPT-5.6-sol raise it to 65.3%. Role-specific post-training lifts the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%, and composing the two post-trained agents reaches 72.6%, an 11.4-point gain over the no-test baseline. No stronger model and no Oracle feedback at evaluation time.
The lesson for anyone building coding agents: your test generator is not a side detail. It is the reward pipeline. When I first ran a test-verify-revise loop, my Test agent wrote tests that passed the broken patch, and the loop converged to nothing. Separating the roles and freezing tests before repair is what forces the feedback to be honest.
Agent SFT: the other half
RL assumes the policy can already produce interesting rollouts. The SFT stage is what gets it there. UltraData-SFT-Agent-2609 ships 483,661 agent trajectories for exactly that, again for MiniCPM5-2B. The breakdown: General-Agent 311,006 (64.3%), Tool-Use 82,760 (17.1%), Code-Agent 69,895 (14.5%), Search-Agent 20,000 (4.1%). Counts are trajectories, not unique tasks, because some tasks were resampled under different harnesses.
Two design choices stand out. Turn-level masks let you exclude low-quality, redundant, or malformed turns from the SFT loss without discarding the whole trajectory. That's the right granularity. And the release keeps some diagnostic failures: timeout, tool-error, and recovery traces, stored as auxiliary supervision, not as success demonstrations. Training on what recovery looks like is how agents learn to recover.
I've trained on static trajectory dumps, and this release has the usual constraints plus a few handled well. Tools are frozen from the sampling harness, so you must render them into your chat template yourself, and many virtual APIs don't exist in real deployments. Web pages, files, dependencies, and APIs are snapshots at sampling time; they go stale. Passing the environment's success check doesn't mean every step was correct, which is exactly why the turn-level mask exists. The release also screens against evaluation sets known at construction time, so run a fresh contamination check before introducing a new benchmark. And some samples originate from publicly accessible course pages: public access is not a grant of redistribution or training rights, and Apache 2.0 doesn't override upstream licenses.
Common pitfalls
Five things I've seen trip people up, in order of how expensive they are.
Deriving code rewards from surface agreement. If you sample k programs and vote on their text, you learn nothing; two correct solutions can share zero tokens. Use execution-based signals: probe consensus as in ERPO, or a real test set.
Letting the same model write the patch and the tests. Correlated errors produce false confidence, and ExecCritic quantifies the cost: bad tests cost 3.9 points over no tests at all, while a trained test role adds 11.4. Split test generation from repair, freeze the tests, then fix the code.
Training on prompts with no advantage gradient. All-correct and all-wrong groups both produce zero signal, and uniform sampling burns 39% of rollouts on them. Calibrate difficulty against your own initialization checkpoint, exclude mastered items, and schedule hard-but-valid ones online.
Trusting a consensus reward without drift control. Spurious consensus is real; rank masking and an entropy ceiling exist because PCR alone is hackable. If your reward is behavioral agreement, pair it with a constraint that stops the policy from collapsing onto the agreeing pattern.
Treating agent trajectory data as a live environment. Frozen tools, stale snapshots, virtual APIs that don't exist in deployment, trajectories that can't be replayed. Verify what was frozen, render the tools field into your template, and remember that verification is not process quality.
One thing to remember
Every method in this cluster works upstream of the optimizer. ERPO fixes the reward, ThinkPrior fixes selection, DATPO fixes exploration, ExecCritic fixes the verifier, and the UltraData releases fix the data. GRPO itself barely changes. The remaining headroom in RLVR for reasoning and coding is in the reward pipeline and the rollout budget, not in the policy update rule.
Summary of recommendations
Four bottlenecks, four fixes.
| Bottleneck | Method | Mechanism |
|---|---|---|
| No reward signal for open-vocabulary programs | ERPO | Probe consensus reward with rank-masked, entropy-regularized updates |
| Cold-start prompt selection | ThinkPrior | Zero-rollout difficulty prior from an external anchor, Beta posterior |
| Stalled pass@k coverage | DATPO | Difficulty-adaptive tree rollouts with sentence-entropy forking |
| Weak or self-confirming agent tests | ExecCritic | Separate test generation from repair, fail-closed harness |
- If you're post-training a small on-device model for reasoning and coding, mirror the UltraData recipe: verifiable outcomes, consensus-checked labels, difficulty calibrated against your own init checkpoint. That's the setup behind MiniCPM5-2B going from 61 to 86 on AIME 2025 at 2B parameters, small enough to run on a phone.
- If you're building a coding agent on repository tasks, split test generation from code repair and train both roles. ExecCritic shows the same 35B-A3B backbone goes from 57.3% to 72.6% on SWE-bench Verified once the test role learns to distinguish correct from incorrect patches.
- If your RLVR runs keep burning compute, add a ThinkPrior-style difficulty prior before the first target-policy rollout. It won't improve final accuracy on a fixed pool, but 10.6% fewer rollouts at the same update budget is real money, and you can spend the savings on DATPO-style tree exploration to grow pass@k for test-time scaling.
One thing to watch: test generation is becoming a trained capability rather than a byproduct of patch training, and the difference between a useful and harmful test generator is now measured in whole benchmark points. Expect test-specific RL to be a standard agent-training stage within two release cycles.