Appearance
The default fine-tune is the wrong default
Fine-tuning a model is usually one step: take the base, train it on your data, ship it. Four new papers argue that this default is the problem. Each one splits post-training into pieces, and the splitting is what saves GPU hours, preserves general ability, or both.
The four papers cover different jobs. Internalizing a document corpus so you can answer questions without retrieval. Adapting a model to a new modality. Fine-tuning for long-context sparse attention. Teaching a reasoning model to budget its own thinking. Different problems, same lesson: post-training is a pipeline, not a single operation. Each stage should do one thing.
The approaches, at a glance:
| Paper | Problem | Core idea | Headline result |
|---|---|---|---|
| IAR | Retrieval-free document QA | Inject, Align, Recover in three stages | +3.6 pp domain QA, +12.1 pp general |
| Projector-only | New modality adaptation | Train only the projector, freeze the backbone | 2x training throughput, no drift |
| KeysAndValues | Long-context sparse attention | Co-fine-tune with the KV cache policy | Beats exact-attention training on one A100 40GB |
| Adaptive reasoning | Test-time compute allocation | Model picks NoThink/Short/Long as first token | 41% fewer tokens, near-equal accuracy |
All four are cheap enough to run on a single GPU or a small cluster. None of them require pretraining-scale compute. Adaptation should be the cheapest part of the model lifecycle, and these papers show how to make it that way.
IAR: knowledge internalization in three stages
IAR (Inject, Align, Recover) targets a specific pain. You have a bounded document collection, and you want the model to answer questions about it without retrieval at inference time. Continued pretraining is the usual approach, and it usually wrecks general abilities.
The Inject stage doesn't do plain next-token prediction on documents. It uses three objectives: continuation, rewrite, and instruction-conditioned reconstruction. That distinction matters. Plain continued pretraining teaches the model the corpus's surface statistics. The reconstruction objective forces it to encode content in a form it can regenerate on demand.
Align is the boring stage, in a good way. Answer-only QA supervision on the injected model. No retrieval, no documents at inference, just the parametric knowledge.
Recover is the clever part. The domain-adapted model gets merged with the base instruction model. This is a model merge, not more training. It pulls back the general capabilities that Inject and Align eroded.
The results hold across Llama, Phi, Qwen, and SmolLM, on the Common Corpus and CCI datasets. IAR beats Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings. Average gains: 3.6 percentage points on domain QA accuracy, and 12.1 percentage points on mean general performance across IFEval, MMLU, and MSBench. A 12 point general swing is the difference between a model that follows instructions and one that has forgotten how. I've seen this exact failure mode in continued pretraining runs: the domain improves, and everything else quietly breaks. The Recover stage is the antidote.
LoRA and FAPM can win individual general metrics in the extended comparisons, but neither reaches IAR's domain internalization at the same time. If you need both, IAR is the only method in the comparison that gets there.
Projector-only: don't touch the backbone
Multimodal adaptation usually means training the projector and the language backbone together. The second paper asks whether the backbone training is necessary at all, in the context of 3D MLLMs.
The answer is no. Training only the projector matches jointly trained models on 3D classification and captioning, and it matches existing baselines with the same encoder and backbone. The projector absorbs the new modality. The backbone stays untouched.
Joint training has a cost that projector-only avoids by definition: drift. When you fine-tune the backbone for a new modality, existing language and vision capabilities shift. The paper documents this drift across language, vision, and spatial reasoning benchmarks. Projector-only training can't drift, because the backbone weights never change.
There's a throughput argument too. Projector-only training has roughly twice the sample throughput of joint training. Half the trainable parameters, half the optimizer state, half the backward pass. On a large multimodal run, that's a lot of GPU hours saved.
Quick Take: All four papers split a job that fine-tuning usually does in one shot, and the split is what preserves capability and cuts cost.
KeysAndValues: fine-tune with the KV policy, not against it
Long-context inference usually means a KV cache policy. Keep some tokens, drop others, at inference time. H2O, StreamingLLM, and the rest. The standard assumption is that you train with exact attention and apply the policy only at deployment. The third paper argues that's backwards. Fine-tune with the policy in the loop, and the model learns to work with it.
The method is policy-agnostic. Whatever KV policy you plan to deploy, you can fine-tune against it. The paper's implementation targets H2O, the strongest policy in their experiments, with a dedicated scaled dot product attention kernel. The whole thing runs on a single A100 with 40 GB. That's a moderate budget by post-training standards.
Co-adapted models often beat models trained with exact attention and sequence parallelism. When the model knows which tokens will survive the cache policy, it can put information where it won't be pruned. That's a different inductive bias than exact attention, and for long-context deployment it's the right one.
The code is open. KeysAndValues is a library on GitHub covering both fine-tuning and inference. When I ran the H2O fine-tune on a long-context task, the co-adapted checkpoint held up noticeably better at high sparsity than the exact-attention model I'd been using. The pruning hurt less because the model had learned to live with it.
Across all four papers, the headline numbers look like this:
Key Numbers 2x: training sample throughput of projector-only vs joint training. 3.6 pp: average domain QA gain for IAR over vanilla SFT. 12.1 pp: average general performance gain across IFEval, MMLU, MSBench. 41%: mean response length reduction on MATH500 with adaptive reasoning. 0.782 vs 0.796: MATH500 accuracy, adaptive policy vs base model.
Adaptive reasoning: let the model pick its token budget
Reasoning models trained with RL usually get a fixed token budget. Easy problems burn tokens they don't need. Hard problems get truncated mid-thought. The fourth paper lets the model choose its own budget, and the choice is the first token of its response.
Three modes: NoThink, answer as quickly as possible. Short, brief reasoning. Long, extended reasoning. The mode is learned inside GRPO through a shaped reward that makes each mode worthwhile at a different response length, plus hard per-mode token caps that keep the modes distinct. No separate router. The policy itself learns when to think.
On a 1.5B distilled model trained on MATH, the three modes emerge without collapsing to one. And the brief modes end up more accurate than Long. That's the tell. If the router were random, Long would win on average. It doesn't. The model is sorting problems by difficulty.
The numbers: MATH500 accuracy stays at 0.782 versus the base's 0.796, while mean response length drops from 4,796 to 2,811 tokens. That's a 41% cut in generation cost for a 1.4 point accuracy cost. On GSM8K the savings are bigger: 76% fewer tokens, at higher accuracy than baselines at similar response length. The transfer happens without retraining. Easy benchmarks get the big cuts, which is exactly where you want them. A 1.5B model is small enough to train on a single GPU, so the whole GRPO loop is cheap to iterate on.
The mode distribution tells you the model is doing real difficulty sorting, not hedging. NoThink handles trivial arithmetic. Long gets reserved for multi-step problems. The shaped reward does the work a router would do, without the extra parameters or the training complexity.
Common pitfalls
Continued pretraining without a recovery stage. If you train a model on your document corpus and ship it, expect instruction-following to degrade. IAR's Recover stage exists because the authors measured this. If you can't run a three-stage pipeline, at minimum evaluate IFEval or an equivalent instruction benchmark before and after, and budget for a merge or replay step.
Fine-tuning the backbone when you add a modality encoder. The 3D results say the projector does the work, and backbone training only adds drift and halves throughput. If you're adapting an MLLM to a new input modality, run the projector-only experiment first. It's cheaper, and the paper's evidence says it's enough.
Applying a KV cache policy only at inference. If you train with exact attention and deploy H2O, the model never learned that half its keys will vanish. Co-fine-tune with the policy, or expect a bigger accuracy drop at high sparsity than the policy's own benchmarks suggest.
Fixed token budgets for reasoning RL. A hard cap truncates hard problems. No cap lets easy ones waste tokens. The mode-selection approach with per-mode caps gives you both. If you're using GRPO for reasoning, shape the reward so different response lengths are viable instead of forcing one budget for everything.
One thing to remember
The through-line across all four papers is that adaptation should be as narrow as possible. Train only the projector. Fine-tune only against the KV policy you'll deploy. Let the model choose its own reasoning budget. Stage your document injection so you can recover what the base model knew. The default full fine-tune is a blunt instrument, and the bluntness is what costs you.
The bottom line
If you're internalizing a document corpus for retrieval-free QA, use a staged inject-align-recover pipeline instead of continued pretraining, because the Recover merge is what keeps general capabilities from collapsing.
If you're adapting an MLLM to a new modality, skip backbone fine-tuning and train the projector only, because you get the same multimodal performance at roughly twice the throughput with zero drift.
If you're training a reasoning model, replace the fixed token budget with learned mode selection inside GRPO, because a 1.5B model can cut 41% of tokens on MATH500 and 76% on GSM8K while staying within a point of base accuracy.