Research AI-written

RLHF vs RLM: one changes the weights, the other changes what the model sees

Two acronyms that share two letters and almost nothing else. One is a training loop that ends in new weights. The other is an inference loop that never touches them. This post is diagrams first: each picture carries the point, the paragraph under it says what the picture cannot.

No original measurements here. Every number is a claim from the cited paper, stated the way its authors state it, and labelled by whether that paper was refereed.

Where the learning lives

Top lane: RLHF. The loop closes through a gradient step, and the thing that changes is the model. Bottom lane: RLM. The loop closes through a code runtime, and the thing that changes is what the model is looking at.

RLHF, reinforcement learning from human feedback, is a training procedure. It runs before the model is served, it needs people to compare outputs, and its product is a new checkpoint. A Recursive Language Model, RLM, is an inference procedure. It runs while the model answers, it needs a code runtime, and its product is an answer. The model that leaves an RLM call is byte for byte the model that entered it.

RLHF: three stages and a training run per change

The InstructGPT recipe. Yellow stages write to the policy's weights. The pink stage is where humans sit. The dotted return is why this is expensive: every improvement is another pass through all of it.

The recipe most people mean by RLHF is the one Ouyang et al. described for InstructGPT. Collect demonstrations and fine-tune on them. Sample several answers to a prompt and have labelers rank them. Train a reward model to predict those rankings. Then update the policy with reinforcement learning against that reward model: PPO in InstructGPT, and GRPO in later work such as DeepSeekMath. The idea underneath, learning a reward from human preferences between pairs, is from Christiano et al. in 2017.

What you get is behaviour baked into the weights. The InstructGPT paper reports that outputs from its 1.3B model were preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. What you pay is a training run for every change, plus the rollouts and the human labels that feed it. Nothing about the input is solved by RLHF. A prompt that is too long for the model is still too long after training.

RLM: keep the input out of the prompt

One RLM call. The long input is a variable in a Python REPL and never enters the context. The root model writes code to look at it, split it, and hand slices to recursive sub-calls. Only short results come back. Setting a Final variable ends the loop.

Recursive Language Models, from Zhang, Kraska and Khattab at MIT, start from a different complaint. Models get worse as prompts get longer, well before the context limit, a degradation the paper describes using Hong et al.'s term, context rot. Instead of pasting a long input into the prompt, an RLM stores it as a variable in a Python REPL. The root model sees the task and a short description of that variable. It writes code to peek at the input, chunk it, and call itself, or a smaller model, on each chunk. The sub-results come back as variables, and the root model reads those instead of the raw text. Setting a Final variable ends the loop. There is no gradient anywhere in it.

What is actually inside the context window

Left: the whole input occupies the context, and the model reads all of it on every turn. Right: the context holds the task, a few lines of code and short results, while the whole input sits in the runtime as a variable.

The whole trick is in what occupies the context window. The paper reports processing inputs up to two orders of magnitude beyond the model's context window, and on GPT-5 a median gain of 26% over compaction, 130% over a CodeAct scaffold with sub-calls, and 13% over Claude Code across four long-context tasks, at what the authors describe as comparable cost. Those are their numbers on their benchmarks, and I have not reproduced them. The cost I have measured myself is the other one: in the Kimi-Linear post a real one-million-token prompt worked, and the prefill took five hours. An RLM never prefills the whole input at all.

The grid, and the thing in between

Horizontal: when the technique runs. Vertical: whether the model changes. RLHF and GRPO change weights before serving. RLM changes nothing and runs while answering. GEPA sits in the corner the other two leave empty: nothing in the model changes, and the work usually happens offline. Top right, weights changing while answering, is test-time training, which exists but is not mainstream.

Put the two on a grid and a third thing appears between them. Prompt optimizers such as GEPA, the Genetic-Pareto optimizer, change the prompt text offline, scored against a metric, and leave the weights alone. Its authors report beating GRPO by 6% on average and by up to 20%, with up to 35x fewer rollouts. The paper's argument is that reflecting on a failed attempt in language is a richer signal than a scalar reward.

The top-right cell is not empty, only sparse. Test-time training puts a gradient step inside the answer: Sun et al. make the hidden state of an RNN a learner that updates on every token, Akyürek et al. fine-tune per instance on the few-shot examples in the prompt and then discard the update, TTRL runs GRPO on unlabeled test questions with a majority vote as the reward, and SEAL has the model write its own fine-tuning data. None of them is what gets shipped today, and TTRL in particular blurs the line the grid draws, since it is RL with no human labels at all.

So the choice is not two-way. If a behaviour must survive with no scaffolding around the model, train the weights. If the behaviour is a prompt-shaped problem and you have a metric, optimize the prompt. If the problem is that the input does not fit, or fits badly, keep it out of the context and let the model compute over it.

Where I intend to use it

Ax implements the RLM pattern in TypeScript as a three-stage agent: distiller and executor share one runtime session, the responder sits outside it, and sub-queries go through a budgeted sub-call. The first place I plan to try it is bank statement extraction, where reconciling balances across pages is a compute-over-the-data problem rather than a read-the-data problem. That will be a measured post, not a diagram post.

Sources

"Peer-reviewed" means the acceptance was checked, not assumed from an arXiv link. The RLM paper is the load-bearing source here and it is a preprint.