Where the learning lives
RLHF, reinforcement learning from human feedback, is a training procedure. It runs before the model is served, it needs people to compare outputs, and its product is a new checkpoint. A Recursive Language Model, RLM, is an inference procedure. It runs while the model answers, it needs a code runtime, and its product is an answer. The model that leaves an RLM call is byte for byte the model that entered it.
RLHF: three stages and a training run per change
The recipe most people mean by RLHF is the one Ouyang et al. described for InstructGPT. Collect demonstrations and fine-tune on them. Sample several answers to a prompt and have labelers rank them. Train a reward model to predict those rankings. Then update the policy with reinforcement learning against that reward model: PPO in InstructGPT, and GRPO in later work such as DeepSeekMath. The idea underneath, learning a reward from human preferences between pairs, is from Christiano et al. in 2017.
What you get is behaviour baked into the weights. The InstructGPT paper reports that outputs from its 1.3B model were preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. What you pay is a training run for every change, plus the rollouts and the human labels that feed it. Nothing about the input is solved by RLHF. A prompt that is too long for the model is still too long after training.
RLM: keep the input out of the prompt
Recursive Language Models, from Zhang, Kraska and Khattab at MIT, start from a different complaint. Models get worse as prompts get longer, well before the context limit, a degradation the paper describes using Hong et al.'s term, context rot. Instead of pasting a long input into the prompt, an RLM stores it as a variable in a Python REPL. The root model sees the task and a short description of that variable. It writes code to peek at the input, chunk it, and call itself, or a smaller model, on each chunk. The sub-results come back as variables, and the root model reads those instead of the raw text. Setting a Final variable ends the loop. There is no gradient anywhere in it.
What is actually inside the context window
The whole trick is in what occupies the context window. The paper reports processing inputs up to two orders of magnitude beyond the model's context window, and on GPT-5 a median gain of 26% over compaction, 130% over a CodeAct scaffold with sub-calls, and 13% over Claude Code across four long-context tasks, at what the authors describe as comparable cost. Those are their numbers on their benchmarks, and I have not reproduced them. The cost I have measured myself is the other one: in the Kimi-Linear post a real one-million-token prompt worked, and the prefill took five hours. An RLM never prefills the whole input at all.
The grid, and the thing in between
Put the two on a grid and a third thing appears between them. Prompt optimizers such as GEPA, the Genetic-Pareto optimizer, change the prompt text offline, scored against a metric, and leave the weights alone. Its authors report beating GRPO by 6% on average and by up to 20%, with up to 35x fewer rollouts. The paper's argument is that reflecting on a failed attempt in language is a richer signal than a scalar reward.
The top-right cell is not empty, only sparse. Test-time training puts a gradient step inside the answer: Sun et al. make the hidden state of an RNN a learner that updates on every token, Akyürek et al. fine-tune per instance on the few-shot examples in the prompt and then discard the update, TTRL runs GRPO on unlabeled test questions with a majority vote as the reward, and SEAL has the model write its own fine-tuning data. None of them is what gets shipped today, and TTRL in particular blurs the line the grid draws, since it is RL with no human labels at all.
So the choice is not two-way. If a behaviour must survive with no scaffolding around the model, train the weights. If the behaviour is a prompt-shaped problem and you have a metric, optimize the prompt. If the problem is that the input does not fit, or fits badly, keep it out of the context and let the model compute over it.
Where I intend to use it
Ax implements the RLM pattern in TypeScript as a three-stage agent: distiller and executor share one runtime session, the responder sits outside it, and sub-queries go through a budgeted sub-call. The first place I plan to try it is bank statement extraction, where reconciling balances across pages is a compute-over-the-data problem rather than a read-the-data problem. That will be a measured post, not a diagram post.