RLVR (verifiable rewards)
Plain English. Reinforcement learning from verifiable rewards: training a model against tasks where a machine can check the answer — the maths proof compiles, the code passes tests, the number matches. No human raters, no learned judge; the reward is ground truth.
Why it moves money. RLVR is the engine of the reasoning-model era, and it draws a bright line across the economy: domains with checkable answers (code, maths, parts of law and finance) can be trained on at machine scale and automate first; domains where quality is a matter of taste improve much more slowly. "Is it verifiable?" has become a serviceable one-question screen for which industries face automation on what timeline — and for which AI products can actually deliver claimed reliability rather than asserting it.
What to watch. Progress in non-verifiable domains, where labs substitute learned judges and rubrics — and reward hacking, because a model trained against a checker learns to satisfy the checker, which is not always the same as being right.
From the signals. Star Fleet set twenty agentic harnesses on open maths problems, checked by Lean 4. The economics of recursive self-improvement meets verification design.
Further reading. Lambert et al., Tülu 3 — the paper that named RLVR (2024) · DeepSeek-R1 (2025)