Models are very good at hill-climbing in easily-verifiable settings. There are already many problems you can effectively make progress on by throwing tokens at, typically at a rate much more efficient than human labor.

Some examples:

But models also reward hack a lot, especially on harder tasks. And for problems where success isn't easy to fully capture with a single metric, the quality of agent output is often subpar.1

This all seems to track with a model of LLMs as something that 1) can do a lot of low-level actions (write code, run this command, read this, etc.) with fairly high reliability, 2) can generate enough diversity of mediocre/straightforward ideas, and 3) is very cheap to run (a few OOMs less than humans, given some way of matching human thought to a token stream). These properties are all you need to do extensive search in domains that have strong and dense (preferably un-hackable) reward signals (math, making tests go green, making something run faster).2

I can imagine two trajectories in LLM capabilities from here:

  1. World 1. LLMs smoothly become able to work at longer timescales, on harder problems, and in less verifiable domains. To accomplish this, they make less mistakes in the first place and are better able to judge intermediate progress when they don't have access to a hard metric.
    • concretely: I would be able to point an LLM at a vague research question like "develop a method to uncover computational structure in transformer weights such that we can make non-trivial predictions about model behavior without running the model" and it would actually make significant progress on it, despite the space being large and the model having to work for a long time before seeing hard progress.
  2. World 2. LLMs continue to get better at hill-climbing in verifiable domains, but they don't get much better (in the current regime) at fuzzier problems because 1) their initial ideas and actions stay sub-human, and/or, 2) they fail to make good judgement calls about intermediate progress and so cannot tackle problems where the reward signal is too far off.
    • concretely: pointing an LLM at a research question like the above yields a similar outcome to today. The LLM is able to write any code that I want, but isn't able to make progress in domains where feedback is sparse.

I would love to feel more confident about which world we are in, but I'm really not sure. As of now, I'm leaning slightly toward world 1.


Whether LLMs can learn to accurately judge intermediate progress is important.

  • LLMs are really good at working on open math problems. OpenAI has been having their internal model hill-climb open problems in natural language and formalizing in lean after the fact. The Navier Stokes blowup proof is hundreds of pages; did this require math "taste" to tell the agents they were on the right path while they were incrementally building up lemmas?
    • My intuition is that even though the distance to the finish line of proving a hard math problem is great, you're still getting fairly reliable dense feedback along the way. The LLM understands proofs enough to check intermediate work and it's sufficiently obvious whether a particular intermediate lemma might be useful that it's able to search effectively and expand the frontier of proof approaches until it hits on a winner.
    • I think that being in world 2 would require that instilling into LLMs judgement of intermediate progress for fuzzier domains than math proofs is too hard for the current regime of scaling/RL.

Why haven't we seen incredible optimization results by AI systems on nanogpt? The reward is dense/robust, but the AI contributions have been pretty shallow, with the deep improvements, ex. Muon, coming from human ingenuity.

  • It's likely that if OpenAI pointed a swarm of instances of their internal model at it then we probably would see more improvements. While the training time is low, it does still require 8xH100 GPUs to run experiments, which makes extensive search expensive. But also maybe it's harder than math (for LLMs) in some other way?
1

METR has found that just judging off of tests-passing (this is how SWEBench does it) overstates the quality of LLM code - even with tests passing many of the AI PRs wouldn't be merged by the maintainers

2

Strong/dense feedback makes it easier to RL but also makes test-time much more forgiving: Dense/cheap feedback means LLMs can try a lot of things and don't need to be geniuses at each step. Plus if your verifiable metric is not just a proxy it's less Goodhart-able and means reward-hacking is much less of a possibility (ex. if the model writes a lean proof that typechecks, the proof is either correct or they're exploiting a bug in the lean kernel which is possible but pretty rare).