- LLM capabilities are spiky
- this is partially due to architectural constraints (math without a CoT) and partially due to training incentives (good prose?)
- LLMs are hacky, but also bad at design decisions/taste
- this is why LLMs currently shine when you have strong feedback, as well as lots of patterns they've seen before they can compose together (coding, formal verification, math)
- but is "taste" something that will get better with scale? something we can SFT/RL in?
- what is taste: judgement / smell generally useful for designing experiments, arguments, new types of solutions to an open problem, etc.
- My priors on this
- on the one hand, LLMs still aren't good at high-level research in my experience. They can come up with ideas for experiments/methods that superficially look decent, but often aren't very informative
- on the other hand
- LLMs have gotten better at tons of stuff that we used to consider higher-level and starts to look like taste (deciding whether to use a subagent or not, deciding a next action in a long-horizon debugging session)
- research/design stuff is hard (because there is less prior work to pattern match on), and maybe LLM capabilities are still just smoothly increasing up the difficulty curve
- whether "taste" is just another point on the same curve we've been climbing or if it's something that we won't crack with current methods determines which of two very different worlds we'll be in for the next 5 years
- world 1: taste is just another point on the curve, a little bit past where we are now.
- world 2: taste is something we won't crack with current methods for some reason (requires too much extrapolation, too hard to generate data, learning general taste is hard, etc.)
- In world 1, LLMs will probably just replace humans more and more (though it'll keep being smooth, and it will be a bit laggy due to the speed of diffusion, culture change, etc.). In world 2, LLMs provide increasingly good action abstractions, but don't get to the point where they can formulate interesting research directions, write good prose, etc.
- what data would make me update / be more confident about which world we're in?
- example of lab curating data for taste -> LLM learns taste in that domain
- seems like you could SFT on taste traces from a human expert, but idk how you'd RL on it? And SFT wasn't enough to improve decisions historically
- we can RL on research tasks that have some concrete metric, but these are the "easy" cases the model can already hill-climb on (easy for the model, but hard for impatient humans who get bored and need to sleep, eat, touch grass, ...)
- maybe all research questions can be framed with a metric that's far enough out (ex. solving interp = being able to make non-trivial formal claims, as this probably requires a good understanding of the structure of computation), but measuring this to RL (and to hill-climb) seems hard. Maybe there is some way to bootstrap it though?
- some refinement of whether taste is something qualitatively different
- benchmark on taste in some area showing growth, with good argument that this will continue
- example of lab curating data for taste -> LLM learns taste in that domain
Conceptual Reasoning Index
August 12, 2026 anthropic post, index leaderboard
Some folks at Redwood Research/Anthropic made some benchmarks meant to measure "conceptual reasoning" capabilities, i.e. their performance on tasks that "lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains."
They're motivated by similar reasoning: "Current AI training depends heavily on abundant data and reliable feedback on the model's performance. Models are therefore typically worse at tasks that cannot be empirically or mathematically verified."
They measure (via 3 separate benchmarks):
- models' ability to rate arguments "on a diverse range of topics, including decision theory, philosophy, and risks from advanced AI" compared to ratings from a few human experts (almost all ratings came from one person which is a little sus).
- models' consistency in beliefs/preferences. Ex. if a model says event A has probability P(A), then do they say
$P(A, B)\le P(A)$? They argue that lack of consistency would be evidence we cannot trust a model's conceptual reasoning - whether models can answer decision-theoretic questions correctly.
- you can see some example questions if you go to this paper and scroll to page 3. I'm not very convinced that performance on these questions closely tracks conceptual reasoning capability as they seem to test recall to me (and of the three this one seems pretty much saturated, which tracks)
Scores on the index have been increasing linearly since late 2024:

verdict: Ignoring the decision theory benchmark, this is evidence that models are becoming somewhat more consistent over time, and that their agreement with a few researchers on decision theory, philosophy, and some other topics is increasing as well. To me this is weakly positive evidence that LLMs are getting better at taste-relevant skills.
TASTE (The AI Safety Taste Evaluation)
August 28, 2026 anthropic post
In a similar vein, some anthropic fellows made a benchmark testing how much LLM reviews of pairs of AI Safety research proposals agree with "experienced human researchers".
They have claude generate 92 pairwise comparisons of research proposals (via seed prompts reverse engineered from anthropic fellow research proposals), and they have the human reviewers discuss for a while to try to get higher agreement on results. Before discussion agreement between researchers was 53% (~chance); after discussion + some filtering, this number jumps to 77%.
(side note: the fact that initial agreement between the smart researchers at anthropic was so low is extra evidence that hard signals of taste are hard)
So how well do models do?
Just around chance. And there doesn't seem to be a strong correlation with model release date. Though I'll admit this whole setup is noisy and they give some reasons models might close the gap to human performance soon ("models are doing poorly because they're making silly mistake X, which will presumably go away soon").
verdict: Superficially seems like a negative result on whether models are learning taste, but considering the initial 53% researcher agreement I don't feel comfortable drawing any conclusions from this.