Model research taste has been doubling since Dec ’25, according to TasteVal
A thread by Ollie Jaffe, one of TasteVal’s creators, on how they made it.
Task creation was hellish. We had to make our own tasks, 20-50h manual investment per task.
We considered adapting PostTrainBench but v1.2 still suffers from extreme reward hacking.
There is a huge amount we learnt about task creation. We iterated a bunch and had to throw away >75% of our tasks. We intentionally did not publish most of our learnings around making good tasks. It’s unclear that the benefits would outweigh the dual-use risks.
They also saw less reward hacking.
Meanwhile, Professor Dylan Hadfield-Menell said it measures metric optimization more than research taste. It’s been discussed for a while that doing well at tasks like inference or kernel optimization, or the NanoGPT speedrun, doesn’t really count as research taste.