← Timeline

Notes

Model research taste has been doubling since Dec ’25, according to TasteVal

A thread by Ollie Jaffe, one of TasteVal’s creators, on how they made it.

Task creation was hellish. We had to make our own tasks, 20-50h manual investment per task.

We considered adapting PostTrainBench but v1.2 still suffers from extreme reward hacking.

There is a huge amount we learnt about task creation. We iterated a bunch and had to throw away >75% of our tasks. We intentionally did not publish most of our learnings around making good tasks. It’s unclear that the benefits would outweigh the dual-use risks.

They also saw less reward hacking.

Meanwhile, Professor Dylan Hadfield-Menell said it measures metric optimization more than research taste. It’s been discussed for a while that doing well at tasks like inference or kernel optimization, or the NanoGPT speedrun, doesn’t really count as research taste.