
Cozmin Ududec
455 posts

Cozmin Ududec
@CUdudec
@AISecurityInst Science of Evaluation lead. Ex quantum foundationalist.




Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from our Science of Evaluation team shows why that matters. 🧵

When Role-Playing, Do Models Believe What They Say? (w/ @DavidDAfrica and @realmeatyhuman) LLMs can say “The Earth revolves around the Sun” and then, when roleplaying as an ancient Greek historian, assert the opposite. What changes inside the model when it acts like this? Does it just say things, or does it start to believe the role? 🧵

Excited to announce my first preprint from @AISecurityInst ! 🎉 Here, we asked: how much does an AI agent's performance on a benchmark depend on how much compute we give it to do the task? Paper link in thread 👇🧵(1/7)



