
I'm excited to share our new paper introducing EvalDetectBench for measuring evaluation awareness in frontier LLMs. We also found that the answer depends a lot on which questions you ask and which model generated the transcripts being judged
Open and Inspect-compatible. 🧵👇
Ram Bharadwaj@arbdwj
Can we reliably measure whether frontier models know they are being evaluated? 🔬 New paper from me, @xinningli6, @levanto_0,@AlexandraSouly, @_robertkirk: EvalDetectBench, a benchmark for measuring evaluation awareness in frontier LLMs 🧵
English