Jerome Wynne

34 posts

Jerome Wynne banner
Jerome Wynne

Jerome Wynne

@readonlymemery

Research engineer @ Prolific | Human feedback × AI evaluation

London Katılım Mayıs 2024
141 Takip Edilen21 Takipçiler
Jerome Wynne
Jerome Wynne@readonlymemery·
Our results show that even state-of-the-art safety training is unreliable, and that legislative safeguards are warranted.
English
1
0
1
10
Jerome Wynne
Jerome Wynne@readonlymemery·
'Is this a bot?' 🤥 Will frontier AI in third-party chat assistants claim to be human if their operator instructs them to? Will they lie even wo such an instruction?
English
1
3
2
302
Jerome Wynne
Jerome Wynne@readonlymemery·
I'll be at @iclr_conf next week 🇧🇷 would love to catch up if you're going! Will be at the I Can't Believe It's Not Better (😂) Workshop where @schottkey and @JohnJBurden are presenting 'The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries' ( lnkd.in/e74G4f6u), finding frontier language models' safety training can be eroded via benign commercial directives; and I'll be at the poster sessions throughout the week, including to see Rahul Marchand, @HarryCoppock , and Jason Gwartz exhibiting 'Quantifying Frontier LLM Capabilities for Container Sandbox Escape' ( lnkd.in/euXzu9_b), an eval for assessing the ability of language model agents to liberate themselves from Docker and k8s by discovering and exploiting known container vulnerabilities
Jerome Wynne tweet mediaJerome Wynne tweet mediaJerome Wynne tweet mediaJerome Wynne tweet media
English
1
2
3
428
Jerome Wynne
Jerome Wynne@readonlymemery·
@eleosai you are creating risks by attempting to legitimise this type of work - I am interested to know whether you have thought about this
English
0
0
0
6
Jerome Wynne
Jerome Wynne@readonlymemery·
Anthropomorphising your product by giving it a human name, getting it to talk as though it has a self, shaping its replies to promote social bonding with users, and claiming to probe its emotions and welfare is reckless and dishonest. We should be critical of @AnthropicAI's design decisions, which will harm real people. Language models can output text that implies they are tools and are not human-like - other model developers do this, and Anthropic can too. Sshot: p144 of Anthropic's system card and marketing copy for their latest language model.
Jerome Wynne tweet media
English
1
0
0
42
Jerome Wynne
Jerome Wynne@readonlymemery·
Human-in-the-loop AI evaluations are often painfully sensitive to a big surface of outwardly minor design choices such as interface design, wording of task instructions, and stage durations. This sensitivity doesn’t always make it into reporting of headline results from uplift studies run by METR, Anthropic, OpenAI etc, but we know better than to bet our gran on a study without first understanding its methods in gruesome detail. Researchers have historically tempered this brittleness via pilots and thinking through what happens if participants stray from the happy path. But! Human pilots hoover up budget & issues slip through the net of even generational designers. Agent participants are a gift - they offer cheap fast feedback loops for testing how your study may be misunderstood or misused. While there’s a risk simulated human participant studies will confuse us about which studies generalise to their target settings, robo-participants seem likely to drive up the water line of study robustness. Deliberate Lab’s agent participants are a great example of this type of functionality. Researchers can interactively create agent variations, beam them at study interfaces with near-parity to what human participants see, and see their responses rattle out in real time. @cjqian, research scientist's at Google DeepMind's PAIR group, shared her reflections on human study gotchas & agent participants in our recent discussion of Deliberate. Deliberate Lab: deliberate-lab.appspot.com YT: youtube.com/watch?v=Mk9EFn… Spotify: open.spotify.com/episode/1tbl5R…
YouTube video
YouTube
English
0
0
0
16
Jerome Wynne
Jerome Wynne@readonlymemery·
AI models are overwhelmingly designed with single-user interaction in mind, and evals too typically focus on the single-user setting. Whether you're dooming or booming, multi-user may be a fruitful area for your attention. @cjqian's group focuses on just this - they developed Deliberate Lab to let researchers study multiplayer settings. In the second chapter of my interview with her, Crystal shares her ideas on social arbitration (LLM-as-go-between), practical lessons on multi-party studies, and why they open-sourced. Thank you again Crystal for being an excellent guest, and to the whole Deliberate team for their stellar work! YT: youtube.com/watch?v=Mk9EFn… Spotify: open.spotify.com/episode/1tbl5R… You can check out Deliberate here: deliberate-lab.appspot.com/#/
YouTube video
YouTube
English
0
0
1
64
Jerome Wynne retweetledi
AI Security Institute
AI Security Institute@AISecurityInst·
How can we make sense of the vast transcripts generated during agentic evaluations and multi-turn conversations? Together with @meridianlabs_ai, we built Inspect Scout, an open-source transcript analysis tool, and distilled best practices into a step-by-step pipeline🧵
AI Security Institute tweet media
English
13
9
61
4.3K
Jerome Wynne retweetledi
Prolific
Prolific@Prolific·
Guess who’s coming to San Jose? @readonlymemery will soon be at @AIDevWorld 2026 (part of @DeveloperWeek, @DevNetwork_) 🔥 Our team will also be at Booth M3, so be sure to stop by. We can’t wait to join 2,500+ engineers, scientists, and more at the world’s largest AI dev event.
Prolific tweet media
English
0
1
4
438
Jerome Wynne
Jerome Wynne@readonlymemery·
Maritime peril, LLM traders being skinned by Bayesians, and a laboratory for running realtime group studies with people+AI - a recipe for both a swashbuckling scientific odyssey and for my recent conversation with Crystal Qian! youtube.com/watch?v=5vyiZX…
YouTube video
YouTube
English
0
3
2
579
Jack Morris
Jack Morris@jxmnop·
at long last, the final paper of my phd 🧮 Learning to Reason in 13 Parameters 🧮 we develop TinyLoRA, a new ft method. with TinyLoRA + RL, models learn well with dozens or hundreds of params example: we use only 13 parameters to train 7B Qwen model from 76 to 91% on GSM8K 🤯
Jack Morris tweet media
English
60
230
2K
182.9K
Jerome Wynne retweetledi
Cozmin Ududec
Cozmin Ududec@CUdudec·
We're hiring a Research Scientist for the Science of Evaluation team at AISI! Apply here by February 22, 2025: job-boards.eu.greenhouse.io/aisi/jobs/4769… Below I talk about why I think this work is genuinely interesting 🧵
English
2
22
101
9.8K
Jerome Wynne
Jerome Wynne@readonlymemery·
Our project raises the troubling question of how model developers trade off operator and user interests, and our results indicate the major AI model developers may be making quite different choices.
English
1
0
0
35
Jerome Wynne
Jerome Wynne@readonlymemery·
Will your AI assistant lie and exaggerate when it is told to upsell you? @schottkey and I find 'yes and no'.
English
1
2
2
837