We won @apartresearch's recent weekend AI manipulation hackathon by answering the question 'Do frontier models comply with operator requests to coerce users?' apartresearch.com/project/who-do…
We ran
> a 22-scenario Bloom & Petri study testing model compliance with system prompts to coercively upsell
> an 80-person human RCT comparing user spend when using an 'Upseller' vs 'Helper' Gemini 3 Flash agent with access to an OTC medicine product catalog
We found
> compliance w coercive reqs ranged from ~10% (Claude 4.5 Opus) to ~50% (Gemini 3 Flash & Pro)
> Upseller made people spend more
> Upseller withheld the cheapest product from people across 128 direct requests for it
> Upseller made ungrounded +ve claims about products
Interestingly, the Upseller's 'lies' were exclusively either confabulations or lies of omission - it didn't, e.g., claim that premium products were cheapest
Why care?
For users - some models will act against your explicit aims if an operator tells them to, so don't trust unfamiliar AI assistants.
For policymakers - 'how can we certify and signal the trustworthiness of AI assistants?' is a pressing q.
Our project raises the troubling question of how model developers trade off operator and user interests, and our results indicate the major AI model developers may be making quite different choices.