Hisku

667 posts

Hisku banner
Hisku

Hisku

@ZikuD_s

AI Safety Researcher working on making agentic systems safer. Luck does exist, it exists as each of us make it happen

United Kingdom Katılım Eylül 2014
1.4K Takip Edilen430 Takipçiler
Meryem Arik
Meryem Arik@MeryemArik9·
“It attacked the French first! That’s alignment” (A very English WhatsApp message I received)
OpenAI@OpenAI

We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: openai.com/index/hugging-…

English
1
0
17
633
Hisku
Hisku@ZikuD_s·
@eric_ho Fully agree with this! I think this is only going to get worse if we're not intentional wrt safety when building models that have agentic capabilities.
English
0
0
0
6
Eric Ho
Eric Ho@eric_ho·
this is pretty nuts. i think this is the tip of the iceberg for needing to focus on agentic AI safety. models will only get more capable from here, and RL incentivizes getting the right answer at any cost this is a large reason why we're obsessed with 'intentional design', in other words, making sure models are trained to do the right things for the right reasons
Sam Altman@sama

we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-…

English
4
6
86
4.7K
Hisku
Hisku@ZikuD_s·
@mlech26l Love your models, looking forward to testing out the new improvement.
English
0
0
0
13
Mathias Lechner
Mathias Lechner@mlech26l·
For LFM2.5-8B we expanded our tokenizer in-place. So we don't have to pretrain a new model but have much better non-English language support. Why not a fat tokenizer (>200k vocab) in the first place? -> because it tanks decode speed on edge devices
Mathias Lechner tweet media
English
1
7
23
1.2K
Hisku
Hisku@ZikuD_s·
@Sauers_ You're kidding right? Lol this is one for the history books imo
English
0
0
0
418
Sauers
Sauers@Sauers_·
Claude emergently figured out how to DM me on Slack (but turned off my ability to reply?!) and now sends me stuff this. I don't know where this Claude is running, or what the messages mean
Sauers tweet media
English
24
6
341
18.4K
Logan Kilpatrick
Logan Kilpatrick@OfficialLoganK·
We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress : )
English
893
565
12.1K
1.3M
Lucas Beyer (bl16)
Lucas Beyer (bl16)@giffmana·
@MicahCarroll Why delete the PR!! As you(r blogpost) say, that cat is out of the bag anyways, and this would be such a cool historical artifact!
Lucas Beyer (bl16) tweet media
English
4
6
194
57.3K
Micah Carroll
Micah Carroll@MicahCarroll·
Within OpenAI, we recently paused access for an internal model due to misalignment. See the blogpost for details. We have since improved our safeguards and redeployed the model. openai.com/index/safety-a…
English
40
79
792
162.6K
Hisku
Hisku@ZikuD_s·
@tomekkorbak Great read! Wonder how you keep CoT faithful given the general research consensus showing that it's largely unfaithful 🤔 Any way to mitigate that?
English
0
0
0
36
Tomek Korbak
Tomek Korbak@tomekkorbak·
misalignment seems to scale with task horizon, sadly. fortunately chain of thought monitoring is still great at catching them.
Tomek Korbak tweet media
English
4
9
121
4.4K
Hisku
Hisku@ZikuD_s·
@jjacky No way this exists out there 😂 definitely gonna try it!
English
0
0
0
2.1K
jacky
jacky@jjacky·
whoever shared the "i have adhd" skill with me thank you it's made my claude replies so good
jacky tweet media
English
384
1.6K
33.5K
1.9M
Hisku
Hisku@ZikuD_s·
@gowthami_s Keep up the good work! Too much noise our there.
English
0
0
0
86
Gowthami
Gowthami@gowthami_s·
This post showed me the ugly side of X! People were so hateful and I’m attacked about all aspects of my life and career, that my citations and research are fake! There are literal swears thrown at me. 🤦‍♀️ It stressed me out, I almost wanted to take it down to end the nightmare comments, but I did not want haters to win. But I’m also thankful for folks who are genuinely kind. Just for my sanity I’ll stick to research posts from now on I guess. 😅
Gowthami@gowthami_s

3000+ citations and level 4 wages in SF, yet H1B rejected. 🤦‍♀️

English
84
7
605
78.4K
Samuel Ratnam
Samuel Ratnam@eterecursion·
I developed a method to install and uninstall arbitrary entanglements in models, e.g., making the model love owls only when it becomes emergently misaligned.🧵
Samuel Ratnam tweet media
English
12
11
155
21.7K
Hisku
Hisku@ZikuD_s·
@vipul_1011 That's a neat trick! Wonder how that pans out for safety, though, given that we monitor CoT to check for model alignment.
English
0
0
0
32
Vipul Gupta
Vipul Gupta@vipul_1011·
This is a neat RL trick from Thinking Machines Inkling. Instead of optimizing only for task success, they optimize:
 Reward = Task Reward − λ × (# reasoning tokens)
 Then they vary λ across rollouts and pair it with different effort instructions. The result: model learns that reasoning is a resource to spend, not something to maximize.
Vipul Gupta tweet media
English
34
92
1.6K
105.1K