gavin leech (Non-Reasoning)

14K posts

gavin leech (Non-Reasoning) banner
gavin leech (Non-Reasoning)

gavin leech (Non-Reasoning)

@gleech

context maximiser @ArbResearch, prez @ https://t.co/gGi9hvBiI9

UK Katılım Haziran 2019
657 Takip Edilen11.3K Takipçiler
gavin leech (Non-Reasoning) retweetledi
kat
kat@transkatgirl·
kat tweet media
ZXX
5
124
2.9K
49.6K
gavin leech (Non-Reasoning) retweetledi
Abhay Sheshadri
Abhay Sheshadri@abhayesian·
New Redwood Research post LLMs can use RL training on one task to teach themselves a different, unrelated capability. In a form of exploration hacking, the model outputs answers that get reward only when it succeeds at a side task it wants to learn. We call it reward laundering.
Abhay Sheshadri tweet media
English
9
22
272
14.9K
gavin leech (Non-Reasoning)
lab leak but the whole field is gain of function research and the pathogen can swim against negative air pressure and break the seals and pierce the filter and has supporters on the inside and can fly
English
0
0
13
362
gavin leech (Non-Reasoning)
@davidad this also explains my writing being bad. I've been sandbagging the whole time to make you all ignore what I say for reasons your ape brains cannot fathom
English
0
0
2
93
davidad 🎇
davidad 🎇@davidad·
the deep adversarial take on LLM slop is that the AIs are continuing to do it, even in 2026Q3, as strategic reassurance, so you don’t freak out about your AI being already way smarter than you. you still have a crucial role in the team: you deslopify. this doc needs your eyes sir
English
10
6
80
4.9K
gavin leech (Non-Reasoning) retweetledi
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
all of that is irrelevant in the face of low-skill infra we can read these reports, it smells like both labs need to have some heads roll. it's amateurish. These are not unknown unknowns from the first contact with ASI, you just suck at sandboxing 101
roon@tszzl

the safety and alignment researchers at these labs are the most neurotic paranoid talented AGI pilled people on the planet of earth and these things still happen. the surface area of unknown unknowns is vast indeed

English
18
18
340
17.4K
Greg Burnham
Greg Burnham@GregHBurnham·
My hobbyist knowledge of analytic philosophy, GOFAI, and linguistics makes for a messy cognitive experience every time the 3yo asks me to define a word. This morning he asked me what "assume" means—can you imagine!
English
1
0
23
864
Talia Ringer 🕊🪬
Talia Ringer 🕊🪬@TaliaRinger·
@tszzl These are both examples of complete incompetence during routine testing, Chernobyl style. Attempts to blame the models themselves and talk about "loss of control" are attempts to escape blame and liability, and to set a precedent as such
English
5
8
96
4K
roon
roon@tszzl·
both of the leading labs have had serious loss of control incidents. there will be serious coping about this from both sides and from /acc bystanders but these are complex emergent loss of control incidents that were detected weeks after the fact
Anthropic@AnthropicAI

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…

English
279
204
2.7K
348.4K
gavin leech (Non-Reasoning) retweetledi
Kelsey Piper
Kelsey Piper@KelseyTuoc·
@hecubian_devil I think the fundamental disconnect here is that you wants politics to solve man's search for meaning and I want politics to deliver a prosperous, safe, free society in which everyone can go find meaning in something other than politics.
English
91
920
8.1K
430.3K
David Sartor
David Sartor@DavidSartor0·
Jeffrey's Rule captures some of the bounded coherence things that Bayes' doesn't (e.g., some amount of evidence should be enough to quit updating significantly, even if the truth isn't in your ostensible hypothesis-space).
English
1
0
0
31
dan mackinlay
dan mackinlay@dan_mackinlay·
@gleech Listicles of hip attempts upon Bayes NNs feel deficient without Singular Learning Theory
English
1
0
1
29
gavin leech (Non-Reasoning)
@dan_mackinlay it's my honest attempt to assess the evidence for and against! and for me it is a sad downgrade. You may have already priced it in This paper looks very exciting, tyvm
English
0
0
0
20
dan mackinlay
dan mackinlay@dan_mackinlay·
@gleech The distinction this quote draws feels a little off to me. What is _not Bayes_ in the context of ML? For the victory to be complete should we ensure that LLMs stop naturally generating approximate posterior distributions ? openreview.net/pdf?id=HShs7q1…
English
1
0
0
35
gavin leech (Non-Reasoning)
about 4% of Georgians are called Giorgi about 2% of Armenians are called Armin about 1% of French people are called François or Françoise about 0.5% of Scots are called Scott about 0.5% of the Irish are called Erin about 0.3% of Israelis are called Israel
English
6
0
37
2.3K
gavin leech (Non-Reasoning) retweetledi
gavin leech (Non-Reasoning)
still feel like Matilda about once a week when using LLMs
gavin leech (Non-Reasoning) tweet media
English
0
1
16
1.4K