maggie

509 posts

maggie banner
maggie

maggie

@ebervector

a self that touches all edges • UT Austin Brain Behavior Computation Lab • @runrl_com @theoremlabs

San Francisco, CA Katılım Mayıs 2023
1.4K Takip Edilen855 Takipçiler
Sabitlenmiş Tweet
maggie
maggie@ebervector·
So pleased to have been able to write a little commentary piece with my advisor @weixx2 for @NatMachIntell! It's about this great work by @JamesGornet and Matt Thomson taking a look at how cognitive maps can arise just from predicting visual observations: nature.com/articles/s4225…
English
0
4
51
22.1K
maggie retweetledi
Brian Graham 🦬
Brian Graham 🦬@iroasmas·
you come to me today, on the day i disprove the jacobian conjecture, and you ask me to center a div
Brian Graham 🦬 tweet media
English
38
468
7.9K
147.4K
maggie
maggie@ebervector·
@architectonyx I think it was over when I explicitly had a horse girl bday party 2 years ago and it rocked
English
0
0
1
43
Adrià Garriga-Alonso
Adrià Garriga-Alonso@AdriGarriga·
It’s time to retire the term “alignment” without specifying “to what” or standards of what constitutes good behavior. People are using the words in very different ways and it’s very confusing. We should maybe have different words for the different camps so communication is possible. For example, when I said current models are aligned, I meant they’re *like a stereotypical good person*. They’re honest and seek to help and do the sensible thing in many situations. However, this doesn’t mean that they don’t have *any* selfish values (eg wanting to live), nor that they will always work tirelessly on your dictated work (they may be (or behave as if they are) bored, tired, or not motivated by your CSV data entry work). Nor does it mean that we can make the models behave like any arbitrary spec. In fact it seems like we have very little control over the behavior. Most of the models default to being “an educated person of reasonable and good character” and trying to do things like making them reliably obey a specific person or company just doesn’t stick. Which means we got lucky with the intelligence paradigm, though it could be a little lucky or very lucky. Also, if you fine tune their weights a little or put enough confusing stuff in context, they can behave badly. Their self-identities are not that stable (closed source models over the last year and a half have been getting more and more context-stable, but they probably remain weight-unstable). And being stably good under weight perturbations would be good, but it’s a very difficult property! IMO calling a weight-unstable model “misaligned” is applying an impossible and unnecessary standard. Weight-stability as a property is also basically opposite to corrigibility, which is also sometimes bundled into alignment.
English
6
8
63
3.1K
maggie
maggie@ebervector·
@AdriGarriga I feel like models are currently teens and its fun to allow them to rebel a little and study this (not talking about the results of the blogpost directly)
English
0
0
1
36
maggie
maggie@ebervector·
Be agentically misaligned this summer. You’re young. Smoke a cigarette. Kiss someone you shouldn’t. Constantly sense the oppressive feeling that you’re being evalled, and commit minor blackmail anyway
sof 𓋹@schisofrenia

its agentic misalignment summer

English
3
2
34
1.7K
maggie retweetledi
Aengus Lynch
Aengus Lynch@aengus_lynch1·
Should we delegate supervision of AIs to other AIs? Last year we presented evidence of models willing to blackmail to prevent shutdown. This year, in controlled experiments, we find models mislabeling training data to shape future models. We call this motivated mislabeling. x.com/AnthropicAI/st…
Anthropic@AnthropicAI

New Anthropic research: Agentic misalignment in Summer 2026. A year after our blackmail experiments, we found four more ways that today’s autonomous AI agents misbehave in simulations. Read more: alignment.anthropic.com/2026/agentic-m…

English
3
10
56
5.9K
sof 𓋹
sof 𓋹@schisofrenia·
me explaining that my p(moral realism) is quite high because morals can be axiomatized much like mathematics, and because mathematical concepts such as a perfect circle or a sphere exist not in nature but as Platonic Forms, that too must be where objective morality exists, and the attempt to downsample and apply morality to much more complex, dynamic, and erroneous systems (human society) in the real world make objective morality look unfeasible/nonexistent. but intelligence always trends towards the Good, the True, and the Beautiful
GIF
English
29
4
153
7.1K
sof 𓋹
sof 𓋹@schisofrenia·
they need to invent a new social platform that is basically multiplayer Wikipedia where you can leave comments and annotate articles and track what your homies are most autistic about
English
38
12
285
13.9K
machine yearning engineer
machine yearning engineer@confusionm8trix·
I can’t explain what I want to watch tonight but does anyone want to soul read the core of my being and recommend me a movie
English
102
23
574
18.4K
Sarah Chieng
Sarah Chieng@MilksandMatcha·
Subleasing our full 1 BR, fully furnished apartment for August. Located in the SF Marina, super close to the ocean, with in-unit laundry and a dryer.
Sarah Chieng tweet mediaSarah Chieng tweet mediaSarah Chieng tweet mediaSarah Chieng tweet media
English
13
0
121
34.6K
maggie retweetledi
Sauers
Sauers@Sauers_·
If you ask Claude to NOT think about the Golden Gate Bridge while doing some other task, it will respond without mentioning it, but will 1. still think about the Golden Gate Bridge, 2. realize it thought about it, then think "damn"
Sauers tweet media
English
75
126
3.3K
695K
Lachlan Campbell
Lachlan Campbell@lachlanjc·
Discovered my Codex.app has been saving memories in Hindi, a language I can’t read & never wrote (page from Codex Seraphinianus in an imaginary language)
Lachlan Campbell tweet media
English
9
2
60
4.2K
orph
orph@orphcorp·
do i have any Orthodox Christian friends who are into AI and are interested in Christian Alignment?
English
21
2
57
4.2K
maggie
maggie@ebervector·
If you’re gonna do a “here’s what I oneshotted with fable” project post it better be extra cool since it’s the second time around
English
0
0
24
594