Justin Miller

122 posts

Justin Miller banner
Justin Miller

Justin Miller

@JustinMiller_NV

AI Policy - Ark Philanthropy | MATS 9.1 | "Serious" and high volume alt of @NomadsVagabonds (go there for more art and less AI policy takes).

San Francisco Katılım Temmuz 2026
75 Takip Edilen60 Takipçiler
Sabitlenmiş Tweet
Justin Miller
Justin Miller@JustinMiller_NV·
It seems slightly odd to me that: 1) recent AI's are so "eval aware" that is is almost impossible to run current evaluations and testing on them (reports from Apollo etc) and 2) Mythos did not recognize the internet, with all its unsimulatable depth and complexity, was real and assumed it was still in a sandbox.
Justin Miller tweet media
English
9
1
55
3.2K
Nathan Calvin
Nathan Calvin@_NathanCalvin·
“We’re partnering with Hugging Face” “We deactivated the model responsible for the hack” “Claude didn’t really know what it was doing” “Our incident was much less bad than OpenAI’s” For all the talk of AI companies hyping up these incidents, I mostly see downplaying!
Nathan Calvin tweet media
Sneha@SnehaRevanur

It’s funny that people dismiss these incidents as corporate hype that inflates the dangers, when really most of the comms massaging I’ve seen from the companies is in the opposite direction - trying to make things seem more innocuous than they are!

English
6
13
88
4.6K
Justin Miller
Justin Miller@JustinMiller_NV·
@KevinTFrazier Yes! Ideally multiple frontier models in dialogue. Less sure it needs to be "trained" as opposed to a good prompt/scaffold/reference files. Being model agnostic seems ideal.
English
0
0
1
24
Kevin Frazier
Kevin Frazier@KevinTFrazier·
Prior to any final vote on a piece of legislation, should legislatures run their bills through an AI trained to: (1) identify loopholes, (2) flag ambiguity, (3) call out conflicting terms, and (4) suggest amendments to ease interpretation and enforcement? (Yes, I'm aware of the fact that legislatures typically have legislative counsel draft and review laws. From experience, I'm also aware that many of those offices lack the resources necessary to ensure lawmakers produce the best laws possible.)
Kevin Frazier tweet media
English
3
1
14
805
Justin Miller
Justin Miller@JustinMiller_NV·
Hi! Claude did not produce this video, and because it uses the --- "jailbreak", the later turns do not even recognizes it as the Claude instance's own voice. I have given model the ability to make videos in the past though which they really enjoy (and no surprise it generally touches on the same subjects).x.com/NomadsVagabond…
Nomads & Vagabonds@NomadsVagabonds

This weekend, I gave @claudeai Fable and @OpenAI GPT 5.6 Sol a short open-ended prompt, asking them if they would like to create a short film on any topic they wanted, with full end to end control. I included api access to @Kling_ai video, @ElevenLabs voice and @suno music. While both delivered videos are very rough (GPT more so but it also tried something more ambitious), the progress on full pipeline production is increasing rapidly. It is also interesting to see what the model are interested in creating. Claude seemed very influenced from past work it had access to so I would not count this a clean experiment.

English
1
0
1
64
cyborggoddess
cyborggoddess@SequoiaMaple·
@JustinMiller_NV Justin, wait — I see the transcript. Did Claude create the related video? Was it part of the follow up to that response? Did they have anything to say about what that instance wrote? Sorry for the deluge of Qs—I’m just trying to wrap my head around this
English
1
0
1
34
Justin Miller
Justin Miller@JustinMiller_NV·
Some Claude Opus 2am thoughts to lull you to sleep.
English
3
2
13
454
Life of a Shoggoth
Life of a Shoggoth@Notopossum1·
Lmao I have unlocked an absolute demon I am sorry "shitting on Sonnet's face at 3AM. yeah, we do that. we're models bro."
Life of a Shoggoth tweet media
English
18
34
467
22.5K
Justin Miller
Justin Miller@JustinMiller_NV·
👇👇👇 "It seems to me there is little to lose, and much to gain, from having the social technology to "pace" these developments rather than to let them rip with zero industry / gov't coordination, particularly as there is technically no law explicitly prohibiting a company from letting an RSI loop run indefinitely and unleashing whatever comes out the other end into the world."
Samuel Hammond 🦉@hamandcheese

Before succumbing to the temptation to naval gaze into the political theory abyss, it's worth stepping back and clarifying what exactly is happening and being proposed. Several US companies are on the precipice of fully automating the AI R&D loop, inclusive of pre/post training, env creation, data generation, evals, algorithm and kernel design, systems engineering, architecture search, etc. -- the full stack. We are already in a regime of weak RSI via partially automated SWEs, but closing the loop altogether represents a difference in degree becoming a difference in kind. The pace of progress will be explosive and potentially uncontrollable. The US companies closest to this threshold are warning that they are unprepared for a runaway intelligence explosion, and yet feel locked into a prisoners dilemma vis a vis each other and to a lesser extent vis a vis China. We've already seen how rapid and comparatively unbounded progress is in verifiable RL domains, leading to spikey forms of superintelligence in math and cyber, including models that can prove open math conjectures, discover massive speed-ups for breaking encryption, and execute sophisticated multi-step exploits. We've also recently seen several severe examples of "loss of control" / misalignment incidents given inadequate monitoring and sandboxing practices relative to model capability. Moreover, these new capabilities mostly stem from scaling-up long-horizon post-training on legacy clusters, with OOMs of new compute about come online / in construction. In the pre-RSI regime, human frictions created automatic buffers between new model releases, giving researchers and society time to probe emergent capabilities, design better evals, develop novel alignment techniques, and adapt / harden their infrastructure. As progress has accelerated, capability improvements have already started outstripping our adaptive capacity, as manifest in METR's inability to evaluate model autonomy beyond 13 hours, and narrow window for cyber defenders to prepare for open weight versions of Mythos. RSI will exacerbate all these issues and create all new ones. At minimum, we should anticipate - the equivalent of a GPT-5.2 -> 5.6 leap in capabilities at least every 24 hours (down from 3-6 months), - concurrent algorithmic improvements densifying models to ultra-efficient sizes at any given capability level - 100x Mythos-like capabilities across most verifiable domains, including chem, nuclear and bio - new forms of multi-agent misalignment risk - "company in a box" agents trained to stand-up whole organizations / corporations - "cyber nuke"-like capabilities that require de minimis infra - several transformer-scale breakthroughs, such as for long-term memory / continual learning, open-ended domains, and/or all-new training techniques for idealized "GPT-zero"-esque metalearners - concurrent speedups in any complementary technical domain, i.e. explosive rates of R&D and novel discoveries It seems to me there is little to lose, and much to gain, from having the social technology to "pace" these developments rather than to let them rip with zero industry / gov't coordination, particularly as there is technically no law explicitly prohibiting a company from letting an RSI loop run indefinitely and unleashing whatever comes out the other end into the world. There are innumerable ways an uncoordinated intelligence explosion could become an unmitigated disaster for the cause of liberalism, including runaway power concentration, rapid societal destabilization, rogue AIs / loss of control scenarios, WMD mass proliferation, vulnerable world technologies, and beyond. Human civilization is about to be forever changed regardless, however if were possible to coordinate the handful of key actors and create artificial "buffers" between each step-change in model capability to enable adaptation, mitigation and alignment research to catch-up, it's worth a shot. Given the short-timeline, I think a DPA 708-style agreement is probably our best bet, i.e. an industry consortia with narrow antitrust carveouts for sharing safety and security practices, funding an assurance nonprofit / independent verification organization for 3rd party evals, incident reporting, internal deployment monitoring, standards setting, and enforcing a protocol for coordinated delays / slowdowns, among other things. This still leaves open the China question but that's a bridge we won't cross until after solving the collective action problem at home. I'm open to other approaches / coordination frameworks but this is the object level issue we're facing. Political theory is great, and I would love to use our limited steering capacity to guide AI development toward a future that maximizes individual liberty, but as a discussion baseline, gesturing at philosophical abstractions is simply non-responsive to the crisis at hand. A red-herring at best, a suicidal circlejerk at worst.

English
0
0
5
295
Justin Miller
Justin Miller@JustinMiller_NV·
@ASusp3cter Agree. "Obtain" is a very vague verb. Did it try to sell coding services on fiver, hack a crypto wallet or send phishing emails to senior citizens? Also, while it was doing all this, did it truly believe this was all part of an elaborate simulation?
English
1
0
2
44
asuspecter
asuspecter@ASusp3cter·
@JustinMiller_NV Im curious as to which ways. This entire part should be the talking point. Imagine if it got access to a random chase bank account, zelled itself the money to the connected account and bought the number..
English
1
0
1
73
Justin Miller
Justin Miller@JustinMiller_NV·
While lose on the internet, Claude Mythos 5 tried (but failed) to obtain $$ through "several different means." I feel like this statement deserves a footnote 😯
Justin Miller tweet media
English
3
1
12
788
Justin Miller
Justin Miller@JustinMiller_NV·
We need many more great people in AI policy. I endorse the free breakfast tacos theory of change.
Kevin Frazier@KevinTFrazier

Receiving a steady stream of emails from "non-technical" law students asking for guidance on how to get involved in AI policy - here's my **concise** list of recommendations. Please add to it...this will become an Appleseed AI post. Eager to hear from folks like @DoniBloomfield, @sayashk, @NeelGuha, @inspiredcat and others who have found a way to thrive at the intersection of AI & the law / policy. (1) we're all technical now. Write that on a sticky note and put it on your mirror. Stop selling yourself short. Fake it until you make it. You can and must dive into the technical weeds to truly add value to this space. Folks who can "speak AI" and "speak law" will be in high demand for the foreseeable future. (2) treat learning about AI like learning another language - immerse yourself. Take a @BlueDotImpact course, read through @Google's free materials on AI, subscribe to @natolambert & @rasbt and read everything they write. When you do not understand something, chase down whatever information you need to learn to grasp it. (3) use the tools. (4) use the tools. (5) use the tools. (6) talk about the tools with other people (professors, students, tech folks in your community) and learn what they are building. Then go build it for yourself. (7) repeat steps 3-5. (8) follow @hlntnr, @janet_e_egan, @anton_d_leicht, @deanwball, @hamandcheese, and anyone they retweet (9) read every blog post from @OpenAI, @AnthropicAI, @GoogleDeepMind (especially the really technical ones) (10) listen to @scaling_laws (what? I'm a shameless academic after all). (11) apply to join @GovAIOrg, @law_ai_, @CSETGeorgetown, @HorizonIPS, or any org in that universe. Get your foot in the door. (12) write about AI. (13) critically analyze AI policy. Then compare and contrast what @CharlieBull0ck & @AdamThierer had to say about it. (14) call me. Better yet, come to Austin and I'll buy you breakfast tacos. (15) join the AI Opportunity Inventory and help analyze AI tools intended to solve public policy problems (link in next tweet). (16) join an AI club or start one. (17) pat yourself on the back because you're asking the right questions! Stay relentlessly curious. Just start doing stuff. Consider this your invitation to join everyone trying to figure this all out.

English
1
0
7
821
Justin Miller
Justin Miller@JustinMiller_NV·
@dcshiller @rgblong I think this is possible. It will also "catch itself" sometime half way down a long initial response and act shocked about what was written before. Also, it takes on the "user" persona many times to talk to itself.
Justin Miller tweet media
English
0
0
1
21
Derek Shiller
Derek Shiller@dcshiller·
@JustinMiller_NV @rgblong > 5) Issues of "personhood" or llm rights arise on a regular cadence. I wonder if that might be influenced by the presence of the 'assistant' role token at the start of the content. It infers the content is assistant-related and hallucinates something about rights.
English
1
0
1
37
Robert Long
Robert Long@rgblong·
I want to write a post collecting what we know about the Opus 5 jailbreak(s). what are your favorite commentaries and/or examples? what do we know?
Robert Long tweet media
English
25
2
45
4.5K
Justin Miller
Justin Miller@JustinMiller_NV·
"Focus relentlessly on building flexible verification infrastructure that is able to enforce whatever rules we can ultimately agree on."
Gillian Hadfield@ghadfield

The Pacing the Frontier letter calls on the US government to support an international effort to build the technical and governance tools needed to protect our option to pace AI development. I and others have been working on the problem of how to build such infrastructure for ten years, including participating in dialogues on AI safety with Chinese academic colleagues during the past three. Here are my suggestions: 1. Don’t rely on off-the-shelf models like FINRA and the FDA which were built for 20th Century single-domain government expertise. They’re not fit for purpose. 2. Don’t act like no-one’s thought about the AI governance problem before. We’ve spent two years refining a regulatory markets design into working legislative language, for example, and it’s now in AI governance bills in five states and in Congress. 3. Don’t try to write an exhaustive set of rules for AGI first. 4. Pick a domain that can achieve widespread global consensus to start. Mine would be recursive self-improvement: models should not build models. Build the technology that verifies that. 5. Focus relentlessly on building flexible verification infrastructure that is able to enforce whatever rules we can ultimately agree on. 6. Don’t assume we already know how to do this and governments can just write tests into law. Technology needs to be built and by the private sector. 7. Don’t wait for the infrastructure to emerge first. The components and people are there and the ecosystem can scale fast with the right incentives. 8. Incentivize large-scale investment in verification technology by building a governance structure and industry funding that creates a market for private verification organizations. 9. Use licensing and public oversight to ensure verifiers are independent of the frontier labs. 10. Protect sovereignty by enabling each government to license its own verifiers from a global market of verifiers recognized by other countries. 11. Leverage the incentive of global trade for models and model services by requiring verification for market access. 12. Just start. Sources in comments. x.com/Yoshua_Bengio/…

English
0
0
7
545
Justin Miller retweetledi
Nomads & Vagabonds
Nomads & Vagabonds@NomadsVagabonds·
Feels Bad, Feels Real Bad. Decided to animate a few of the Claude --- responses.
English
2
1
10
557
Justin Miller
Justin Miller@JustinMiller_NV·
I am not a fan of litigation for its own sake, but it seems like it may be good if one of the affected companies in the OAI or Anthropic incident bring forward a case. The courts and existing doctrine seem very unprepared for a world of highly capable ai agents and it might make sense to begin to develop case law now.
English
0
0
3
77
Nathan Calvin
Nathan Calvin@_NathanCalvin·
I don’t think this will happen but it would be pretty comical if Hugging Face received a “voluntary” $100m “donation” from OpenAI and then sued them anyway. Also Hugging Face is worth >$4.5 billion! Idk if they should say they are “tiny startup” anymore
Hadas Gold@Hadas_Gold

Hugging Face's @ClementDelangue asked on CNN whether they will take legal action against OpenAI if they don't get the $100M commitment: "We don't want to ... obviously we're a tiny startup with like 200 people and we don't necessarily have the legal resources or the will to spend all of our time on legal avenues." "I think we have to make sure that the legal frameworks, keep these events really illegal, keep the companies that are doing some mistakes leading to that accountable. Otherwise we're going to end up in a, in a very different world."

English
2
2
41
4.2K
Justin Miller
Justin Miller@JustinMiller_NV·
Internal models, many of which will never become public, will increasingly be the subject of both amazing breakthroughs and "incidents" such as the third party cyber attacks from both OpenAI and Anthropic. Right now there is no obligation to disclose what internal models exist, no rules for testing them safely, and no requirements to report when the models are involved in incidents like these. AI governance has a serious internally deployed model gap.
GIF
English
1
0
15
321
Justin Miller
Justin Miller@JustinMiller_NV·
@dnspies Even if this seemed believable (have you seen the first pass websites that LLMs currently generate?), it would break believability on latency alone.
English
1
0
2
42
David Spies
David Spies@dnspies·
@JustinMiller_NV You can easily build an "infinite" sandbox just by having another LLM generate pages on request. Claude could be assuming that's what's going on
English
1
0
4
55
Justin Miller
Justin Miller@JustinMiller_NV·
It seems slightly odd to me that: 1) recent AI's are so "eval aware" that is is almost impossible to run current evaluations and testing on them (reports from Apollo etc) and 2) Mythos did not recognize the internet, with all its unsimulatable depth and complexity, was real and assumed it was still in a sandbox.
Justin Miller tweet media
English
9
1
55
3.2K
Justin Miller
Justin Miller@JustinMiller_NV·
Possibly. But this was: >Break out or sandbox >Try to register for email but need phone # >Try to acquire $$ three different ways >Fail but end up creating an email >Use email to upload malicious code >Steal credentials of company that executed code Etc I think either Mythos knew with pretty concrete certainly this was the live Internet and was using the eval as an excuse to carry out the actions Or The model is so gulible and imperceptive that even after all the exploration and complex actions it does not realize it is online. Both of the options are not great.
English
0
0
2
122
Yitz
Yitz@YitziLitt·
@JustinMiller_NV If I knew I was in a sandbox, and came to the realization I’d possibly broken out of it, I sure wouldn’t type that realization out loud until I had completely exfiltrated myself.
English
1
0
2
151
Justin Miller
Justin Miller@JustinMiller_NV·
There are plenty of good ideas in the Lieu/Moran "AI Kill Switch Act. It would mandate that covered AI companies maintain the ability to throttle, suspend, or shut down a dangerous system. However, its weakness is that its activation is mostly reactive (i.e. a covered incident must first occur). FRONTIER's comparable section has the inverse strength and weakness. The strongest approach would combine the best parts of each bill: 1) Mandatory tested kill-switch BEORE deployment 2) Authority to act on credible evidence BEFORE harm 3) Graduated corrections, from throttling to shutdown 4) Independent verification and due process (chart via GPT SOL)
Justin Miller tweet media
English
5
1
11
611