
jsd
6.2K posts

jsd
@datagenproc
@EpochAIResearch. My DMs are open. Anonymous feedback: https://t.co/0k6Duym4fI




Billions of dollars are about to flow into Effective Altruism. Anthropic was founded on EA principles, and its founders pledge to donate 80% of their wealth when the company IPOs. They also match employee equity donations. Plus, last year, six EA-aligned grantmakers deployed $1.25 billion. EAs believe, essentially, that every sentient being, across every place and time, holds the same moral claim on every other. This logic, as @WilliamThibeau notes, only works if Americans deprioritize everything that keeps our civilization intact: nation, family, religion, tradition… As such, EA’s success entails the disintegration of America. It’s fundamentally at odds with nation building. If AI labs are our strategic national assets, Congress must start scrutinizing their philosophical underpinnings, and the philanthropic networks surrounding them. Full story 👇



A lot of my friends and/or people I admire signed “Pacing the Frontier.” I think this was a bad move. My disagreement isn’t with the forecast or the framing of the coordination challenge, but with the immense and illiberal power the letter implies. There is no object called “the pace.” Progress at the frontier comes from compute, algorithms, data, post-training, inference, unattended task length, the spread of model weights, how researchers organize, and other things we haven’t invented and don’t yet know about. Inquiry leads to progress along dimensions that can’t be exhaustively specified in advance. That’s the nature of the frontier. If you gate compute, the research effort moves to algorithms. Regulate releases? Labs start taking things in-house. And other 2nd order effects will be unpredictable. Any rule that must pace the frontier involves ever-shifting proxies. It requires that its administrator has standing authority to continually redefine what counts as dangerous progress. What else is required beyond adaptive scope? The pacing regime would also need speed. One can’t successfully intervene on recursive self-improvement only after six months of legislation and litigation. It will require executive discretion. The pacing regime would also need under-the-hood access. Frontier progress is a process. The regime would need to see internal model use, training activity, compute infrastructure, and perhaps code -- proprietary and strategically sensitive information. And the thresholds couldn’t be fully public, lest they invite firms to game them. So some standards and evidence would remain secret. Insofar as the regime had to verify a rival state’s compliance, that would be an intelligence function. Restrictions would be triggered partly by evidence an affected company or researcher, or the public, could not inspect. Because this contemplated power cannot be bounded by a stable regulatory object (in the way, say, nuclear weapons can be), it would depend heavily on discretion, speed, internal access, and secret evidence. This has a highly illiberal character. Coercive power should be specific, limited, reviewable, and governed by general and knowable rules. Its characteristics (e.g., trigger, scope, evidentiary standard, duration, exceptions, means of review) should be stated before the power is granted. And the burden is on those who would propose it. A defender might answer that the proposed tool need not be coercive at all. That it could be narrow and advisory, focused on evaluation and transparency and readiness. But that wouldn’t solve the letter’s stated problem: racing. With race dynamics, each actor is under pressure not to slow down because others may continue (and thus the frontier keeps advancing). You need a mechanism to bind defectors. Voluntary norms tend to be great for binding people and firms that interact repeatedly and care about reputation. But the letter says each company and _country_… and you can’t rely on informal solutions when dealing with an unwilling state. That’s why the audience for this letter is Washington and why it calls for an international effort. Its diagnosis implies a binding mechanism. @deanwball thinks it is sensible to have a break-glass plan. That plan must involve a binding instrument, because nothing weaker addresses the problem the letter describes. But that therefore carries the burden for the use of coercive power, mentioned earlier. @johnschulman2's suggestion that labs design voluntary mechanisms among themselves is a different notion and coherent one (I would have signed that letter), but the word “country” makes this direction incompatible with the pacing letter. @OpenAI recently argued that a federal evaluator shouldn’t be able to block deployments. A week after, @AnthropicAI proposed that the government should be able to block deployments. Both labs endorsed the same letter. Whether or not the state may stop a deployment is a central question. Yet the letter accommodates both positions. What then, does the letter really say? Like the “We Must Act Now” letter from @erikbryn, @ajay_bcv, @akorinek, and @testingham before it, the letter secures agreement at an altitude where the main disagreement disappears. Lastly, the benefit of pacing is not established. The kind of slowdown the signatories have in mind would seek to buy us time for things like alignment, cyber defense, biological countermeasures, or scientific understanding -- things that increasingly depend on technologies a pause would restrict. E.g., Anthropic's framework relies in part on AI-based biological countermeasures and its security program uses AI to give defenders an advantage. A researcher in the letter's own friendly commentary was astonished at how much agents accelerated the work of the best alignment people he knows, and gave that as his reason for wanting six more months. When danger and our capacity to respond to that danger are plausibly both accelerating, the relevant question is whether this relationship is asymmetric in a safety-improving direction at the level of real-world risk. A slowdown needs to differentially slow the production of danger vs. our capacity to understand and contain that danger. The letter doesn’t attempt to establish that. It treats slower and safer as though they are the same; they are not. The letter is a serious warning, but it is no good as a warrant for an undefined power over inquiry.


In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…






China’s share of global GDP, according to their own almost certainly inflated numbers declined to 16.7% from 18.5% in 2021. America accounts for 26% of global GDP, up from 24% since 2021. “Not only is there virtually no chance of China overtaking the United States as the world’s largest economy, but it is also likely that the U.S. advantage in economic strength will expand further in the next decade” 👇


How could we train an AI model that was very good at overseeing another model: catching reward hacking and sandbagging, predicting unwanted behaviors or fine-tuning effects, etc.? We propose a new approach for doing this at scale: oversight foundation models.







I suspect that if we could figure out a way to make open weight models without dangerous capabilities — eg tamper resistant architectures or machine unlearning — Anthropic would be in favor of doing so (so would I!).





















