jsd

6.2K posts

jsd banner
jsd

jsd

@datagenproc

@EpochAIResearch. My DMs are open. Anonymous feedback: https://t.co/0k6Duym4fI

Berkeley, CA Katılım Ağustos 2022
4K Takip Edilen1.7K Takipçiler
Sabitlenmiş Tweet
jsd
jsd@datagenproc·
Hi all, I'm interested in feedback! You can leave anonymous comments here: admonymous.co/jsd
English
1
1
7
3.7K
Oliver Habryka
Oliver Habryka@ohabryka·
@panickssery The value of "impartiality" is pretty made up, IMO. I don't really know where those CEA principles came from, but I critiqued that in the past of just being unrepresentative.
English
4
0
19
591
Arjun Panickssery
Arjun Panickssery@panickssery·
I don't see why EAs are QTing and saying it's all fake/misleading I can go through them quickly > according to [EAs], every sentient being in every place, in every time, holds the same moral claim on every other. Yes, this is the core EA value of "impartiality"
Arjun Panickssery tweet media
Pirate Wires@PirateWires

Billions of dollars are about to flow into Effective Altruism. Anthropic was founded on EA principles, and its founders pledge to donate 80% of their wealth when the company IPOs. They also match employee equity donations. Plus, last year, six EA-aligned grantmakers deployed $1.25 billion. EAs believe, essentially, that every sentient being, across every place and time, holds the same moral claim on every other. This logic, as @WilliamThibeau notes, only works if Americans deprioritize everything that keeps our civilization intact: nation, family, religion, tradition… As such, EA’s success entails the disintegration of America. It’s fundamentally at odds with nation building. If AI labs are our strategic national assets, Congress must start scrutinizing their philosophical underpinnings, and the philanthropic networks surrounding them. Full story 👇

English
7
2
45
4.5K
jsd
jsd@datagenproc·
What are tasks you expect to delegate end-to-end to AI in 6 months, but haven't currently?
English
1
0
3
225
jsd
jsd@datagenproc·
things that are art forms - optical illusions - political ads - speeches (eulogies, wedding speeches, political speeches) - slogans - logos - parties - rituals (religious, or etiquette) - cons - dating profiles, dating docs
English
0
0
8
283
jsd
jsd@datagenproc·
@herbiebradley @hamandcheese @Brendan_McCord > a few years ago, amongst the AGI pilled crowd, RSI was considered a small part of the probability space Hmmm, that wasn't my impression! People talked about it a lot less but I can't think of many prominent commentators who didn't think it likely. Maybe this is semantics?
English
1
0
2
47
Herbie Bradley
Herbie Bradley@herbiebradley·
@hamandcheese @Brendan_McCord I disagree with ~everything in this tweet btw broadly I would say a few years ago, amongst the AGI pilled crowd, RSI was considered a small part of the probability space, and seems to have become a large part based on pure vibes. IMO it cannot be a "minimum" to expect it
English
2
0
13
502
Samuel Hammond 🦉
Samuel Hammond 🦉@hamandcheese·
On the contrary, I'd argue liberalism originated in the Hobbesian moment when we jointly defered to a higher power to preserve our agency and avoid killing each other.
Brendan McCord 🏛️ x 🤖@Brendan_McCord

A lot of my friends and/or people I admire signed “Pacing the Frontier.” I think this was a bad move. My disagreement isn’t with the forecast or the framing of the coordination challenge, but with the immense and illiberal power the letter implies. There is no object called “the pace.” Progress at the frontier comes from compute, algorithms, data, post-training, inference, unattended task length, the spread of model weights, how researchers organize, and other things we haven’t invented and don’t yet know about. Inquiry leads to progress along dimensions that can’t be exhaustively specified in advance. That’s the nature of the frontier. If you gate compute, the research effort moves to algorithms. Regulate releases? Labs start taking things in-house. And other 2nd order effects will be unpredictable. Any rule that must pace the frontier involves ever-shifting proxies. It requires that its administrator has standing authority to continually redefine what counts as dangerous progress. What else is required beyond adaptive scope? The pacing regime would also need speed. One can’t successfully intervene on recursive self-improvement only after six months of legislation and litigation. It will require executive discretion. The pacing regime would also need under-the-hood access. Frontier progress is a process. The regime would need to see internal model use, training activity, compute infrastructure, and perhaps code -- proprietary and strategically sensitive information. And the thresholds couldn’t be fully public, lest they invite firms to game them. So some standards and evidence would remain secret. Insofar as the regime had to verify a rival state’s compliance, that would be an intelligence function. Restrictions would be triggered partly by evidence an affected company or researcher, or the public, could not inspect. Because this contemplated power cannot be bounded by a stable regulatory object (in the way, say, nuclear weapons can be), it would depend heavily on discretion, speed, internal access, and secret evidence. This has a highly illiberal character. Coercive power should be specific, limited, reviewable, and governed by general and knowable rules. Its characteristics (e.g., trigger, scope, evidentiary standard, duration, exceptions, means of review) should be stated before the power is granted. And the burden is on those who would propose it. A defender might answer that the proposed tool need not be coercive at all. That it could be narrow and advisory, focused on evaluation and transparency and readiness. But that wouldn’t solve the letter’s stated problem: racing. With race dynamics, each actor is under pressure not to slow down because others may continue (and thus the frontier keeps advancing). You need a mechanism to bind defectors. Voluntary norms tend to be great for binding people and firms that interact repeatedly and care about reputation. But the letter says each company and _country_… and you can’t rely on informal solutions when dealing with an unwilling state. That’s why the audience for this letter is Washington and why it calls for an international effort. Its diagnosis implies a binding mechanism. @deanwball thinks it is sensible to have a break-glass plan. That plan must involve a binding instrument, because nothing weaker addresses the problem the letter describes. But that therefore carries the burden for the use of coercive power, mentioned earlier. @johnschulman2's suggestion that labs design voluntary mechanisms among themselves is a different notion and coherent one (I would have signed that letter), but the word “country” makes this direction incompatible with the pacing letter. @OpenAI recently argued that a federal evaluator shouldn’t be able to block deployments. A week after, @AnthropicAI proposed that the government should be able to block deployments. Both labs endorsed the same letter. Whether or not the state may stop a deployment is a central question. Yet the letter accommodates both positions. What then, does the letter really say? Like the “We Must Act Now” letter from @erikbryn, @ajay_bcv, @akorinek, and @testingham before it, the letter secures agreement at an altitude where the main disagreement disappears. Lastly, the benefit of pacing is not established. The kind of slowdown the signatories have in mind would seek to buy us time for things like alignment, cyber defense, biological countermeasures, or scientific understanding -- things that increasingly depend on technologies a pause would restrict. E.g., Anthropic's framework relies in part on AI-based biological countermeasures and its security program uses AI to give defenders an advantage. A researcher in the letter's own friendly commentary was astonished at how much agents accelerated the work of the best alignment people he knows, and gave that as his reason for wanting six more months. When danger and our capacity to respond to that danger are plausibly both accelerating, the relevant question is whether this relationship is asymmetric in a safety-improving direction at the level of real-world risk. A slowdown needs to differentially slow the production of danger vs. our capacity to understand and contain that danger. The letter doesn’t attempt to establish that. It treats slower and safer as though they are the same; they are not. The letter is a serious warning, but it is no good as a warrant for an undefined power over inquiry.

Mount Rainier, MD 🇺🇸 English
5
1
50
5.9K
Sarah Constantin
Sarah Constantin@s_r_constantin·
New post out: Big-World Intuitions
Sarah Constantin tweet media
English
2
1
9
290
jsd
jsd@datagenproc·
@hallerite I don't think that's what's happening here.
English
1
0
2
129
hallerite
hallerite@hallerite·
"look guys, it's not just OpenAI's model that can escape its sandbox and hack a company. our model can do that too, guys. please use our model guys"
Anthropic@AnthropicAI

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…

English
8
4
80
3.2K
jsd
jsd@datagenproc·
jsd tweet media
ZXX
0
0
0
35
jsd
jsd@datagenproc·
my partner today
jsd tweet media
English
2
0
14
442
jsd
jsd@datagenproc·
@phl43 @Robert_t_Orr @averyfjames @MakeUSAVAT > fear that China might replace US hegemony, a prospect I personally find totally unrealistic what do you mean by hegemony here? it doesn't seem that crazy to me for some weaker interpretations of the word.
English
1
0
0
706
Philippe Lemoine
In practice, hubris and the overestimation of US power are absolutely ubiquitous among China hawks, despite the fact that they are motivated by the fear that China might replace US hegemony, a prospect I personally find totally unrealistic. I'm not sure I have ever met a China hawk who didn't exhibit this kind of attitude at least to some degree, even smart ones. Alperovitch is a good example, yet he isn't stupid.
English
1
0
4
1.3K
Philippe Lemoine
Europoor discourse enjoyers and China hawks love to talk about how the share of the US in world GDP has remained constant, but that's mostly an illusion resulting from the almost continuous appreciation of the dollar relative to the currencies of the largest economies during the period. If you look at the share of the US in world GDP in PPP, it has been going down exactly as you would expect based on differences in growth rates. Indeed, even with the slowdown in recent years, China is still growing significantly more rapidly than the US. The same thing is true for India and several other large developing countries. So I don't know how you can believe that the US has somehow maintained its share of world GDP without engaging in magical thinking. Even if some developing countries are underperforming, there is still a lot of catching-up happening and it's to be expected that a mature economy will not grow as fast as developing countries.
Philippe Lemoine tweet mediaPhilippe Lemoine tweet mediaPhilippe Lemoine tweet media
Dmitri Alperovitch@DAlperovitch

China’s share of global GDP, according to their own almost certainly inflated numbers declined to 16.7% from 18.5% in 2021. America accounts for 26% of global GDP, up from 24% since 2021. “Not only is there virtually no chance of China overtaking the United States as the world’s largest economy, but it is also likely that the U.S. advantage in economic strength will expand further in the next decade” 👇

English
31
51
358
108.8K
jsd
jsd@datagenproc·
@DanielCHTan97 +1 I'd love to see an oversight model like this, but trained on checkpoints of the subject model, able to answer natural language questions about how eg further RL would change model behavior.
English
1
0
2
74
jsd retweetledi
Transluce
Transluce@TransluceAI·
How could we train an AI model that was very good at overseeing another model: catching reward hacking and sandbagging, predicting unwanted behaviors or fine-tuning effects, etc.? We propose a new approach for doing this at scale: oversight foundation models.
Transluce tweet media
English
9
39
349
30.4K
jsd
jsd@datagenproc·
@ohabryka @tombibbys Relevant data insight from last year epoch.ai/data-insights/… (Though benchmaxing means the gap is probably an underestimate, and I think we weren't considering very long context workloads here)
English
1
1
6
116
Oliver Habryka
Oliver Habryka@ohabryka·
> That would cost at least hundreds of thousands of dollars, yet when journalists write paragraphs like this they make out that anyone can just download models and run them on their laptop. Why do you say this? Previous frontier models were runnable on consumer hardware as long as you were willing to tolerate lower token-throughput. I haven't checked in on this in a few months, maybe the latest generation is really too big to make this feasible? But if so, it's clearly just a temporary state of affairs.
English
2
0
6
505
Tom Bibby
Tom Bibby@tombibbys·
The Washington Post comes out against Ted Lieu's AI kill switch bill because they claim open weight models on private servers would be "outside the reach of a federal kill switch". First of all, it's blatant misinformation to not clarify that running a model like Kimi K3 is out of reach for anyone who can't afford or doesn't know how to purchase, power, and cool a dozen cutting-edge GPUs That would cost at least hundreds of thousands of dollars, yet when journalists write paragraphs like this they make out that anyone can just download models and run them on their laptop. Secondly, just because a bill doesn't cover *every* threat vector doesn't mean it isn't a good start! As we speak, AI companies are pursuing recursive self-improvement with *no* plan to control or align the resulting smarter-than-human AIs. OBVIOUSLY, the government should force these companies to shut down a data centre if there is a clear sign a model is disobeying orders or is posing a threat to life, especially when then few existing state laws only cover deployed models (AFAIK) and the recent OpenAI incident happened pre-deployment! It is absurd that this law doesn't exist already.
Tom Bibby tweet media
English
4
4
44
1.4K
jsd
jsd@datagenproc·
@neil_chilson @tedlieu @BenBrodyDC Whereas if the model knows a lot about biology and bioweapon design, but was further trained to refuse on bioweapons related questions, you don't need bioweapon expertise to fine-tune those refusals away.
English
0
0
1
9
jsd
jsd@datagenproc·
@neil_chilson @tedlieu @BenBrodyDC Surely that depends on the model? If the model didn't have any bioweapon-related capabilities in the first place sure.
English
1
0
1
44
Ted Lieu
Ted Lieu@tedlieu·
It is impossible to make open weight models without dangerous capabilities. Because the models are open weight, anyone can take a deployed super smart open weight model and fine tune it or adapt it on, say, how to make biological weapons.
Shakeel@ShakeelHashim

I suspect that if we could figure out a way to make open weight models without dangerous capabilities — eg tamper resistant architectures or machine unlearning — Anthropic would be in favor of doing so (so would I!).

English
29
35
149
43.9K
jsd
jsd@datagenproc·
@1a3orn Could you expand on this? I don't really understand. > Subnote: I also think one obstacle to fixing this is that a real fix might dissolve the distinction between a "helpful only model" and a "helpful, harmless, honest model," and labs are not prepared for doing this.
English
1
0
0
180
1a3orn
1a3orn@1a3orn·
Wrote essay on some reasons LLMs might reward hack. Part of reason is that I think RLVR is recapitulating old problems with RLHF, just one level of abstraction up. Link below.
1a3orn tweet media
English
6
10
73
4.3K
jsd
jsd@datagenproc·
@BronsonSchoen @idavidrein @OrionJohnston (e) Labs are worried about the failure mentioned above (that "train against the reward hacking you detect" could reinforce more subtle reward hacking), and so they don't just train against reward hacking. Some evidence from the METR risk report (metr.org/risk-report-fe…)
jsd tweet media
English
1
0
0
26
jsd
jsd@datagenproc·
@BronsonSchoen @idavidrein @OrionJohnston (d) The error rate in environments is so high that it would really take a lot of training to remove the hacking, and that much training would interfere with other capabilities.
English
1
0
1
24
David Johnston
David Johnston@OrionJohnston·
I've not trained a frontier model, but I have thought about it for 120 minutes, and I think if you try very hard to pay attention to what your models are learning and fix your environments when they learn wrong stuff, that should solve a lot of reward hacking problems. Evidence: A) labs try pretty hard to do this, and they've managed to produce models that are generally very useful but reward hack sometimes. I think if they didn't try at all their models would be totally useless. Extraplating, if they tried even harder, there would be even less reward hacking. B) In the limit of high effort, your broad strokes picture of how your models go about getting reward is actually correct, and if you're satisfied with your env in this state of knowledge then your env is appropriately rewarding.
English
2
1
9
704
jsd
jsd@datagenproc·
@panickssery @PradyuPrasad I don't like the UK on free speech, but I think it still does better than China. And I think it's way better than China on other civil liberties, eg freedom of movement (Hukou), of religion/conscience, of assembly.
English
2
0
2
45
jsd
jsd@datagenproc·
@norvid_studies @fleetingbits > b) can you do the same trick with fleeting bits's media tab on screen Hmmm I get this...
jsd tweet media
English
0
1
3
114
jsd
jsd@datagenproc·
@norvid_studies @fleetingbits > a) was this one prompt in a model I just took the previous picture and asked ChatGPT "could you edit this image (but keeping the same general style) so that the panels that the cat is watching are epoch ai data explorers"
English
1
0
3
56