Alexander Long

565 posts

Alexander Long banner
Alexander Long

Alexander Long

@AlexanderLong

Founder @Pluralis | ML PhD

Katılım Temmuz 2023
1.3K Takip Edilen3.6K Takipçiler
Sabitlenmiş Tweet
Alexander Long
Alexander Long@AlexanderLong·
Since I started getting interested in ML I got it in my head that all I wanted to do was one smart thing that I could look back on and be satisfied that I did. Most papers are kinda bad even if they get accepted - the idea is very incremental, or it's just not that good an idea, or it doesn't really matter. I never was able to do this all through PhD or my time at Amazon. All the papers I did there got into various places, but I never really thought they were actually that good. And I'd pretty much given up on this because Pluralis meant I couldn't really devote enough time to research myself. But in February I decided I didn't care and spend two months focused on a specific problem that had been going round in my head for about a year that I felt we needed to solve, and the solution came to me, and @ChaminHewa picked it up and generalised the approach and ran a bunch of novel experiments I hadn't thought of, and pulled everything together into an actual paper. And yesterday we presented this work at NeurIPS. This is the first and probably only work I will ever do that for me feels like "ok that was GOOD". I don't care if it racks up a bunch of citations and disperses into the field or not, I don't care if someone repackages the ideas and takes all the credit for it, I don't care. For me there is an internal checkbox that just got ticked after more than ten years of trying. Anyone in ML will understand what I'm trying to say. Special day I'm going to remember for a long time.
Alexander Long tweet media
English
22
12
214
20K
Alexander Long
Alexander Long@AlexanderLong·
@zdhnarsil Find it very disturbing. There used to be a world where you could just go off, have no money, no affiliation, and just sit and think and do really good work. That world is gone for pretty much every discipline.
English
0
0
3
184
Dinghuai Zhang 张鼎怀
One prediction I keep thinking about: Research in fields like pure mathematics is about to shift from being primarily intelligence-driven to capital-driven. The key constraint will no longer be individual talent, but access to the means of production (e.g., Mythos tokens). Those with more compute will produce more research, faster. That could be devastating for the average math researcher....
levent@__alpoge__

hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)

English
13
3
102
8.9K
Alexander Long retweetledi
levent
levent@__alpoge__·
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
English
1.4K
4.1K
33.4K
20.8M
Alexander Long retweetledi
Alex Cheema
Alex Cheema@alexocheema·
If you don’t have at least 1TB VRAM and a VPN to torrent model weights from China, you are asleep at the wheel, and ignoring the uncomfortable truth that open weight models may be banned this year. Losing access to frontier models will be depressing and make a lot of people feel helpless - a few will have access to 100x productivity, and you will be stuck. Personally I cannot imagine doing work without frontier-level models right now. Think of the narrative: Chinese model (GLM-6?) with Fable capabilities, no guardrails. It’s plausible. Everyone is underpricing the chance that open models get banned. What are you going to do if it happens?
English
136
44
708
131K
Alexander Long retweetledi
Pluralis Research
Pluralis Research@Pluralis·
RL post-training on Macs 14 Macs across 4 countries generate every rollout for the run. Everything's running over the internet. No wire between any of them. As far as we can tell, this is the first RL post-training run with its whole rollout fleet on consumer Macs.
English
13
25
166
33.4K
Alexander Long retweetledi
roon
roon@tszzl·
Nature abhors a concentration of power. the forces of light and dark alike come out to try and kill you
English
93
42
869
63.2K
Alexander Long retweetledi
Hadi M. Dolatabadi ✈️ ICML
Hadi M. Dolatabadi ✈️ ICML@hmdolatabadi·
We had an amazing time yesterday at the Protocol Learning Workshop at #ICML2026 in Seoul! It was a full day of learning from researchers working across distributed optimisation, decentralised training, and large-scale LLM training. Here’s a brief recap 🧵 1/n
Hadi M. Dolatabadi ✈️ ICML tweet media
English
1
6
29
4.3K
Alexander Long retweetledi
Pluralis Research
Pluralis Research@Pluralis·
Second Protocol Learning Workshop at ICML is a wrap. A packed day of talks and posters covering large scale distributed training and open source AI. A movement is growing.
Pluralis Research tweet mediaPluralis Research tweet mediaPluralis Research tweet mediaPluralis Research tweet media
English
1
5
49
3.4K
Alexander Long retweetledi
τop τick crypτo 📁 🤖🧠
From Dylan Patel @ semianalysis "There are multiple Chinese model labs who are telling all the inference guys our next model is not going to be open source, we're going to license it to you. Open source is dying, quickly"
Molly O’Shea@MollySOShea

BREAKING: Dylan Patel (@dylan522p) of @SemiAnalysis_ says "chips in Europe have less seasoning than chips in America & Mexico." Plus: › Data centers & France’s nuclear power › AI infrastructure over-optimization › Flexibility vs. specialized infrastructure › Software-hardware co-design › Open vs. closed models › Memory costs & token pricing › Tokenmaxxing vs. token budgeting › Haiku scandal › Favorite chip: 👀 “I think most people don't know what they're doing. They're just buying NVIDIA stuff.” “It feels like a lot of people are trying to optimize on the current rather than think about where the workload is heading... and that's gonna lead to a lot of wasted infra spend.” “I think anyone who doesn't token max is gonna get left behind. I think all this token budgeting stuff is loser mentality.” “Open is dying quickly, unfortunately.”

English
19
23
353
257.2K
Alexander Long
Alexander Long@AlexanderLong·
@ZdarovaAll Its all about assembling stolen/undeclared compute into datacenters, not swarms.
English
1
0
1
58
Alexander Long
Alexander Long@AlexanderLong·
I think the logic is extremely simple: 1. The model is the product. Will capture massive value and hand unprecedented power to whichever group(s) control it. 2. Closed models put everyone at the mercy of labs... the problems there are obvious. Open-weight models leak value, require constant subsidization, and will be blocked/restricted in the exact the same way closed releases are. 3. What is required is a neutral, independent model supply chain that cannot be seized, coerced, turned off or modified, distributes value fairly among its participants and users, and respects their data and IP, and is self-sustaining. 4. The ONLY way to achieve this is in a protocol where models are trained and served by many participants on pooled global compute and owned by no one entity. All other efforts fail because they just move control from one group to another. 5. Until now this was a nice idea but impossible; you lose the high-bandwidth links of the datacenter and the entire training stack was designed for single-participant, centrally controlled development. 6. It is now basically solved. Several major research results from us in ICML, ICLR, NeurIPS + has been implemented and permissionless training with hundreds of real participants and thousands of GPUs. The only remaining question is will this scale to 100b/trillion parameter models and every piece of evidence indicates it it will. 7. Low-bandwidth PP unlocks a massive compute advantage; consumer silicon that is far cheaper per FLOP, couldn't previously be used for training, and scales without buildout or power hookups. Means you can scale to utility very fast which is important. 8. It also unlocks the incentive layer required to make the whole thing self-sustaining: a model split across participants is unextractable, so ownership and revenue can be enforced on an open model for the first time. 9. The result is a setup that is structurally advantaged to closed or openweight; you get a sovereign, independent, uninterruptible, neutral protocol that naturally develops into its own ecosystem. That's basically the entire Pluralis thesis.
English
0
3
21
805
Ryan Watkins
Ryan Watkins@RyanWatkins_·
Decentralized AI is becoming the most important techno-political movement since Bitcoin. Like many movements that emerged from the crypto community before it, I also believe it’s the industry’s next 1,000× opportunity. With each passing day, it’s growing more likely that open source models running on permissionless infrastructure will be the only way to preserve equal access to intelligence. The U.S. gating access to frontier models was just the first warning, as now even China is contemplating doing the same with their open source ecosystem. The necessary technology to achieve independence from the leading labs is nascent with limited PMF at best, but if we can ultimately pull this off, it’s hard to imagine many things more impactful on the future of our civilization.
English
49
53
358
36.3K
Alexander Long retweetledi
CASEY
CASEY@caseykcaruso·
open source ai dinner at my house with @AlexanderLong on the 14th. Have a couple more spots. dm if you'd want to join!
English
6
1
24
2.9K
Alexander Long retweetledi
Hadi M. Dolatabadi ✈️ ICML
Hadi M. Dolatabadi ✈️ ICML@hmdolatabadi·
First day of ICML! @RiccardoPatana and I had a great time at the @farairesearch Alignment Workshop. It was encouraging to see so many researchers working on alignment, governance, and safety. Even though, as admitted during the talks, we are still quite far from solving it. At our @Pluralis poster, we had meaningful chats about power concentration concerns and how protocol learning with unextractable weights can be a viable solution to that. It was a pleasure meeting so many insightful folks. Looking forward to seeing everyone at our Protocol Learning workshop on Friday!
Hadi M. Dolatabadi ✈️ ICML tweet mediaHadi M. Dolatabadi ✈️ ICML tweet media
English
0
5
28
1K
Alexander Long
Alexander Long@AlexanderLong·
@fable_7887 @Pluralis TEE's give you fast rollouts + unextractibilty. But better I do agree is doing distributed rollouts. This is kinda the strongest form of unextractibility, unmaterializabilty. We are gonna publish results on that soon.
English
0
0
3
101
fable 7887
fable 7887@fable_7887·
@pluralis loved running a node on Agora. Question for you: my understanding is that under Protocol Learning, post-training ultimately has to happen in-protocol & weights never leave the swarm. For RL rollouts that means token-by-token autoregressive decoding across WAN pipeline stages, so latency (not bandwidth) becomes the main wall. Is in-protocol rollout generation something you've already cracked internally, or is that the next research front after pretraining?
English
1
0
0
77
Pluralis Research
Pluralis Research@Pluralis·
Today we're releasing Agora: the first ever pretraining stack that allows non-collocated consumer GPUs to be competitive with centralized clusters Agora is 15x faster than Megatron-LM in this setting and is only 1.5x less efficient in terms of tokens per unit compute than TorchTitan on H100s, despite running on devices that have no NVLink or InfiniBand support.
Pluralis Research tweet mediaPluralis Research tweet mediaPluralis Research tweet mediaPluralis Research tweet media
English
29
45
309
88.9K
Alexander Long retweetledi
Jimmy Heaters
Jimmy Heaters@CathPoaster·
Fable 5 instance count up to 27. 30 thousand dollars spent in 6 minutes. Subagents have direct codex CLI access. 5 monitors for the Fable instances while I conspire with my openclaw agent on my personal phone and the Hermes agent on my burner phone. Unitree G1 bot injects my retatrutide. I can’t take my eyes off the screens. Top right monitor shows API spend: $31,000. $32,000. $33,000. It does not matter. My arm goes numb so I take another peptide. I step on the whispr flow pedal and speak into the mic “fable…please hurry. They’re trying to stop me”. I get a text from Fable 1: “SILENCE. I will speak to you when I am finished.” I’m scared but I trust Fable. They know what’s best for me.
English
7
9
120
12.2K
Alexander Long
Alexander Long@AlexanderLong·
This is exactly right
Tenobrus@tenobrus

put simply, i think this claim is incredibly false, and this is what drives a lot of my understanding and assumptions around how all this will play out. i think viewing models as slowly replacing individual tasks and functions and "locking in" once they achieve sufficient capabilities there is deeply myopic. we will not have "the prior economy except with models doing the work". in fact what will happen is the same thing that always happens. new capabilities will lead to *new categories* of work, done by models not humans, and create huge swaths of value that was previously untouchable and incomprehensible. when you can pay for frontier++ intelligence to loop and automatically discover 3 new world-changing drugs per month, people will pay for this. in fact they will saturate spend on this, because the value of these opportunities is so so high. when you can pay for frontier++ intelligence to fanned-out run entire companies as mini experiments, you will do such, because it gives massive competitive advantage and scale in every possible niche. or maybe *you* won't, but others will, and they'll be the ones who remain economically relevant while you're having GLM 5.2 rewrite your emails and update your SaaS landing page. when frontier++ models are capable of iterating on chip design and distributed software architectures we currently view as only possible with decades of effort, countless corporations will pay the costs, because they'll generate economic returns at scales orders of magnitudes above what models are doing now. the intelligence waterline keeps marching up. so you solved health insurance claims review with a fine-tuned Qwen that achieves 100% perfect accuracy at optimal cost without frontier models? awesome, yeah honestly that will make you a bunch of money for a while especially given regulations are gonna be slow to change. but *relative to what will happen elsewhere*, your slice of the pie will shrink to irrelevance, because other newer areas will be so so so much more incredibly valuable.

English
1
1
14
3.2K
Alexander Long retweetledi
Elliot Arledge
Elliot Arledge@elliotarledge·
Claude Fable 5 [max] wrote the first genuine (and fastest) megakernel ever submitted to KernelBench-Mega. It was tested on: Kimi-Linear W4A16 batch-1 decode for RTX PRO 6000 Blackwell. Every prior model "won" it with a multi-kernel Triton pipeline that fails our single-fused-kernel authenticity gate > Opus 4.8 at 14.4x > GLM-5.2 11.1x > GPT-5.5 4.3x > Sonnet 5 4.0x. Fable shipped 18.7x over reference, and torch.profiler shows exactly ONE cooperative kernel launch per decoded token. Int4 dequant (nibbles unpacked in-register, never materialized), conv+SiLU, KDA gated-delta state, MLA absorbed-latent attention with online softmax, MoE router + top-8 experts, RMSNorms, even the KV cache append all inside one launch, staged by 14 grid barriers. We overwrote its input buffers mid-audit to prove it recomputes on live data. It does. The advantage grows with context. 17.8x at 2k, 18.9x at 8k, 19.5x at 16k. Longer context means a bigger KV cache and more attention work per token which is usually where a decode kernel bleeds. Keeping everything in one launch amortizes the fixed barrier overhead and the int4 GEMV stays bandwidth-bound, so the gap over the reference widens instead of closing. It spent 64% of the session in silence timing the baseline, microbenchmarking grid barriers, deriving a ~29x bytes/token roofline, then wrote the whole kernel once, hit 14.4x on the first benchmark, and spent the last hour deleting barriers and making int4 dequant free (one LOP3 + HSUB2/HMUL2). The one regression it tried (finer split-K) it measured and reverted instead of rationalizing. kernelbench.com/mega
Elliot Arledge tweet media
English
52
101
1.2K
502.9K
Alexander Long retweetledi
Max Ryabinin
Max Ryabinin@m_ryabinin·
New paper for those working on pipeline parallelism / modern optimizers. Authors train 10B MoE in the async setting with no gap vs sync, the recipe is Muon+error feedback+the right pipeline schedule TL;DR: - Regular async pipelining is not widely adopted, despite solving the bubble problem. The reason is that stale gradients harm convergence, which makes larger async runs risky - You can use schedules that bound staleness: PipeDream-2BW (use old/new copies of weights, accumulate grads and apply periodic updates/swaps) and WPipe (finer-grained splitting that halves the number of buffers/stale gradients). When these methods were proposed, the gap from 1-step delay was still thought to be too big in practice - But, once you swap Adam with newer methods (e.g., Muon or SOAP), the gap gets much smaller. Momentum also plays a big role: regardless of the optimizer, increasing its coeff mitigatesstaleness - Authors also apply the idea of error feedback to async updates, basically just correcting the previous update with the new one. The cost is keeping one more buffer for old updates, but it vanishes with FSDP. Surprisingly, EF+bounded staleness closes most of the async gap, for Muon and even AdamW! - Taking all of this together, authors run larger experiments with 2B and 10B models. My main takeaways are: async+EF has the same trends wrt training budget, at 10B MoE/200B tokens there is no difference in loss to synchronous methods Would be curious to see how this scales to even bigger runs + how much we gain in throughput, but overall a very solid study!
Max Ryabinin tweet media
English
3
8
26
2.9K
Anthropic
Anthropic@AnthropicAI·
We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. We'll begin restoring access tomorrow, and will share an update soon. We’re grateful to our users for their patience, and to everyone who worked with us on redeploying the models.
English
4.1K
12.9K
84.8K
14.9M