Joe Cole - e/acc

7.7K posts

Joe Cole - e/acc banner
Joe Cole - e/acc

Joe Cole - e/acc

@joecole

RLVR for expert judgment. Founder @tacitco. Prev: Fusion Sport (acquired). e/acc x h/acc.

The future Katılım Mayıs 2007
18.5K Takip Edilen16.7K Takipçiler
Joe Cole - e/acc retweetledi
Joe Cole - e/acc retweetledi
Joe Cole - e/acc retweetledi
dax
dax@thdxr·
used a trick @jlongster came up with agents can control browsers but you can also ask it to record network requests into a HAR file then it can derive a client for any website which is more efficient than browser controlling it every time made it build a quick uber eats cli
dax tweet media
English
181
238
5K
431.5K
Joe Cole - e/acc retweetledi
Dillon Mulroy
Dillon Mulroy@dillon_mulroy·
actually an insane thing for openai’s head of strategy to publicly say
Dillon Mulroy tweet media
Dean W. Ball@deanwball

Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.

English
357
981
12.4K
1.4M
Joe Cole - e/acc retweetledi
Y Combinator
Y Combinator@ycombinator·
Why do even our best AI models need tens of thousands of examples to learn skills that a human picks up in a handful of tries? Solving this problem is one of the great open challenges in modern AI. World models, which give AI an internal simulation of its environment, are one of the most promising paths forward. In this episode of Decoded, YC's @agupta and @FrancoisChauba1 discuss the intuition and math behind world models, new research, and current applications in self-driving, robotics, and more. 01:45 — What would perfect efficiency look like? 05:10 — World models in the human brain 09:20 — Control theory & the drone example 14:30 — When physics breaks down 17:45 — Chess, Go & the action space problem 24:10 — Why AlphaGo can't scale 28:00 — Monte Carlo tree search explained 34:00 — Self-Driving: state space is infinite 40:30 — Model-Free vs. Model-Based RL 44:00 — Why robotics is the hardest case 48:20 — World models that actually work 54:10 — JEPA & latent space tricks 59:00 — Open problems remaining 1:04:30 — Does this pass the squint test?
English
27
40
400
104.3K
Joe Cole - e/acc retweetledi
GiGadgets
GiGadgets@gigadgets_·
AS_3D Volumetric LED Display from Aoshow
English
15
296
2.4K
153.7K
Joe Cole - e/acc retweetledi
Madhu Guru
Madhu Guru@realmadhuguru·
The reason enterprises struggle to go beyond basic chat bots is the talent gap to build harnesses and evals. 1. Evals: do you clearly understand your use cases and can you replicate that in the form of offline and online evals. Do the evals express your ambition and do they push the jagged frontier of the models? Do they help you pick the right models on the quality-cost-latency curve? 2. Harness: do you have a system that manages routing, multiagency orchestration, context management, tool calling, memory that is independent of the models? 3. Talent: do you have the talent to build this all out on the frontier? This is the scarcest piece.
Nikesh Arora@nikesharora

Interesting take. The challenges today: 1. Most enterprises don't know how to make AI game changingly effective. As they embark on the journey, the use cases being addressed are "80%" single shot, semi deterministic use cases with multiple guardrails or humans in the middle. We are a little ways away from mass adoption of custom models. 2. More complex cases which require any multi agent orchestration and context retention are beginning to be conceived and tested. These cases will make model portability harder, requiring new evals and harnesses. One will have to commit to one structure and also commit to constantly updating and retraining your model.. 3. CIOs and CEOs aren't sure if the ultimate architecture is single stack, multi model - interoperable orchestration and context/harness/eval, or a custom model. Uncertainty causes slowdown on longer term decisions, which in this case is perhaps right. 4. Custom and Opensource come with the need to deploy on either your own GPUs or public cloud. "Interesting fact - if token prices fall as I hope - it will be cheaper to run frontier LLMs than open source on your own GPUs" 5. Generally horizontal solutions that can serve tens of thousands of customers make more money than vertical custom solutions, but maybe this time it's different? Even if we solve the model conondrum, the enterprises need to redefine workflows, collect more training data on each use case and rebuild the application in a simple UI flow, not everything will be done in a conversational window. But time will tell, this will continue to be a space to watch, lots of minds at work to solve this.

English
13
27
324
35K
Joe Cole - e/acc retweetledi
shirish
shirish@shiri_shh·
Kimi K3 vs Fable 5
Arena.ai@arena

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!

HT
93
132
3.3K
384.9K
Joe Cole - e/acc retweetledi
andrew chen
andrew chen@andrewchen·
This is what I want to say: Been playing w GLM 5.2 for the past week and it’s legit. If I can get this in a box sitting on my desk, I’m not sure I need much else But this is probably what’s actually true: Can imagine a guy running on a Pentium PC for the first time and being like, wow, I don’t need anything else. But actually we invent all sorts of cool new things we can use compute. So each frontier/adjacent LLM is cool but we’ll always want to use the next, better one
English
34
10
133
38.1K
Joe Cole - e/acc retweetledi
Sean Cai
Sean Cai@SeanZCai·
Its been no secret that benchmarking has long been outpaced vastly recently by model improvements, but the infrastructure around it breaking means raw performance becomes more unwieldy and expensive to measure. The slowdown in infrastructure to benchmark/eval effectively is the hinderance to most enterprise AI adoption. Enterprise AI adoption strays away from much model post-training efforts not only because of perceived high cost/know-how constraints, but because the act of post-training itself is highly subjective in how it translates to business KPIs in lack of custom evals. That much of data markets remains a game of telephone in translating task realism from contrived data producer —> post-training regime —> unwieldy benchmark —> real world application means that post-training in today’s regimes with today’s benchmarks scarcely adapts one’s data to actually relevant processes. In a period where benchmarks break constantly, it is useful to explore certain approaches of certain benchmarks whose construction behavior we should encourage. To name a few: @cognition FrontierCode's FP/FN analysis @harvey LegalBench's model kickoff prompts that avoid tasks sounding like instructing someone with amensia @OpenAI Healthbench's "consensus" mechanisms where extreme rigor is placed on aligning LLM as a judge with real world expert’s opinions (+ -10/+10 reward rubric grading) @AnthropicAI BioMysteryBench's superhuman quesiton generation via controllable properties of data And all of the benchmarks who've started listing infrastructure specs, as infrastructure specs become larger determinants of model performance at long horizons. Altogether, multidimensionality of unverifiable verification approaches, as well as overtures from the cost-latency side of the Pareto curve threaten the validity of most benchmarks today. Just as AI engineering become an overnight skill in 2023, eval creation shall become one in the latter half of this year as a subset of that.
English
8
4
92
8.9K
Joe Cole - e/acc retweetledi
@jason
@jason@Jason·
When tokens go down 90% by the end of the year and then another 90% next year, everyone's opinions on artificial general intelligence and superintelligence are going to change radically I'm currently on an unlimited GLM 5.2 bittensor subnet and I can tell you your behavior changes radically when token prices plummet
English
259
222
4.5K
568.9K
Joe Cole - e/acc retweetledi
elie
elie@eliebakouch·
really interesting research direction imo, claude and gpt have very similar benchmark scores on pretty much everything, but yet are so different in the way they interact with humans
elie tweet media
Mira Murati@miramurati

Today we share the worldview behind our mission. Human values don't average out. Local knowledge can't be centralized. The good future has many AIs, raised in different places, shaped by the people they serve, disagreeing with each other the way we do. thinkingmachines.ai/blog/the-futur…

English
4
5
121
13.2K
Joe Cole - e/acc retweetledi
Sasha Rush
Sasha Rush@srush_nlp·
No idea what Thinking Machines is working on, but this line goes hard.
Sasha Rush tweet media
English
36
60
1.1K
64.1K
Joe Cole - e/acc retweetledi
Soumith Chintala
Soumith Chintala@soumithchintala·
What do we do at @thinkymachines: Personalization/sovereignty, Human Participation, Decentralization. Democratize AI and make it useful for people. All three of them reduce society's dependence on centralized AGI companies (including ours when we get important), and that is a future worth aiming for. You've seen a preview of this with Tinker, Interaction models and our research openly published on Connectionism. A **lot** more to come very very soon...
Thinking Machines@thinkymachines

We're building AI that people and organizations can shape and make their own. AI should extend our will and judgment instead of neglecting it; enabling that is the technical challenge we are working to solve. thinkingmachines.ai/blog/the-futur…

English
43
60
829
212K