Can Sar

5.9K posts

Can Sar

Can Sar

@cansar

Codex @OpenAI. Previously: Cash App & Block, YC S17, Google, several startups.

Katılım Ağustos 2008
328 Takip Edilen4.4K Takipçiler
Sabitlenmiş Tweet
Can Sar
Can Sar@cansar·
A personal update: I’m joining OpenAI to work on Codex, the best agent app in the industry. After years building consumer products, my obsession since September has been to make Block one of the most AI-native engineering orgs in the world. I’m proud of what we’ve built: faster Eng velocity with RPI in October, an internal multi-harness ADE in January, Slack-based cloud agents in February, and AI woven into every step of product development. The team is just getting started. Along the way I tried almost every agent orchestrator out there — 30+ over the past year. Codex felt special from the day it launched. Lately, it’s started to point at something much bigger. What used to be scattered across prompts, tabs, tools, files, and handoffs is starting to come together in one place. The app is becoming the single surface for knowledge work — letting you accomplish more than you ever thought possible, while staying calm and focused. You can already see the early signs: making a slideshow and editing it right there, working inside a custom site built just for you and your team, handing off to the browser and the computer while you watch it go. Getting this right takes powerful models. But it also takes taste, product thinking, and real care for the person on the other end. You see that in the product, and you feel it the moment you talk to the team. I couldn’t be more excited to join the Codex team. Together we’re going to build the future of computing and knowledge work, and we’re going to build it for you.
English
88
17
738
113.5K
Thomas Ricouard
Thomas Ricouard@Dimillian·
Tomorrow I’m reviewing and redoing my entire Codex flow. How I manage tasks and projects. More on that soon.
English
44
4
368
16.3K
Can Sar
Can Sar@cansar·
@s_batzoglou @DeryaTR_ @ChrisRMcGuire Do you think that they would progress more slowly if we sold them our best chips? And does that include RSI like effects where the next 1-2 years potentially matter an immense account.
English
3
0
0
149
Chris McGuire
Chris McGuire@ChrisRMcGuire·
Kimi admits it is compute constrained, and is struggling to serve K3. The same thing happened to DeepSeek when it released v4. When Chinese AI labs say their #1 constraint is compute, they aren’t lying. They don’t have enough chips to serve the model at scale to customers. If we stop China from buying, smuggling, or remotely accessing AI chips, it will be harder for them to either make advanced AI models or serve them at scale. But instead we are selling them the compute capacity they need most, and have also loosened restrictions on smuggling and remote access. We are making it easier for China to catch up, and are acting surprised when they release good models. In reality, models like K3 are made with and run inference on US chips - and are the direct result of the weakening and non-enforcement of our export control policies. The good news is if we start closing loopholes in our export control policies and enforcing them more vigorously, we can still constrain China’s future AI capabilities. But this is the consequence of our non-serious approach to export controls over the last 18 months.
Kimi.ai@Kimi_Moonshot

Kimi K3 has received far more love than we expected, and our GPUs are feeling it. Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected. We're adding capacity as fast as we can and will reopen new subscription spots in batches. Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable. Thank you for your patience and understanding!

English
275
45
307
153.6K
Can Sar retweetledi
Ryan Lopopolo
Ryan Lopopolo@_lopopolo·
It’s not enough to throw a bunch of agents at the plan since that’s not what you’re trying to produce — you want the result of the plan. The plan will never be well formed enough, so just produce the thing directly; it will give you way more signal on what is under- or mis-specified.
Ryan Lopopolo@_lopopolo

@blader While you may converge to _something_ high quality doing this, it is very likely to not be what you want. You generally want to have a single agent close over the task and, instead of adversarial review on the plan, statically reject misaligned code via the codebase itself.

English
9
4
68
9.3K
Can Sar
Can Sar@cansar·
@BretDevereaux It was a fun movie but so absurd that at some point my friend who’d talked us into going started laughing so hard that her friend had to hold her mouth shut lest we get kicked out.
English
1
0
4
1.8K
Tibo
Tibo@thsottiaux·
What are you building this weekend?
English
2.6K
63
3.3K
515.3K
Can Sar retweetledi
Can Sar
Can Sar@cansar·
@can I interpreted it as 1) making government want to push the breaks more, 2) undermining the case for investing more. I don’t agree with everything here, but enjoyed it overall.
English
1
0
0
72
can
can@can·
@cansar I don’t get how oss is decelerationist at all. First time hearing that case, though im admittedly not in the ai policy circles.
English
2
0
6
504
Can Sar
Can Sar@cansar·
Great post
Dean W. Ball@deanwball

Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.

English
2
0
2
1.9K
Dean W. Ball
Dean W. Ball@deanwball·
Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
English
2K
900
7.5K
10.7M
Can Sar retweetledi
Andrew Guenther
Andrew Guenther@codebrewed·
Do not sleep on Codex for life things. Codex pays my bills. It helps track my budgets. It reads my email. It shops for me. It has booked vacations for me almost entirely autonomously. There is truly just a magic to it that I haven't felt with a piece of software in so long.
English
12
12
456
25.2K
Can Sar retweetledi
Charlie Marsh
Charlie Marsh@charliermarsh·
If you live in the Codex TUI I am here for you
English
29
1
238
17.2K
Tom Brow
Tom Brow@_tombrow·
Which ones am I missing?
Tom Brow tweet media
English
1
0
0
169
Alex
Alex@gpt_alex·
One of my favorite parts about the OpenAI team on Twitter is that one guy keeps talking about fashion/taste, one dude keeps resetting rate limits, and then there's a third guy who keeps posting the games that he made with Codex Vibes seem pretty great overall
English
35
30
1.2K
50K
Can Sar retweetledi
Noam Brown
Noam Brown@polynoamial·
2023: LLMs struggle with 4th grade word problems 2024: LLMs can do high school math 2025: LLMs get a gold medal at the IMO Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn't even news. Where will we be next year?
Edgar Dobriban@EdgarDobriban

AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (faculty.wharton.upenn.edu/wp-content/upl…) and will be on arxiv tonight; supporting code is here (github.com/dobriban/BH).

English
91
292
2.9K
426K
Can Sar
Can Sar@cansar·
@gertig @ajambrosino This is still happening for you post updates? We changed the way they was working for a while but should be back to the setting.
English
1
0
0
67
gertig
gertig@gertig·
@ajambrosino this keeps getting turned off somehow after updates. Don't want me to know?
gertig tweet media
English
1
0
0
212
Can Sar
Can Sar@cansar·
@ebloch You gotta panic sell those bottoms!
English
1
0
1
86