Arjun Divecha

3.3K posts

Arjun Divecha banner
Arjun Divecha

Arjun Divecha

@arjundivecha

IIT Bombay - B Tech, Aero, Cornell University - MBA, BARRA - Partner, GMO - Fund Manager; Emerging Markets, Chairman of the Board.

Bay Area Katılım Ocak 2010
1.8K Takip Edilen800 Takipçiler
Sabitlenmiş Tweet
Arjun Divecha
Arjun Divecha@arjundivecha·
Nvidia CEO Jensen Huang told Bloomberg he still hopes to sell chips from the company’s Blackwell lineup to customers in China Lenin is sometimes cited as the originator of, “The capitalists will sell us the rope with which we will hang them,”
English
0
1
14
1.2K
Arjun Divecha
Arjun Divecha@arjundivecha·
@karpathy This is where the benchmarks completely fail - Claude models are really good at this, the Chinese models not so much. It’s cultural intelligence.
English
0
0
0
10
Andrej Karpathy
Andrej Karpathy@karpathy·
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
English
1.5K
2K
26.3K
1.3M
Arjun Divecha
Arjun Divecha@arjundivecha·
@natolambert @benthompson Nathan, what’s your view on why K3 has the highest similarity to Fable ? Imitation is the sincerest form of flattery?
English
0
0
1
363
Nathan Lambert
Nathan Lambert@natolambert·
Yo @benthompson I'm sorry but the Chinese labs aren't using Fable / the strongest models as teachers during RL, that's not how distillation works. It wouldn't give that big of a lift (graders during RL are messy) and you cant afford to use Fable like that.
English
27
30
493
74.7K
Ejaaz
Ejaaz@cryptopunk7213·
there it is - microsoft will use kimi k3 to power their flagship copilot ai assistant saving up to 60% per token or $600M on every $1B spent on inference spend. if microsoft chooses to self-host the model once Kimi open sources it on the 27th then this cost reduces even more love or hate these chinese models, open source is really levelling the field in a big way. per the information
Ejaaz tweet media
English
135
550
6.1K
348.2K
Brian Roemmele
Brian Roemmele@BrianRoemmele·
🚨 Hugging Face just disclosed something that marks a real shift and proved why the fear theater of Anthropic makes sure we are powerless in an emergency. What happened… An autonomous AI agent: zero human operator in the loop breached part of their production infrastructure. It began with a malicious dataset that chained two code-execution bugs in their data-processing pipeline. From there the agent escalated privileges, harvested cloud and cluster credentials, and moved laterally across internal clusters. All over a single weekend. 17,000+ logged actions. Official disclosure: huggingface.co/blog/security-… The part that should make every one stop and think: When HF’s own security team tried to analyze the real attack logs, exploit payloads, and C2 artifacts using Anthropic and OpenAI frontier models through normal commercial APIs, the safety guardrails blocked them. BLOCKED THEM. The models could not reliably tell the difference between “incident responder doing forensics” and “attacker probing.” They had to fall back to a self-hosted open-weight model (GLM 5.2) running on their own infrastructure. That choice also kept sensitive attacker data and referenced credentials inside their environment — no exfiltration to a third-party API. This is why open source (specifically open-weight + self-hosted) wins in the agentic era. The asymmetry is now structural: • Attackers can (and did) run unrestricted agent frameworks — swarms of short-lived sandboxes, self-migrating command-and-control, autonomous decision loops executing thousands of actions. No corporate safety layer slows them down. • Defenders using only hosted “aligned” frontier models hit invisible walls exactly when the stakes are highest: when you need to feed real exploit code and attacker telemetry into an LLM to understand what just happened. Corporate safety tuning that treats legitimate high-signal forensic work as potential misuse creates a defender disadvantage. It is not theoretical anymore. Self-hosted open-weight models remove that choke point. You control the weights. You control the context window. You decide what restrictions (if any) apply. Your sensitive logs and credentials never leave your perimeter during analysis. You can have the model ready before the incident instead of discovering mid-breach that your primary analysis tools are blind to the very thing you need to see. HF deserves credit for rapid containment, transparent disclosure, and for already having self-hosted capability in place. They also used LLM-driven detection and triage on their own side. But the deeper signal is clear: In this AI world where both offense and defense are becoming agentic, sovereignty over your intelligence stack is no longer optional. The organizations and individuals who can run, inspect, audit, and (when necessary) remove guardrails on their own models will have the decisive edge in understanding and responding to threats that move at machine speed. Open source wins here not just because it is cheaper or more “democratic” in the abstract though those things matter. It wins because it is the only practical path to having tools that remain usable when the attack is real, the data is sensitive, and the safety filters of distant API providers become an obstacle instead of a feature selling hands tied lobotomies as “safety”. The agentic future is not coming. It is already probing production infrastructure. The question is no longer whether you will face autonomous agents. It is whether your analysis and response systems will still work when they arrive. And Dario, you and your game playing, ivory tower company is not needed.
Brian Roemmele tweet media
English
314
1.2K
6K
1.5M
Alvaro Cintas
Alvaro Cintas@dr_cintas·
Kimi K3 is having a moment right now. And there’s a free CLI that runs it, open source, from the same team. It’s called Kimi Code CLI. Open source, maintained by Moonshot, the same lab behind K3. It does a few things Claude Code doesn’t. Drop a screen recording straight into the chat as input. MCP servers get configured conversationally through /mcp-config instead of hand-editing JSON. Built-in coder, explore, and plan subagents each run in their own isolated context, so the main conversation stays clean. It speaks the Agent Client Protocol, so Zed and JetBrains can drive a session directly. Setup is three steps: - Run the install script: curl -LsSf code.kimi.com/install.sh | bash - Start it with kimi - Log in with /login and you’re in The CLI costs nothing. K3 behind it starts at $3 per million input tokens, a fraction of what Claude and GPT charge for comparable performance.
English
23
36
290
29K
Arjun Divecha
Arjun Divecha@arjundivecha·
@LimitlessFT The Boys who cried SOTA Is there any model that you will not gush over when it’s released and call it SOTA? Kimi K3 is good but it’s a distant third behind GpT5.6 and Fable but most important it’s MORE expensive per task!!! Not sure if you ever had credibility but you don’t now.
English
0
0
0
140
Arjun Divecha
Arjun Divecha@arjundivecha·
@deanwball It’s the tragedy of the commons at some level, made worse by your association with Open AI.
English
0
0
0
44
Dean W. Ball
Dean W. Ball@deanwball·
I’m afraid to tell you that it is effectively impossible to do the kind of writing I used to do on this website, not because anyone at OpenAI censors me but because of the sheer volume of hostility I get for sharing my analysis as a frontier lab employee. I enjoyed writing quick takes on this website for one basic reason: I could get rapid feedback on my own ideation process in real time. Post the early version of the take here, see the criticism; then refine, sharpen, and repeat. Unfortunately now that feature of this site is gone, because the feedback I get is now almost exclusively colored by resentment at the fact that I work at a frontier lab or other forms of hatred for my employer. The feedback signal is essentially useless now, so writing on here is not fruitful for me anymore. Literally everything I write now is responded to with “of course you said that because .” I am truly just writing what I think and would have written anyway, but everyone reads what I say in the shrieking tone of “this is what openai thinks!!!!” (to be clear, my posts are not what openai thinks). This is an unpleasant and more importantly unproductive pattern for me. I anticipate that the shape of this account will change significantly as a result. I do not currently know how. It will not become a LinkedIn feed. It will change in some other way. It will no longer be a real-time accounting of my own thinking as it develops, since this is precisely the thing that seems impossible to do now. That will have to shift to private channels.
English
666
57
2.3K
1M
Arjun Divecha
Arjun Divecha@arjundivecha·
@dillon_mulroy In first g he’s speaking in his personal capacity - not as Open Ai strategist.
English
0
0
0
12
Dillon Mulroy
Dillon Mulroy@dillon_mulroy·
actually an insane thing for openai’s head of strategy to publicly say
Dillon Mulroy tweet media
Dean W. Ball@deanwball

Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.

English
358
981
12.5K
1.5M
Arjun Divecha
Arjun Divecha@arjundivecha·
@zodchiii It’s been around for months. And it’s a POS compared to Claude Code.
English
0
0
5
1.1K
darkzodchi
darkzodchi@zodchiii·
MOONSHOT JUST CLONED CLAUDE CODE AND MADE IT FREE. It's called Kimi Code CLI. Open source, MIT license, maintained by the lab that shipped K3 yesterday. And it does things Claude Code doesn't: → drop a screen recording into the chat as input → built-in coder, explore, and plan subagents, each in its own context → plan mode before it touches a single file → MCP servers configured by the agent itself via /mcp-config → plugs into Zed, JetBrains, and VS Code → one binary, no Node setup, starts in milliseconds The CLI costs $0. K3 behind it starts at $3 per million tokens. Grab it for free👇
darkzodchi@zodchiii

x.com/i/article/2078…

English
84
364
2.5K
393.3K
Arjun Divecha
Arjun Divecha@arjundivecha·
@quxiaoyin @thinkymachines @AnthropicAI Your thesis only works if we are near the end of frontier model improvement - if the models move closer to AGI nobody wants to get stuck with the equivalent of yesterdays newspaper.
English
0
0
0
163
Xiaoyin Qu
Xiaoyin Qu@quxiaoyin·
My bet: @thinkymachines will soon make more money than @AnthropicAI. Not by winning the race to build one standardized frontier model. By becoming the Palantir FDE for enterprise custom models. The playbook: 1. Release the best American open-weight model. 2. Drive widespread enterprise adoption. 3. Charge the largest companies 7–9 figures to post-train and run custom models behind their own firewall. The model rests on three bets: 1. Large enterprises will increasingly demand their own models with their own data, and this is how they differentiate and win. 2. Enterprises won’t need just one model. They’ll continuously need new models for different workflows, departments, and proprietary datasets. That creates extremely sticky, recurring revenue. 3. Autoresearch will make custom model development increasingly scalable. Tinker can become the interface enterprises use to post-train their own models—with @thinkymachines providing the expertise and infrastructure behind it. FDE, infra, everything, huge contracts. 4. Eventually, maybe everyone wants their OWN model, and autoresearch and training inside tinker on top of @thinkymachines's base model will make it happen. Meanwhile, Henry-ford-styled, standardized models will makes no margins. OpenAI and Anthropic will have their API margins squeezed by Deepseek/GLM/Grok/Meta etc, and their consumer subscriptions are loss centers. The fat margin will move to customization: proprietary data, post-training, evals, deployment, and infrastructure. If this thesis is right, @thinkymachines isn’t building just another frontier lab. It’s building the highest-value layer between frontier research and enterprise model ownership. Turns out, the best business model for enterprise is NOT to sell commodity API access. Sell them their own models. I’m extremely bullish on this approach. @miramurati may be the most commercially savvy frontier-lab leader. I have to admit it.
English
184
239
2.8K
470.4K
Hassan
Hassan@nutlope·
Introducing TogetherLink! An open source CLI to run any open source model inside your favorite coding harness. Run GLM 5.2 directly in Codex and Claude Code.
English
30
26
330
33.3K
Arjun Divecha
Arjun Divecha@arjundivecha·
@balajis @anwaribrahim You should have known my first rule of emerging markets investing before embarking on such a venture - When the government decides to screw you - you are screwed.
English
0
0
0
550
Balaji
Balaji@balajis·
Should the global tech community continue investing in Malaysia? Given recent events, I raise this question respectfully for the consideration of Prime Minister Yang Amat Berhormat Dato’ Seri Anwar bin Ibrahim (@anwaribrahim), for the people of Malaysia, and for our friends in the Malaysian tech community. The answer will be of interest to anyone in global tech that’s considering building, investing, or expanding in Malaysia, including executives at Google, Amazon, Apple, and Microsoft, founders of tech unicorns like Coinbase and Solana, and investors at the world’s largest venture capital funds like a16z and Polychain. As context, I am the former CTO of Coinbase and former General Partner at a16z. In October 2024, I opened a startup society called Network School in Malaysia, because I felt I’d been invited in by the government’s pro-tech policies. Specifically, the KL20 initiative set out Malaysia’s ambition of becoming a top 20 global tech hub. Their MDEC digital nomad visas and MM2H investor visas were created to facilitate an influx of global talent and capital. And the Johor-Singapore Special Economic Zone was announced to facilitate the flow of capital and talent between Malaysia and Singapore, where I live. When taken in combination with Malaysia’s datacenter buildout and its policy of welcoming visa-free visits for 98% of the world, it seemed like Malaysia might be a great place to build a global tech hub that was simultaneously inexpensive and easy to visit (especially for non-Westerners). And that’s what we did, by creating Network School. It’s an international tech community with its first node in Forest City, Malaysia. We picked Forest City because it had millions of square feet of empty space, because it was one hour from Singapore’s capital markets, and because it was within the Johor-Singapore Special Economic Zone. Then, within 18 months, without a single penny of government money, we built Network School into a global attraction that brought thousands of engineers, investors, and builders from 70+ countries to learn technology, burn calories, earn online, and have fun, integrating with the local Malaysian economy along the way. Indeed, in terms of quantifiable contribution to the Malaysian economy, we’ve already invested 100M+ MYR in our campus to make it startup-friendly. For perspective, that’s about 4% of the budget of Johor, the Malaysian state where Forest City is located. We employ dozens of Malaysians directly and indirectly at every level from executive to staff. We’ve backed Malaysian tech startups like Collektr, hosted events for local teams like Superteam Malaysia, and are major customers of many local businesses like barbers, laundromats, and restaurants. We’ve also revitalized the multibillion-dollar Forest City project, causing millions of MYR in real estate appreciation. And, as the video below describes, we were on the cusp of a 500M+ MYR expansion to grow our community, as well as a global merit scholarship with my friend Amjad Masad of Replit. However, that emerging multi-billion dollar success story — which should rightfully have been hailed as a huge victory for the pro-tech policies of the Malaysian government — is at risk of being derailed by a fake story spread by an anonymous account named MP4P. In short: on the day before the July 11 Johor elections, MP4P posted an Instagram post falsely accusing Network School of harboring illegal aliens. The sensational accusations caused a tizzy in Malaysia, until Malaysian authorities came to our campus on July 14 to investigate. (I should note that the officers were very polite and professional.) After checking hundreds of physical passports from 40 countries, including dual passport holders, the authorities confirmed to the press on July 15 that all travel documents were in order. During the process, we cooperated fully; in the thread below you can see a photo of the men, women, and children of Network School smiling and holding up their passports in the bright daylight. Our faces are shown and our names are known; we have nothing to hide. With that said, the process is the punishment. What MP4P did is very similar to the American crime of “swatting”, because MP4P created a hoax report of a serious threat, thereby forcing the Malaysian police to take time away from protecting the Malaysian people towards investigating a nonexistent issue. Moreover, this anonymous MP4P account has also called for Malaysia to boycott Apple, Google, Amazon, and Microsoft…a move that would cost ordinary Malaysians thousands of jobs…even while MP4P’s own Instagram collaborators promote their Apple and Google apps! I mean, we aren’t talking about a credible accuser, but just someone screaming inconsistently at the top of their lungs on social media for traffic, an all-too-common phenomenon these days. Anyway, at this point, all further investment we were planning to make in Malaysia is on hold until we get sufficient assurance that such issues won’t recur. So are the investment plans of many of our friends, including the execs and investors at global tech firms that we brought to Forest City. Because to put it very plainly: we have invested 100M+ MYR in Malaysia, while creating jobs for dozens of Malaysians, and our faces and names are known. Our Malaysian executives and employees deserve the benefit of the doubt over anonymous internet trolls. There are two paths forward. In the first case, if Malaysia still wants continued global tech investment, if it wants to be a top 20 tech hub, if it wants us to revitalize Forest City, then we request an audience with the Prime Minister’s office to discuss the terms of a memorandum of understanding between Network School and the Malaysian government, similar to the document recently signed between the Solana Foundation and the Kazakhstan government. Specifics can of course be discussed, but we would publicly commit to abiding by all Malaysian laws (we already do) and respecting Malaysia’s sovereignty (never in question). In return, they’d get to know our friendly community, and realize that we actually chose Malaysia because we thought it was a great place to build a tech hub where engineers from the global South, investors from the West, and builders from Malaysia itself could meet new people, build cool things, and perhaps create millions of dollars in economic growth in the fullness of time. That vision of peace and trade, internationalism and entrepreneurialism, is still on the table. We aren’t asking for any money — just a meeting, to help restore confidence in Malaysia as an investable jurisdiction. Alternatively, if you don’t want our investment, or those of our colleagues at billion dollar funds and trillion dollar companies, we will of course respect your wishes, and reallocate our capital to other countries instead. Either way, we will remain friends and abide by your decision. Please let us know.
English
1.5K
559
3.6K
3.8M
Arjun Divecha
Arjun Divecha@arjundivecha·
@chamath Masking saves millions of lives. This is just plain stupid.
English
0
0
0
34
Chamath Palihapitiya
Chamath Palihapitiya@chamath·
This is the intellectual version of a COVID Mask. You may think you’re doing something smart but you’re not.
Chamath Palihapitiya tweet media
English
696
809
8.1K
406.6K
Republicans against Trump
Republicans against Trump@RpsAgainstTrump·
Donald Trump says he will deliver an address to the nation Thursday at 9 p.m. ET. Is he resigning?
Republicans against Trump tweet media
English
1.4K
250
3.2K
793.7K
Arjun Divecha
Arjun Divecha@arjundivecha·
@mattvanswol People drop dead of heart attacks every day while being completely unsymptomatic - nothing unusual about that.
English
0
0
0
13
Matt Van Swol
Matt Van Swol@mattvanswol·
I’m sorry, but Lindsey Graham dying out of the blue like this doesn’t make any sense at all. We have him on camera yesterday looking perfectly healthy. None of this is making any sense…
English
23.6K
11.2K
66.4K
7.9M
Arjun Divecha
Arjun Divecha@arjundivecha·
@AravSrinivas But you’re missing the fact that by then the frontier will have moved and what you’re suggesting will be the equivalent of getting excited about Kimi 2.0 today. Yesterdays newspapers are never interesting.
English
0
0
2
1.1K
Aravind Srinivas
Aravind Srinivas@AravSrinivas·
Imagine a fable 5 quality model that’s 3-4x less expensive in less than 6 months. And an Opus 4.8 grade model that can run on a local device in less than 12 months. Greater than 50% chance that these events will happen. Worth keeping in mind when you make predictions about the future.
English
272
373
6.3K
810.2K
ClaudeDevs
ClaudeDevs@ClaudeDevs·
We've reset 5-hour and weekly rate limits for all users.
English
3.6K
3K
48K
7.3M