
Devansh
80 posts

Devansh
@devanshtandon_
AI research & product @meta tbd, @instagram | before: @GoogleDeepMind @YouTube @Google news/discover, search ranking, ads | CS & Econ @yale
san francisco Katılım Ekim 2017
2.2K Takip Edilen645 Takipçiler

@realmadhuguru cloud margins reselling kimi tokens will be much worse than gemini.
ideally a strong gemini/model layer enables google to capture value at the app layer with the best margins
without a competitive model, reselling kimi falls back below model layer to the cloud/infra layer
English

note: benchmarks don’t tell the whole story, i used artificial analysis for a clear frontier chart but the real performance is jagged – and depends on the applications being considered. Coding / long horizon benchmarks follow similar trends.
from vibe checks, i’m most impressed by Kimi K3’s grit & persistence to keep going on longer horizon tasks. the model is quite verbose and could be more token efficient.
English

@SemiAnalysis_ @Kimi_Moonshot US/China frontier chart based on Artificial Analysis: devansh-kimi-frontier.netlify.app
English

@GavinSBaker Kimi K3 closed the US/closed v China/open frontier gap from 6 months behind to 1 month.

English

Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.
Rationale:
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.
Time will tell on both points. And likely fairly quickly.
Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.


English

note: benchmarks don’t tell the whole story, i used artificial analysis for a clear frontier chart but the real performance is jagged – and depends on the applications being considered. Coding / long horizon benchmarks follow similar trends.
from vibe checks, i’m most impressed by Kimi K3’s grit & persistence to keep going on longer horizon tasks.
English

You see, what fascinates me about “remember to not smile” is that he needs his pain to be witnessed. He suffers at 4am, okay, fine, many people suffer. But then, immediately, he makes a slide! He gathers the entire company! He writes an article! Is it not clear what is happening here? Suffering without an audience simply does not count. This is ideology at its purest.
We see here the old Protestant trick in its Silicon Valley form: success alone is vulgar, almost pornographic. My god, you cannot simply succeed. It must first be purified through unnecessary suffering. So you get this obscene reversal where pain stops being a cost of the product, and instead becomes the product itself. The developer documentation is, how do you say, merely the byproduct, the excrement of the true production, which is the suffering. As Hegel already knew, and here I think even Lacan would agree, although perhaps not, the true obsession is always with deserving success.
Here I must tell you this old joke from Soviet Union. A worker says: "we pretend to work, they pretend to pay us.” Now here it is inverted, which I claim is much worse. They are paid perfectly well, wonderfully well, and so they must pretend to suffer. Do you see the perfect madness of this?
And now, the ultimate obscenity. Imagine the company succeeds while everyone leaves at 5pm. Catastrophe! Total ontological catastrophe! In this worldview, I am tempted to say, failure would almost come as a relief. Failure means you simply did not suffer enough, go back, suffer more. But success at 5pm? This cannot be explained.
So when he says "remember to not smile," this is a desperate command of the superego. If you smile, we will all discover the pain was never necessary.

Han Wang@handotdev
English

good model, low price, will be powering llm-native recommender systems soon
Mark Zuckerberg@finkd
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
English


my paper won an award at icml 😁
some thoughts:
• this work was rejected from NeurIPS. i cleaned it up a small amount and it got great reviews from ICML! don't give up
• ICML received 24k submissions and only gives out 7 awards, which is crazy. feeling grateful
• i distinctly remember sitting at my desk two winters ago wondering if i would ever finish this project. most of all this is the product of sitting down and forcing myself to keep working for several months straight. the results emerged from running the experiments over and over and fixing a long sequence of tiny details. eventually, the curves looked like that 👇
• also happy that the insights in this paper are becoming more widely accepted: 3.3 bits/param, thinking about capacity "LLM as flashdrive" mentality
• the method here is used successfully for selecting midtraining data at least one frontier lab, which is cool!
• i am grateful to my collaborators, but Meta is no longer a great place for academic research imo and this almost never got published for a number of reasons. i shall not elaborate further
• for future work, i think analyzing the implications of on-policy algorithms on capacity, as well as LoRA and things like it, are fruitful potential research directions
• sadly i'm not in Korea but am following the conference online from california and happy to chat!
a nice end to one phase of my research career :)
Jack Morris@jxmnop
new paper from our work at Meta! **GPT-style language models memorize 3.6 bits per param** we compute capacity by measuring total bits memorized, using some theory from Shannon (1953) shockingly, the memorization-datasize curves look like this: ___________ / / (🧵)
English

I'm speaking @aiDotEngineer today, for the LLM-RecSys track.
The biggest consumer apps in the world are content feeds (Instagram, YouTube, Facebook, TikTok), and recommendation model improvements drive most of their engagement & revenue growth.
These recommendation systems are being transformed and rebuilt into LLM-native recommendation systems. I'll share frameworks + a recipe to train LLM-native recs models. LLM-native models enable new product experiences: steerable & interactive recommendations that put users in control of their algorithm.
LLM-Recs is the largest application of AI for consumer apps, and remarkably token efficient when compared to AI chat apps.

English

WhatsApp may be the statistically best app we've ever made.
It's a product manager's dream with its unparalleled addictiveness (DAU / MAU) of 87% and stickiness (M1 retention) of 86%, both #1 in the world while having the #1 most monthly active users for any non preinstalled app of 2.7B users.
Here are the top 25 most used apps in the world by MAU on both these metrics. Some surprising observations:
— There are now 15 1B+ user apps in the world, 8 by Google, 4 by Meta.
— The 3 that aren't are TikTok, Telegram (!) and ChatGPT
— Telegram has more users than Spotify, Pinterest, Netflix, Amazon, Snap!
— ChatGPT's one month retention is #5 after WA, Instagram, Chrome and Youtube. 2yrs ago, it was been #20 by M1 retention
— Shopee, a shopping app in southeast Asia, is huge and retains users better than Amazon!
Useful way to break down consumer businesses especially within certain categories (X vs Reddit vs Threads is a good one). It's shocking how few new apps have been able to break through in the past 10yrs.

English

