Shrey Modi

365 posts

Shrey Modi

Shrey Modi

@ShreyModi13

Infra @FireworksAI_HQ, https://t.co/LkqyNDPcga CS @iitbombay, @uchicago. Poker. RL and Continual Learning research.

San Franscisco Katılım Kasım 2020
1.7K Takip Edilen391 Takipçiler
Shrey Modi
Shrey Modi@ShreyModi13·
a new wave of 2T+ parameter models? waiting for deepseek's launch. I am especially excited about the 27B parameter model, Qwen has been excellent with their smaller model releases. They're really pushing the frontier there.
Qwen@Alibaba_Qwen

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:github.com/qwen-code-dev-… - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: qwen.ai/blog?id=qwen3.8 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.… ⚡ API: qwencloud.com/models/qwen3.8…

English
0
0
5
285
Shrey Modi
Shrey Modi@ShreyModi13·
I wanted to check out the state of post-training auto-research so I analysed model traces on PostTrainBench (@posttrainbench ). Models are given 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. 1. Fable 5 is the leader with a 40% benchmark score as compared to human trained instruct models standing at 50%. 2. The benchmark tests post-training on famous benchmarks like AIME, GPQA etc. The ideal dataset mix and training recipes for which are already published in papers online, most of the models directly used these datasets without having to curate their own dataset. They used already human-researched training recipes and just implemented those. This is not going to be the case on unknown benchmarks that enterprise care about. 3. They spent around 25% of their tokens fighting environment issues which points towards a need for an auto-research harness which abstracts away the training infra complexity. They were also too scared of causing OOMs and didn't know the extend to which they could scale up training. They strongly preferred running over small datasets to get a signal. 4. Models showed a lot of reward hacking to the extent that they even submitted the final instruct models on huggingface as their own post-trained model. To bring auto-research to the enterprise, a lot of work on prevention of reward hacking would need to be done. 5. All models gave up before utilising their time budget. I see the future as humans specifying their $ and time budget and models fully utilising it. 6. Models were really good at implementation and running experiments without making any errors. They also first checked and tried to align the harness of evaluation and training which is a common mistake people run into, especially looking chat templates! 7. GPUs suffered about 20% of idle time which can be improved to allow a higher throughput and lower cost for experiments! Great work by the @posttrainbench team for open-sourcing the model traces! I see auto-research systems coming to the enterprise very soon. The cost of experimentation just has to go down and the certainty of training recipes to work has to increase for enterprises to invest time and money in these systems.
English
2
2
73
5.3K
Raj Jagiasi
Raj Jagiasi@raj_jagiasi·
America is the most ambitious project in human history. Two and a half centuries of pure dominance. Until somewhere along the way we got complacent and things began to fall. The new age is about rebuilding the state, restoring dominance, and helping the country get back in order. The government has a clear vision and the will to change. People behind the scenes work far too hard to help this country win, but are handed tools which are far too slow and old. This was okay in the 80s, but it stands no chance when AI and robotics is moving this fast. @Stratum_Inc is building the software stack for that new age. AI agents that accelerate every agency of the United States government - so the institutions that run the country define the pace at which the country moves. Today, @VihaanJagiasi and I are announcing that we have joined @ycombinator to help America accelerate into the future. Thanks to @snowmaker @raphaelschaad and @ycombinator for taking a bet on us. This is the work of the next century and we're just getting started. God bless America.🇺🇸🦅
Raj Jagiasi tweet mediaRaj Jagiasi tweet mediaRaj Jagiasi tweet mediaRaj Jagiasi tweet media
VJ@VihaanJagiasi

x.com/i/article/2082…

English
4
3
16
2.4K
Hemant Mohapatra
Hemant Mohapatra@MohapatraHemant·
Announcing a San francisco office and great to bring in @dchaplot as a hands on advisor to help @SarvamAI build frontier models from India.
Hemant Mohapatra tweet media
English
25
62
1.5K
69K
Kevin Jiang
Kevin Jiang@jiang_kev·
We just raised $5.7M for @PolarBrowser, the AI browser that beats Anthropic and OpenAI on every major web agent benchmark. - 4,500,000+ actions taken for users, automating sales, recruiting, and ops - One company cancelled Clay and saves 25+ hrs/wk per person - Team is from MIT, YC, Prod, Citadel, Jane Street, Perplexity, Modal, and Apple 100 hours of work. From 30 seconds of typing. Download the world's most powerful AI browser: polarbrowser.com
English
230
119
1.4K
978.6K
Shrey Modi
Shrey Modi@ShreyModi13·
I was curious to figure out what the last 20 years of IOI gold medalists are doing right now. So I pulled all 440 golds (2006–2025) and tracked them down. Two things stood out: 1. Only 35 competed for the US, but 141 now live there. 2. AI has caught up to quant. ~39 are now at AI labs (OpenAI, DeepMind, Anthropic, Cognition, DeepSeek) vs ~29 in trading (Jane Street, HRT, Citadel, Jump). A decade ago the default was a hedge fund now it's a model lab.
Shrey Modi tweet media
English
2
0
28
5.1K
Molly O’Shea
Molly O’Shea@MollySOShea·
BREAKING: Devin now writes ~95% of Devin's code Scott Wu (@ScottWu46), CEO of @cognition says the abundance era of AI is real & coming soon.. FULL INTERVIEW Cog Stats: › $2.5B raised, $26B valuation › $500M+ run rate, up from $37M LY › Customer usage is up 11x-12x in 6 months › Total code shipped inside Cognition has grown 7x › Devin writes ~95% of Cog's code, up from 89% at May Series D › The entire Windsurf acquisition was negotiated over 1 weekend, ~40 people buying a 200 person company with $82M of ARR › Goldman Sachs, Mercedes-Benz, Citi, Dell, Santander, NASA, the US Navy & the US Army all run Devin R I P : "The tokenmaxx era lasted from January to May of 2026, roughly." "I'm excited to have AI that allows me to spend all of my time & focus on the things that I care about & take care of the rest. I think the abundance era is real, & I think we'll be there pretty soon." Recorded 7 July 2026. On 23 July 2026 Cognition acquired Poke maker The Interaction Company 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Scott Wu, Co-Founder & CEO at Cognition (01:00) The Identity crisis of building Devin (01:42) How Cognition 7X-ed its output in 6 months (02:55) The Metric every AI company gets wrong (04:18) Cognition's 12x growth curve (07:46) The Windsurf acquisition (13:26) Scaling overnight: Cognition's hardest week (15:11) The strategy behind the integration of Cognition products (17:23) Cognition's M&A playbook (20:01) Why the DMV is Cognition's next big target (22:59) Why Cognition bet against AI centralization (25:09) Why enterprises refuse to bet on a single AI lab (28:32) Scott's honest take on Superintelligence (32:36) Is Silicon Valley losing its mind over the AI gold rush? (33:26) Devin's origin story (37:17) The question Scott thinks nobody's asking yet (39:59) Poker with Peter Thiel & Napoleon ? (41:24) Scott's vision for the next year
English
13
38
258
669.3K
Shrey Modi
Shrey Modi@ShreyModi13·
what are some good benchmarks that test ML auto-research capabilities of a model other than PostTrainBench?
English
0
0
11
1.6K
Shrey Modi retweetledi
Alex
Alex@avrecum·
Maybe in the future one way to do robotics data collection will be to ask a frontier VLM to slowly and carefully tweak controls and look at the result, then speed up + smoothe that trajectory and train on it, thus turning test time compute into new data
Jay Chooi@chooi_jeq

We gave Opus 5 robot arms. It can stack bowls.

English
2
5
12
1.8K
Shrey Modi
Shrey Modi@ShreyModi13·
Just as manufacturers like potato chip companies hedge commodity prices in the commodity market, enterprises will soon start hedging compute costs in the compute market. compute derivatives is going to be huuuge!
English
0
0
3
213