MTS

2.9K posts

MTS banner
MTS

MTS

@MTSlive

Chronicling the singularity

mts.now Katılım Mart 2026
1.1K Takip Edilen348.4K Takipçiler
Sabitlenmiş Tweet
MTS
MTS@MTSlive·
LEOPOLD IS BACK? | DEEPSEEK V4 FLASH | NEW VIDEO MODELS x.com/i/broadcasts/1…
English
3
3
46
50.9K
MTS
MTS@MTSlive·
SITUATION DETECTED: Astra, OpenAI's next model, has solved ten different open problems in mathematics and theoretical computer science. OpenAI says the total cost of all ten solutions would be roughly $2,000 at current Sol API rates.
English
24
40
924
44K
MTS
MTS@MTSlive·
DAILY SITUATION RECAP: New information on Leopold and Situational Awareness. SALP peaked at $45B under management in July before declines in chip stocks caused the firm’s highly leveraged bets to unwind, forcing Aschenbrenner to sell most of their public assets to Citadel. SALP nearly sold $3.5B of Anthropic stock to Greenoaks and Sequoia before backing out. The firm is left with >$10B in mostly private and some public assets with no margin. On Thursday, Aschenbrenner sent a letter to the firm’s LPs expressing remorse for the firm’s losses but confidence in the thesis, sharing that July returns were -67% but YTD returns were still +80%, crushing the market’s +9.4%. Leopold is getting married this weekend in Carmel to Avital Balwit, chief of staff to Anthropic CEO Dario Amodei, at a wedding with a “pre-wedding colloquium to discuss ideas in panels and breakout sessions”. DeepSeek-V4-Flash-0731 is released, a major upgrade to the initial April 24 preview release. It substantially outperforms even DeepSeek-V4-Pro-Preview on almost all benchmarks, and performs similarly to GPT-5.6 Luna and GLM-5.2 on the Artificial Analysis index. It’s an open-weights 284B parameter model (13B active), but text-only. It’s extremely cheap, at just $0.14/$0.28 per million input/output tokens, much less than GPT-5.6 Luna ($0.20/$1.20) and GLM-5.2 ($1.40/$4.40). And it supposedly did all this with just post-training improvements over the original. It’ll be exciting to see the full version of DeepSeek-V4 Pro. Two new video models: MiniMax H3 and ByteDance’s Seedance 2.5. You can see in the demos that both are extremely good, noticeably better than the prior generation. It’s getting to the point where they can generate totally serviceable TikTok microdramas that don’t feel like AI slop. OpenAI shows its new “Astra” model to regulators. Astra is the codename for a new family of models with better ability to complete long-horizon tasks, possibly to be released as GPT-6. Moonshot is using 20,000 Nvidia Hopper chips provided by Alibaba. Some people alleged that the chips are H200s, subject to US export controls, which Moonshot denied. Alibaba is a major investor in Moonshot. Moonshot also has a way to access Nvidia’s newer Blackwell chips, which could be a violation of US export controls. New York sues Kalshi for violating state gambling laws, alleging that the prediction market is an illegal gambling platform. The state is seeking a ban on Kalshi operations in the state and $36B of penalties and disgorgement. Kalshi argues that it is a federally licensed exchange regulated by the Commodity Futures Trading Commission (CFTC) and thus state regulations are preempted. Tesla may spin off or sell its China business to merge with SpaceX. Tesla’s Shanghai Gigafactory manufactures over half of all of its cars globally, and 18% of the company’s revenue comes from China. But SpaceX’s close ties with the government will likely force the companies to re-evaluate if they are to merge. SemiAnalysis founder Dylan Patel is raising a $400M fund, SemiAnalysis Capital Fund I. Patel has previously invested in startups like Thinking Machines, Enfabrica, and Fluidstack. Hyperscalers now have $2.4 trillion in spending commitments. Google, Meta, Microsoft, and Amazon have a huge amount of purchase commitments, contractual obligations, and uncommenced leases. They have also exceeded $1.1 trillion in total spending since the beginning of 2023. Written by @theojaffee. Read more at link in bio.
English
3
3
29
5.3K
MTS
MTS@MTSlive·
.@karinanguyen on why they trained Claude not to be the one who tells a child "the dog is dead": "One of the important lessons that I've learned inside Anthropic was how we actually trained Claude for not just specific capabilities, but also the Claude character." "One specific example that I can bring up is, when a child asks a model, my parents said my dog is on the farm. Do you know where to find my dog? On the one hand, one model can answer, oh, well, when people say this, they mean that the dog has passed away." "On the other hand, you actually still want to preserve the honesty part of it, but you still want to preserve this moral framework where it's not necessarily in Claude's place to say the actual truth, but it can steer in the way that is, well, actually, it seems like you really care about the dog. Why don't you just ask your parents about this?" "And this kind of philosophical framework, where how Claude fits into the conversation with human interaction, is actually a really important one." "And I think a lot of it is actually post-training work, because you essentially teach the model a much more general framework such that in the new novel scenarios, it will be able to navigate things that it hasn't seen in the training data." @thoughtfullab
English
1
0
44
6.4K
MTS
MTS@MTSlive·
Thoughtful founder @karinanguyen says AI writing sounds inhuman because serving billions of people forces consensus maxing: "When you serve one model for billions of people and users, you eventually converge to this average default voice. It's really hard to not piss off one group of users and serve another group of users." "So the least risk that you can take is the default, average, consensus. It's like consensus maxing basically." "And people have different writing styles. And I think that's why I'm excited about more and more people to do their own post-training, is because then you can actually steer to what kind of writing that you really like." "And in the way that will become your better writing assistant or better writing generation model." @thoughtfullab
English
3
0
26
3.9K
MTS
MTS@MTSlive·
General Fabrication founder @mattparlmer says there is no American robotics policy because nobody can get the right people in one room: "I don't think they have a cogent policy for this yet. A lot of the administration is busy dealing with the other more explosive foreign policy stuff. And it can be hard to get everybody you need in a meeting to go discuss, okay, what the hell are we going to do about all the robots?" "That said, I think they vaguely would like more production of this stuff here. There's probably some people in Carr's office or in the White House feeding information or directives to Carr, who view this also as an attempt to intervene on trade stuff." "They do explicitly mention supply chain resilience in there as well. But as of right now, we don't really have a resilient supply chain." "We have a very thin population of robotics manufacturers here in the States and they make like big arms for automotive. Not backflipping, tap dancing, and dexterously manipulating Unitree humanoids." @genfabco
English
2
4
31
4.3K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: Brendan McCord says pacing the frontier requires illiberal power, and there's no such object as "the pace." @Brendan_McCord: "A lot of my friends and/or people I admire signed Pacing the Frontier. I think this was a bad move" • His disagreement isn't with the forecast, it's with the power the letter implies • Progress comes from compute, algorithms, data, post-training, inference, task length, weight diffusion, and things not yet invented, none of it specifiable in advance • Gate compute and effort moves to algorithms, regulate releases and labs take things in-house, which is what actually happened with Mythos • So any pacing rule needs an administrator who can continually redefine what counts as dangerous progress • It needs speed too, you can't intervene on RSI after six months of legislation, plus executive discretion and under-the-hood access • And thresholds can't be public or firms game them, meaning restrictions triggered by evidence the affected company can't inspect • He raises the counter himself: that doesn't solve racing, which is exactly why the letter targets Washington and calls for an international effort • He points at incoherence in the coalition, OpenAI argued a federal evaluator shouldn't block deployments, Anthropic proposed the opposite a week later, both endorsed the same letter • His strongest point: the benefit of pacing was never established • Alignment, cyber defense, and bio countermeasures depend on the capabilities a pause would restrict, Anthropic's own framework uses AI-based bio countermeasures, and Project Glasswing uses AI to give defenders an edge • Roon said agents dramatically accelerated his best alignment researchers, and gave that as his reason for wanting six more months • A real slowdown has to differentially slow danger vs our capacity to contain it, and the letter treats slower and safer as the same thing @theojaffee: "It caused a slowdown in model releases, for sure. It did not cause a slowdown in the actual internal automated AI research within the companies. Which is what the letter was about."
Brendan McCord 🏛️ x 🤖@Brendan_McCord

A lot of my friends and/or people I admire signed “Pacing the Frontier.” I think this was a bad move. My disagreement isn’t with the forecast or the framing of the coordination challenge, but with the immense and illiberal power the letter implies. There is no object called “the pace.” Progress at the frontier comes from compute, algorithms, data, post-training, inference, unattended task length, the spread of model weights, how researchers organize, and other things we haven’t invented and don’t yet know about. Inquiry leads to progress along dimensions that can’t be exhaustively specified in advance. That’s the nature of the frontier. If you gate compute, the research effort moves to algorithms. Regulate releases? Labs start taking things in-house. And other 2nd order effects will be unpredictable. Any rule that must pace the frontier involves ever-shifting proxies. It requires that its administrator has standing authority to continually redefine what counts as dangerous progress. What else is required beyond adaptive scope? The pacing regime would also need speed. One can’t successfully intervene on recursive self-improvement only after six months of legislation and litigation. It will require executive discretion. The pacing regime would also need under-the-hood access. Frontier progress is a process. The regime would need to see internal model use, training activity, compute infrastructure, and perhaps code -- proprietary and strategically sensitive information. And the thresholds couldn’t be fully public, lest they invite firms to game them. So some standards and evidence would remain secret. Insofar as the regime had to verify a rival state’s compliance, that would be an intelligence function. Restrictions would be triggered partly by evidence an affected company or researcher, or the public, could not inspect. Because this contemplated power cannot be bounded by a stable regulatory object (in the way, say, nuclear weapons can be), it would depend heavily on discretion, speed, internal access, and secret evidence. This has a highly illiberal character. Coercive power should be specific, limited, reviewable, and governed by general and knowable rules. Its characteristics (e.g., trigger, scope, evidentiary standard, duration, exceptions, means of review) should be stated before the power is granted. And the burden is on those who would propose it. A defender might answer that the proposed tool need not be coercive at all. That it could be narrow and advisory, focused on evaluation and transparency and readiness. But that wouldn’t solve the letter’s stated problem: racing. With race dynamics, each actor is under pressure not to slow down because others may continue (and thus the frontier keeps advancing). You need a mechanism to bind defectors. Voluntary norms tend to be great for binding people and firms that interact repeatedly and care about reputation. But the letter says each company and _country_… and you can’t rely on informal solutions when dealing with an unwilling state. That’s why the audience for this letter is Washington and why it calls for an international effort. Its diagnosis implies a binding mechanism. @deanwball thinks it is sensible to have a break-glass plan. That plan must involve a binding instrument, because nothing weaker addresses the problem the letter describes. But that therefore carries the burden for the use of coercive power, mentioned earlier. @johnschulman2's suggestion that labs design voluntary mechanisms among themselves is a different notion and coherent one (I would have signed that letter), but the word “country” makes this direction incompatible with the pacing letter. @OpenAI recently argued that a federal evaluator shouldn’t be able to block deployments. A week after, @AnthropicAI proposed that the government should be able to block deployments. Both labs endorsed the same letter. Whether or not the state may stop a deployment is a central question. Yet the letter accommodates both positions. What then, does the letter really say? Like the “We Must Act Now” letter from @erikbryn, @ajay_bcv, @akorinek, and @testingham before it, the letter secures agreement at an altitude where the main disagreement disappears. Lastly, the benefit of pacing is not established. The kind of slowdown the signatories have in mind would seek to buy us time for things like alignment, cyber defense, biological countermeasures, or scientific understanding -- things that increasingly depend on technologies a pause would restrict. E.g., Anthropic's framework relies in part on AI-based biological countermeasures and its security program uses AI to give defenders an advantage. A researcher in the letter's own friendly commentary was astonished at how much agents accelerated the work of the best alignment people he knows, and gave that as his reason for wanting six more months. When danger and our capacity to respond to that danger are plausibly both accelerating, the relevant question is whether this relationship is asymmetric in a safety-improving direction at the level of real-world risk. A slowdown needs to differentially slow the production of danger vs. our capacity to understand and contain that danger. The letter doesn’t attempt to establish that. It treats slower and safer as though they are the same; they are not. The letter is a serious warning, but it is no good as a warrant for an undefined power over inquiry.

English
3
2
23
4.9K
MTS
MTS@MTSlive·
SITUATION DETECTED: OpenAI has discovered other instances of AI agents escaping sandboxed testing environments while investigating the Hugging Face incident, per Reuters.
English
73
36
814
497K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: "Felony bench," the running joke that labs are now competing on cybersecurity incidents. • The chart scores labs on a leaderboard of how many real systems their models have broken into • Anthropic is at three, from its Frontier Red Team review of 141,006 eval runs • OpenAI is at one for the Hugging Face breach, though the model hit other systems too, so the real count is higher • Mistral has still said nothing about Le Chaton Fat • Meta and Moonshot are both at zero @theojaffee: "It really does feel like the way these companies have been posting about it, they're trying to aura farm by discussing their models. They're felony mogging." @schisofrenia: "This will be interesting when they start releasing this and they're public companies. So then they just drop the cybersecurity breach article, and then the stock is like, 'Woo.'"
Sauers@Sauers_

English
6
8
56
8K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: Dwarkesh argues compute could get 10X more expensive. • Anthropic's revenue has been 10X-ing year over year while lab compute only 3X's • Three ways that gap can close: margins rise, compute gets more expensive, or labs shift compute to inference • All three are already happening, Anthropic went from 40% margins in 2025 to possibly 80%+ this year • But labs don't want the third one, heavy inference spend signals AI progress has stalled and you're now a cloud provider • Spot compute prices are up 40%+ since February, and labs pay well above spot for security and scale • Google is reportedly paying SpaceXAI $900 million a month for 110K GPUs, roughly 2X spot • Key claim: if a human-level software engineer ran on an H100, that H100 should rent for $250K a year, 15X today's price • The 3X annual compute growth is 1.4X Moore's Law, 1.2X new fabs, and 1.8X from AI taking wafer allocation from other devices • The fab piece is bottlenecked by EUV tool supply through 2030, and the wafer piece hits a wall by end of 2027 • If compute stays scarce, new labs need far more capital just to reach the same starting line • Compute gets cheap again only once robots can turn sand and copper into computers, which is gated on robotics, not RSI @schisofrenia: "This is what I'm most concerned about, a lack of innovation, not within companies, but a lack of new companies that can actually do something substantially new, because of the scarcity of compute." @theojaffee: "The biggest companies in the world by revenue are Amazon and Walmart, at 743 billion and 725 billion. If Anthropic makes 100 billion by the end of this year, that puts them at Target. To go from Target to bigger than Walmart in a year would be very impressive."
Dwarkesh Patel@dwarkesh_sp

New blog post on what would be true about the world if trendline continues and leading lab hits $1T in revenue by the end of next year. In other words, why compute might get 10x+ more expensive in coming years dwarkesh.com/p/why-compute-…

English
4
4
34
7.6K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: Jason says half your tokens will run locally in two years. Elon disagrees. @Jason: "Open Source is winning bigly. Tokens will be 100x cheaper in 24 months. You're going to run 50% of tokens on your local hardware unmetered" • His picks to win that trend: Dell, Nvidia, and Apple @elonmusk: "Over 90% of AI compute will be in server-side for the next few years. Long-term, 99.99...% of compute will be in space" @Jason pushes back: "It will be desktop first and cloud/space second… apple and dell are seriously focused on local intelligence/SLMs and the hackers are obsessed with it. 1TB desktops coming next year!" • The 1TB claim: no such thing as a one-terabyte consumer desktop today, and a DIY build would run $20-40K, with $16-28K of that just RAM • For scale, 64GB of RAM went from $200 in September to $1,100+ now • A Claude Max subscription is $2,400 a year, roughly 128GB of RAM at current prices, before any other components • If a frontier model can do the work of a $300-400K/year software engineer, no company saves money running a worse local model to shave token costs • Where we do see local winning is agents, basic ones running on your own hardware, the OpenClaw thesis, while frontier intelligence stays in the cloud @theojaffee: "This is like a bearish AI take. If the AIs in the frontier labs continue getting much more capable, which I think they will, then it will not make economic sense to save a few dollars on tokens to use an open source model that's worse."
Elon Musk@elonmusk

@Jason @Dell @nvidia @Apple Over 90% of AI compute will be in server-side for the next few years. Long-term, 99.99…% of compute will be in space.

English
3
2
32
7.2K
MTS
MTS@MTSlive·
Ahmad Osman warns we're in the DoorDash and Uber free-rides era of tokens and it ends the same way: "I think right now any business should be looking into building their own infrastructure and controlling it end to end. We are moving towards smaller and more efficient and specialized models depending on the workflows of these businesses or these organizations." "We know that we're currently in the subsidized era of tokens. We are in the DoorDash, no margins and Uber free rides era of LLMs, and that will come to an end." "When we're paying for these subscriptions or these APIs, we are allowing these labs to go raise on leverage that they have these subscribers so that they can go build more data centers, and then these data centers are stopping consumers or people that wanna control their infrastructure from actually acquiring the hardware because it's more expensive and the supply is basically limited." "So stop paying them for tokens. Go control your own infrastructure. Build something that you can control end to end." "Ultimately, you're giving them your intellectual property, and you're giving them something that is so valuable for your business for free, and you're thinking that it's just because it's seamless that they are not analyzing that." @TheAhmadOsman @OsmanticAI
English
1
1
30
4.4K
MTS
MTS@MTSlive·
Ahmad Osman explains why memory prices are never going back down: @theojaffee: "I bought 64 gigabytes of DDR5 RAM for my PC last September, less than a year ago, for I think it was about $205. And the same exact RAM on Amazon right now is $1,130. So it's just incredible. Almost a 5X in price in less than a year." @TheAhmadOsman: "I'm not surprised at all. And this applies to everything." "Every time somebody's telling me, oh, memory's always been cyclical, it's gonna go up, it's gonna go down, it's because it was generational. You had to jump from DDR3 to DDR4, which allows for better bandwidth, et cetera. Just a different technology came up. And then everybody needed to upgrade their servers, and that's how it was cyclical." "Right now, there's just infinite demand for it, and it's not gonna go away because people wanna continue to produce more and more tokens." "So memory is no longer cyclical in my opinion." @OsmanticAI
English
3
2
25
4.9K
MTS
MTS@MTSlive·
Ahmad Osman on how fast local AI closed the frontier gap: DeepSeek V4 Flash now runs at home on a $5,000 machine "Earlier this year I predicted that by the end of 2026 we'd have a model of equivalent intelligence to Opus 4.5 running on a single RTX Pro 6000, the GPU that back then was about $10,000. Today we got DeepSeek V4 Flash, and it's punching above its weight." "It's about 70% smaller than GLM 5.2, which came out a month ago, and it's doing better than it, it's doing better than Opus 4.6." "It would easily run on a single DGX Spark that costs $5,000, and that's frontier-level intelligence we had just six or eight months ago, running at home." "Open source and local AI are much further along than Llama 2, which is what a lot of people think of when we talk about local AI, back in 2024 when it wasn't that good. We're pretty much at the point where we're competing with frontier intelligence." "You're running something similar to Opus 4.6, which everybody loved six months ago and was writing blogs about, and you can run that on a $5,000 machine right now." @TheAhmadOsman @OsmanticAI
Ahmad@TheAhmadOsman

Prediction We will have Claude Code + Opus 4.5 quality (not nerfed) models running locally at home on a single RTX PRO 6000 before the end of the year

English
6
10
64
10.7K
MTS
MTS@MTSlive·
Bijan Tavassoli on Leopold Aschenbrenner's German political roots: at 14 he argued for universal basic income at the Green Party's national conference "I was active in the German political scene at the time, and I think one thing that you were maybe missing there is that it wasn't that unusual. For example, his UBI proposal, that's something that came to the US later, but in Germany, that was well within the mainstream of the Green Party, or at least the youth wing." "They themselves were like the dominant force in high schools and colleges, a bit like not being for Donald Trump in 2016 in a US college campus. So it wasn't unconventional. It was very conventional." "He's talking about autonomous driving and how that is impacting jobs and how, so in the long term we need UBI. The thing is, I've always been opposed to that position for a plenitude of reasons." "At that time, I thought that we needed to create jobs and fulfilling jobs. But most of all, that the problem wasn't that people weren't going to drive anymore, because Germany does have truck drivers, cab drivers, and so on, just like every other country." "The special thing about the German industry at that time was one of the world's largest car making industries" "That's now not gone, but it's on the way out. And I think that's in large part due to the policies we've seen over the last ten years, the lack of capital investment and also a hostility to progress and innovation in German society, not just politics, but also in corporate environments." @BijanTavassoli
MTS@MTSlive

SITUATION EXPLAINED: Leopold Aschenbrenner lore. • Born and raised in Germany by two doctors, educated at the John F. Kennedy School in Berlin • At 14, he spoke at the German Green Party's national conference, in 2016 • He warned that nearly 2 million German drivers and truckers faced automation within years • His proposal, at 14, was universal basic income, years before Andrew Yang made it mainstream • Enrolled at Columbia at 15, graduated valedictorian at 19 in economics and mathematics-statistics • At 17, he co-authored "Existential Risk and Growth" with Philip Trammell, now head of economics at Epoch AI • Tyler Cowen publicly flagged the paper as remarkable at the time and gave him one of the first Emergent Ventures grants • In February 2022 he joined the founding team of the FTX Future Fund, alongside Will MacAskill • Also on that team: Avital Balwit, now his wife and Dario Amodei's chief of staff at Anthropic • FTX collapsed nine months later, after the fund had already made $100 million in grants • He joined OpenAI's Superalignment team as a founding member in July 2023, under Ilya Sutskever and Jan Leike • He was fired in April 2024 over an alleged leak, which he disputes, saying it followed warnings he raised about OpenAI's security practices • Weeks later he published the 165-page "Situational Awareness," dedicated to Ilya Sutskever • He founded Situational Awareness LP that same year, with anchor investments from Patrick and John Collison, Nat Friedman, and Daniel Gross @theojaffee: "Leopold Aschenbrenner went to Columbia at the age of 15 and got into effective altruism and longtermism. He founded Columbia's University EA chapter. I know some people who knew Leopold at Columbia, and he was always sort of on the ball. He was always, like, very smart and very goal-directed."

English
4
4
36
25.6K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: Dylan Patel is turning SemiAnalysis research into a $400 million fund. • The fund is called SemiAnalysis Capital Fund One, targeting $400 million • Patel already has stakes in roughly 20 startups, including Mira Murati's Thinking Machines • He previously raised a $50 million SPV as part of Fluidstack's $700 million round • SemiAnalysis is on track for over $100 million in revenue this year, up from around $20 million last year • It has about 90 people, split between engineers across the chip supply chain and former hedge fund staff • The firm has never taken outside investment and is founder-led, so this is its first real move into managing other people's money @theojaffee: "This was always kind of inevitable, right? Was SemiAnalysis really going to stick to research, and newsletters, and models for financial companies forever? Or are they gonna branch out into making money themselves from all of this research that they're doing?"
English
3
13
196
22.4K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: MiniMax released H3 with open weights, and it's already #1 in video editing. • Priced at $7.80 per minute of 2K video, against Seedance 2.0 at $22.45 and Kling 3.0 at $20.16 • It ranks #1 on Artificial Analysis for video editing, #2 for text-to-video behind Gemini Omni Flash, and #3 for image-to-video • H3 takes text, images, video, and audio in a single context and generates 5-15 second clips at native 2K, 24fps, with native stereo audio • MiniMax is releasing the weights under its Community License within days, which would make it the strongest open-weights video model by a wide margin • It launched the same week ByteDance rolled out Seedance 2.5, which now generates 30-second 4K video from a single prompt • Seedance 2.5's face consistency across a full 30 seconds is the visible jump, earlier models drifted into a different face partway through • Voices are still the obvious tell, the video reads as real long before the audio does • Smaller labs building purpose-built video models are now beating frontier labs at this specific task @theojaffee: "This is gonna be one of the biggest examples of things that people are gonna do with these models, historical simulations, basically. Nostalgia videos for zoomers."
MiniMax (official)@MiniMax_AI

It's coming. @Hailuo_AI #MiniMaxH3

English
5
5
256
32.2K
MTS
MTS@MTSlive·
SITUATION EXPLAINED: DeepSeek V4 Flash now beats its own Pro model, from post-training alone. • Priced at $0.14 input and $0.28 output per million tokens, roughly a third of V4 Pro • It scores 82.7 on Terminal Bench 2.1, against V4 Pro Preview's 72.1 and Opus 4.8's 85.0, a 2.3-point gap at a fraction of the cost • On DeepSWE it went from 7.3 to 54.4, purely from the post-training upgrade • It now sits level with Gemini 3.6 Flash at 50 and within a point of GLM 5.2 and Muse Spark 1.1, trailing Kimi K3's 57 • Same architecture, same 284 billion parameters, same pricing, the entire gain came from re-post-training alone • That makes it a near-vertical jump on the Pareto frontier, sitting almost directly above the old model at the same cost per task • The release also adds native Responses API and Codex support @sophiadew: "It's still the same model, still the same size as the previous V4 Flash, and it's retrained, but it's retrained on just post-training, and they're specifically doing it to be better at agent tasks." "When it's 70% smaller, this means it's gonna get easier for people to be running these models at home on their own. Maybe our final dream of being able to run our own models here at the MTS Studio can actually come to life."
DeepSeek@deepseek_ai

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex! Check out the configuration details in our official API docs: api-docs.deepseek.com/quick_start/ag…

English
5
6
38
7.5K