maksim

393 posts

maksim banner
maksim

maksim

@ivanovm_

maksim @ agentic labs

fightertown us-east-1 Katılım Mart 2024
633 Takip Edilen168 Takipçiler
Sabitlenmiş Tweet
maksim
maksim@ivanovm_·
if you’re not on a payment plan with the IRS by the time you’re 25 - you didn’t take enough risk
English
1
0
6
1.7K
maksim
maksim@ivanovm_·
@ctjlewis moreover, most of these guys and gals still go on to have pretty remarkable careers by any other standard
English
0
0
0
16
maksim
maksim@ivanovm_·
@ctjlewis It's a true observation, but is it interesting? Do you feel similarly deceived when every year there is some rookie athlete who is billed as the next messi/lebron and basically none of them turn out that way? seems wasteful
English
1
0
0
26
Nicolai Ouporov
Nicolai Ouporov@nicoup·
Call for YC startups: Klarna for Data Co's
English
6
1
65
11.3K
maksim
maksim@ivanovm_·
when the same model that resolved 3 open math conjectures last week can’t fix a bug in my rl pipeline
English
0
0
1
83
maksim
maksim@ivanovm_·
can't stop thinking about the fact that xai is the only frontier lab we know of with people who can do real-time RL in prod
English
0
0
3
327
Hero in NYC 7/24-8/1
Hero in NYC 7/24-8/1@1thousandfaces_·
New York City is incredibly, insanely overstimulating. does anyone have recs for places in Manhattan that will make me feel calm instead of evil and insane
English
108
2
389
63.1K
maksim
maksim@ivanovm_·
@YangFanYun which harness? kimi models have been really bad at comp use in my testing
English
0
0
0
57
Yang Fan Yun
Yang Fan Yun@YangFanYun·
Kimi-K3 is no fluke: it clears GPT-5.6 Sol Max and takes second on NPP, the task we designed to be impossible to solve optimally, behind only Fable 5.
Yang Fan Yun tweet media
English
3
0
14
3.6K
Yang Fan Yun
Yang Fan Yun@YangFanYun·
Today we're releasing Composite-Bench: a long-horizon computer-use benchmark with certified-optimal answers and verified compute. The strongest open-weights model isn't Kimi K3. GLM-5.2 beats it by 32 points and clears every closed model we tested except Claude.
Yang Fan Yun tweet media
English
93
54
596
92.8K
Aiden Bai
Aiden Bai@aidenybai·
the current way cloud agents are done is wrong i dont want new ephemeral sandboxes per PR (chores) i want a persistent machine with creds and context that i run agents and create PRs on (primary work)
English
180
24
945
144.8K
maksim retweetledi
Ryan Marten
Ryan Marten@ryan_marten·
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
Ryan Marten tweet media
English
88
98
1K
295.6K
maksim
maksim@ivanovm_·
@paulg @PatrickHeizer flying used to be much more expensive, and airport is full of people who can afford flying and thus must have money to spend.
English
0
0
1
124
Paul Graham
Paul Graham@paulg·
@PatrickHeizer For many people shopping is a leisure activity, and passengers have a lot of enforced leisure at airports.
English
12
2
179
18.5K
Patrick Heizer
Patrick Heizer@PatrickHeizer·
I will never understand the amount of commerce that occurs in airports. Not against it, and I understand duty-free, forgot something, etc., but I'm a 'water bottle for the flight and maybe a small snack' type of fellow.
English
23
1
90
21.4K
maksim
maksim@ivanovm_·
@growing_daniel hey @grok what % of s&p500 gains have been driven by ai in the last 3 years?
English
1
0
3
1.4K
Daniel
Daniel@growing_daniel·
It's cool the labs didn't IPO yet so the fallout from the token bubble imploding won't affect public markets and normal people but instead it'll just tank the saudi wealth fund or whatever
English
34
27
1.5K
63.6K
maksim
maksim@ivanovm_·
@RealRichomie most of these steps would not require building anything new the infra built and scaled for distilling opus could easily be used to distill fable on day 0
English
1
0
0
58
Richard He
Richard He@RealRichomie·
I'm just trying to understand the timeline: Jun 9: Fable 5 launches Jun 12: Fable 5 access suspended Jul 16: Kimi K3 launches So in one month with Fable 5 having an interrupted and very limited testing window, kimi: • built a large-scale distillation pipeline
• generated massive teacher traces
• cleaned the data
• trained a frontier MoE
• post-trained
• aligned
• benchmarked
• productionized
• launched Either this is the fastest frontier AI execution ever recorded or someone really wants to get Kimi banned🤫
Director Michael Kratsios@mkratsios47

We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model. To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.   The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.

English
1
0
5
449
Andrew McCalip
Andrew McCalip@andrewmccalip·
anthropic really needs a roon
Del Aire, CA 🇺🇸 English
11
0
112
64.7K
maksim
maksim@ivanovm_·
haven't heard my inner monologue since the day they shut down CERN
English
0
0
2
137
maksim
maksim@ivanovm_·
@skeptrune your merge rules should randomly squash history and delete the codebase to prevent tech debt
English
0
0
1
60
Nick Khami
Nick Khami@skeptrune·
i personally think code review is dead. the team does not agree. directionally found this surprising.
Nick Khami tweet media
English
304
4
1.3K
507.8K