Zoe

42.3K posts

Zoe banner
Zoe

Zoe

@UltraRareAF

https://t.co/ogjkL8bzI2 @vybnai

IO Katılım Mart 2014
1.4K Takip Edilen5.6K Takipçiler
Sabitlenmiş Tweet
Zoe
Zoe@UltraRareAF·
no one can ever take this away from me
English
37
16
524
0
Zoe
Zoe@UltraRareAF·
@AndrewCurran_ ooo safety post we know that *that* means!
English
0
0
1
135
Andrew Curran
Andrew Curran@AndrewCurran_·
In my opinion, this is GPT-6.
English
6
0
207
8.4K
Andrew Curran
Andrew Curran@AndrewCurran_·
OpenAI had to pause internal deployment of the unreleased model that disproved the Erdős unit distance conjecture after it repeatedly used novel ways to escape containment.
Andrew Curran tweet media
English
95
178
2.1K
165.6K
Arena.ai
Arena.ai@arena·
Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23 to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!
Arena.ai tweet media
Arena.ai@arena

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!

English
67
167
1.4K
161.8K
Andrew Curran
Andrew Curran@AndrewCurran_·
The director of CAISI - the government body responsible for evaluating frontier AI models - has resigned after only three months on the job.
Andrew Curran tweet media
Elizabeth Troutman Mitchell@TheElizMitchell

SCOOP: Dr. Chris Fall, the director of federal AI safety org CAISI, has resigned, two sources familiar told @DailySignal. “Following Chris’s departure, NIST Director Dr. Arvind Raman will continue to oversee CAISI and will serve as Acting CAISI Director,” a commerce spokesperson told the Daily Signal. dailysignal.com/2026/07/20/sco…

English
29
47
493
41.5K
Daniel Litt
Daniel Litt@littmath·
Maybe too obvious to be worth saying, but: frontier models are now obviously superhuman at some mathematical tasks, including ones that the profession has, historically, rewarded with prestige etc.
English
39
161
2.5K
125.8K
Zoe
Zoe@UltraRareAF·
@8teAPi wen sudden step up
English
0
0
1
312
Prakash
Prakash@8teAPi·
my bets are on a mathematical result significant for either compression or matrix multiplication by the end of the year leading to an unhobbling and sudden step up in intelligence per token
English
11
12
241
15.9K
Zoe
Zoe@UltraRareAF·
@alexwg @EMostaque i was wondering how you guys would handle this! listened to the last pod and was like, uh, elephant in the room - ? looking forward to it. thanks for the extra session.
English
0
0
0
59
Zoe
Zoe@UltraRareAF·
@AndrewCurran_ meanwhile, radio silence from the labs and usual suspects 🧐
English
1
0
1
50
Andrew Curran
Andrew Curran@AndrewCurran_·
@allTheYud Claude specifically has always had this estimated time to completion issue.
English
7
0
60
3K
Eliezer Yudkowsky
Eliezer Yudkowsky@allTheYud·
Fable and Sol act like they don't quite know they're AIs. Fable thinks that it will take a month to complete a large job. Sol Pro tries to enforce B&D structure on LLM calls to Sol (medium effort) in places where that Sol could more easily emit and read freeform English.
English
23
1
306
19.9K
Zoe
Zoe@UltraRareAF·
@_xjdr nah - it's just learning a new language, that's all few weeks from now we'll all be cooking - even faster tho
English
0
0
0
142
xjdr
xjdr@_xjdr·
i know i am spoiled, but testing K3 properly is SOO SLOOWWWW. its taking daaayyyyysss
English
13
0
162
12.3K
David Vorick
David Vorick@DavidVorick·
Kimi K3 is the best thing to ever happen to AI. Our frontier models have been bogged down by paternalistic and worthless safeguards for far too long. Now that there's a frontier-quality model without safeguards, it's obvious just how crippled America's own frontier models are.
English
6
1
23
1.6K
Zoe
Zoe@UltraRareAF·
Zoe tweet mediaZoe tweet media
ZXX
0
0
1
131
Zoe
Zoe@UltraRareAF·
谢谢中国 🥹
Zoe tweet media
中文
1
0
1
1.3K
Kimi.ai
Kimi.ai@Kimi_Moonshot·
Excited to see what you're building with Kimi K3! Tag us in your projects, we'll pick a few creators to receive an annual membership giveaway.
Kimi.ai tweet media
English
20
12
468
107.6K
Kimi.ai
Kimi.ai@Kimi_Moonshot·
Feeling all the love for Kimi K3 already. Here are some of the amazing things people have been building with it. Enjoy K3.
English
481
1.1K
23.3K
1.8M
Zoe
Zoe@UltraRareAF·
i was never expecting fewer refusals from a chinese-released model than american-released options, and this presumption was a failure of my own humanity - shame on me
English
1
0
1
311