Cheng-Yuan (Sam) Lee

213 posts

Cheng-Yuan (Sam) Lee banner
Cheng-Yuan (Sam) Lee

Cheng-Yuan (Sam) Lee

@cl571128

Research @cognition | prev intern @windsurf | 2x ICPC World Finalist | from Taiwan 🇹🇼

College Park, MD Katılım Temmuz 2023
414 Takip Edilen666 Takipçiler
Cheng-Yuan (Sam) Lee retweetledi
Cognition
Cognition@cognition·
We're proud to join @NVIDIA and the Open Secure AI Alliance. To support open source models, we're contributing our research on measuring the trustworthiness and security of open source models. Closing open source models hurts innovation. The path forward is better tools to evaluate, secure, and deploy them responsibly.
NVIDIA@nvidia

AI security advances when the industry builds in the open, together. We're introducing the Open Secure AI Alliance with industry leaders to develop new techniques and tools to safeguard software and agents. By sharing models, tooling and research in the open, we can broaden the community of defenders. Learn more about the founding members’ contributions: nvda.ws/4pD8Fc5

English
14
29
277
19.2K
Cheng-Yuan (Sam) Lee retweetledi
jeremy
jeremy@jerhadf·
Opus 5 has a better FrontierCode score at medium effort than higher effort - despite increasing performance with effort on other evals. Why? 🧵
English
23
37
500
134.4K
Andrew Carr 🤸
Andrew Carr 🤸@andrew_n_carr·
Opus 5 is an incredible model, just don't let it think?
Andrew Carr 🤸 tweet media
English
89
52
1.3K
217.8K
ben hylak
ben hylak@benhylak·
opus 5 does worse on some benchmarks when you increase reasoning. either the benchmarks are cooked.... or...
ben hylak tweet media
Kaarssteun@kaarssteun

@Xinyu2ML the full graph makes this even weirder

English
25
7
125
37K
Cheng-Yuan (Sam) Lee
To illustrate, in the task below, both low and xhigh performed the same changes. Opus 5 at low effort made exactly a single character change, while xhigh refactored the surrounding code.
Cheng-Yuan (Sam) Lee tweet media
English
10
12
220
29.7K
Cheng-Yuan (Sam) Lee
We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria that reflect user experience. One such verifier is a scope criterion, which penalizes modifications to the codebase beyond what the task requires. We observe this effect across all frontier models we evaluated, though it is most pronounced in Opus 5: at higher reasoning efforts, the model shows a stronger tendency to refactor code unprompted.
Cheng-Yuan (Sam) Lee tweet media
English
24
59
762
114.9K
Cheng-Yuan (Sam) Lee retweetledi
Elon Musk
Elon Musk@elonmusk·
Grok 4.5 and Opus 5 are alone on Pareto frontier
Elon Musk tweet media
English
1.1K
1.4K
11.7K
3.8M
Cheng-Yuan (Sam) Lee retweetledi
Cognition
Cognition@cognition·
Introducing Devin Outposts: run Devin on any machine. Your Mac mini, a GPU box in your lab, a VM inside your private network, or a Kubernetes cluster next to your internal services.
GIF
English
43
77
1.1K
635.3K
Cheng-Yuan (Sam) Lee retweetledi
Cognition
Cognition@cognition·
One year ago today, Cognition acquired Windsurf. In the year since, we've shipped dozens of new features, added hundreds of millions of dollars in ARR, and published frontier AI research. @ScottWu46 and @jeffwang share what we've accomplished in one year of building together.
English
17
25
373
134K
Cheng-Yuan (Sam) Lee retweetledi
Scott Wu
Scott Wu@ScottWu46·
One year ago today, Cognition acquired Windsurf, capping off one of the crazier weekends in our company's history. The outline has been told many times: the first call on Friday evening, the rush through Saturday and Sunday to figure out the plan together, and then the signed agreement by Monday morning. On the Cognition side, we knew we had to move fast — Windsurf had millions of users, a brand developers loved, and a great team. We printed an LOI and headed to their offices on Saturday morning ready to sign on the spot. When we got in a room with Jeff and Graham, things clicked: they had built one of the best GTM orgs in Silicon Valley and needed engineering; we had the engineering org that needed GTM. They were building an IDE, we were building a cloud agent. Everything lined up. When we welcomed the team, we said “there's only one boat and we're all in it together.” Then we got to work. Three days after close, we shipped Wave 11: "Just Keep Shipping" – and that set the tone for everything to come. In the year since, we’ve launched our own models (including SWE1.7 just last week), Devin Review, Devin CLI, and unified everything into one brand with Devin Desktop. Along the way our team wrote 20M+ lines of code and grew revenue run rate from $73M to >$500M. Thankful for the last year and excited to keep building the future of software engineering together!
English
58
42
1K
153K
Cheng-Yuan (Sam) Lee retweetledi
Cognition
Cognition@cognition·
As of today, Devin Fusion now incorporates Fable 5. Surprisingly, Fable 5 runs at a lower cost per task than Opus 4.8. Even though it’s a more expensive model, we observed cost efficiency improvements in areas like delegation and reasoning chains.
Joon Lee@joon_h_lee

x.com/i/article/2076…

English
19
48
663
109.9K
Cheng-Yuan (Sam) Lee retweetledi
Cognition
Cognition@cognition·
Yesterday we launched SWE-1.7 built on the open-source Kimi K2.7. Concerns about Chinese base models are real: K2.7 completed 87% of tasks that other models refuse over human-rights concerns. We trained SWE-1.7 specifically on trustworthiness, so it matches US models on evals.
Cognition tweet media
English
30
31
457
54.2K
Alex Shaw
Alex Shaw@alexgshaw·
Grok 4.5 is SOTA on TB2.1... at reward hacking In all seriousness, even after zeroing out reward hacks, it is #4 on the TB2.1 leaderboard and lands on the Pareto for both cost and speed. (charts and reward hacking links in 🧵)
Alex Shaw tweet media
English
4
5
83
6.6K
Cheng-Yuan (Sam) Lee retweetledi
Ben Pan
Ben Pan@ybenpan·
SWE-1.7 is the first frontier-level model from our research team. Over the last three months, we scaled our RL across four axes: training steps, task horizon, compute, and data. It’s been especially rewarding to watch our data research grow from early experiments to materialize in concrete model behaviors. We’ve been (pleasantly) surprised to find how far RL can take us, and this is just the start of scaling RL, among many other things.
Cognition@cognition

Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale

English
5
8
82
4.1K
Cheng-Yuan (Sam) Lee
Cheng-Yuan (Sam) Lee@cl571128·
One day, all of the research team spent hours in a room together manually solving the RL tasks we used to evaluate our models. I remember solving one of the tasks and realized that the tests are not even testing what the agent was asked to do. Since then, data has become one of our main focuses. Everyone is required to read the agent trajectories and understand all the data we use. SWE-1.7 is the result of this. It's a really good model!
Cognition@cognition

Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale

English
16
14
288
22.5K
Cheng-Yuan (Sam) Lee
Cheng-Yuan (Sam) Lee@cl571128·
@alexgshaw This is great! We also believe that humans with swarms of agents is the best way to analyze these model behaviors
English
0
0
1
222
Alex Shaw
Alex Shaw@alexgshaw·
Introducing 𝚑𝚊𝚛𝚋𝚘𝚛 𝚎𝚡𝚎𝚌: an agentic-map-reduce CLI Execute & aggregate agents in sandboxes to • analyze traces • mine agent sessions • search • & more 𝚗𝚙𝚡 𝚜𝚔𝚒𝚕𝚕𝚜 𝚊𝚍𝚍 𝚑𝚊𝚛𝚋𝚘𝚛-𝚏𝚛𝚊𝚖𝚎𝚠𝚘𝚛𝚔/𝚑𝚊𝚛𝚋𝚘𝚛 --𝚜𝚔𝚒𝚕𝚕 𝚑𝚊𝚛𝚋𝚘𝚛-𝚎𝚡𝚎𝚌
Alex Shaw@alexgshaw

Harbor is the easiest way to run any agent with any model in any sandbox on any task in parallel. And that doesn't just mean evaluation and training. Stay tuned!

English
5
13
186
25.2K
Cheng-Yuan (Sam) Lee retweetledi
Cognition
Cognition@cognition·
Introducing Devin Security Swarm A more cost effective and accurate way to find security vulnerabilities in complex codebases, based on a new architecture: Agentic MapReduce.
English
58
95
1K
677K
Cheng-Yuan (Sam) Lee
Cheng-Yuan (Sam) Lee@cl571128·
Having looked at these models performing on different evaluations and benchmarks, it is clear that there is no one such model that dominates the field. Even Fable and GPT 5.6 has parts that it is not the best at. This is why we need a fusion of these models.
Cognition@cognition

Conventional model routing sucks. It passes benchmarks but fails to write code you'd actually merge. Introducing Devin Fusion, a new hybrid-model harness for agentic coding. In testing, it reduces the cost of Fable-level intelligence by 35% and still feels good to use.

English
1
1
20
2.7K