Yitong chen

88 posts

Yitong chen

Yitong chen

@1997yrrr

Katılım Ağustos 2024
625 Takip Edilen7 Takipçiler
zhuzilin
zhuzilin@zhuzilinallen·
感觉持续学习就是 llm 界的世界模型....
中文
6
5
76
8.2K
crazyphage
crazyphage@crazyphage·
做了开源几年,关了几万个 Issue,合了很多几万个 PR 之后我觉得我还是能说上两句的。目前开源社区的主要矛盾就是:无限扩张的 AI PR 供给,同开源维护者有限而宝贵的注意力之间的矛盾。
中文
51
4
67
8.4K
Yitong chen
Yitong chen@1997yrrr·
@docmilanfar it's kind of hard to make this theorem really work in diffusion model
English
0
0
1
518
Peyman Milanfar
Peyman Milanfar@docmilanfar·
What do polar coordinates, polar matrix factorization, & Helmholz decomposition of a vector field have in common? All are implied by Brenier’s Thm: a cornerstone of Optimal Transport theory. It’s a fundamental decomposition result that’s not widely known 1/6
Peyman Milanfar tweet media
English
7
56
512
27.1K
Vic Zhang
Vic Zhang@RealVicHere·
前几天看到了 Kimi 关于 tool-calling 工程优化的分享,大意是为了避免过多的工具信息把 context 撑爆,做了一个 tool-search 工具,让 AI 能够在需要使用 tool 的时候,按需查找、调用。 今天在管理安装的 Claude Code Skills 时,我看着超多的 Skills 突然意识到,我们也需要一个 skill-search 工具。 可以在本地搭建一个简易后端,用于让 Agents 能够在需要调用 Skills 的时候,按需查找、调用。
中文
77
4
40
19.7K
Yitong chen
Yitong chen@1997yrrr·
@NikStarodub RAE outlier tokens seem to come from the outliers already in DINO. Since there are no visual token norm heatmaps for JiT and SiT, I guess the outlier tokens are less obvious than in RAE.
English
0
0
0
26
Yitong chen
Yitong chen@1997yrrr·
@NikStarodub RAE also is a class-conditional models, not MM-DiT. The author didn't show the JiT/SiT internal token nomr heatmap, but they indeed see the the improvement across them
Yitong chen tweet mediaYitong chen tweet media
English
3
0
0
42
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
Register tokens were introduced to fix artifacts in Vision Transformers (ViTs) (arxiv.org/abs/2309.16588) In Diffusion Transformers (DiTs), we don't observe such outliers, so registers shouldn't help. But they do... And especially for pixel-space DiTs. Simply adding a few empty tokens can improve FID from 3.52 to 2.69.
Nikita Starodubcev tweet media
English
5
20
132
8.5K
Yitong chen
Yitong chen@1997yrrr·
@fiapp_pro qwen 最大号的model感觉一直有 benchmarkmax 的感觉,有点像 gemini
中文
0
0
0
537
索螺丝
索螺丝@fiapp_pro·
如何评价 Qwen3.8?
中文
45
0
16
18.2K
Yitong chen
Yitong chen@1997yrrr·
@wangzhian8848 k3 实力很强,模型大小是 glm5.2 的三倍多,从各个层面来说训出来都不容易
中文
1
0
0
8.4K
王局志安
王局志安@wangzhian8848·
那个Kimi k3到底好用不好用?
中文
152
4
149
220.9K
Chris Jericho
Chris Jericho@chrisjericho603·
@giffmana What's the connection between Curry finding a dead bounce and LLM solving ghost font?
English
1
0
12
5.5K
熊师傅 weight decay 了吗
熊师傅 weight decay 了吗@bigeagle_xd·
the fifth should be replaced with “不要比烂”. think not what you've done better than someone else, but what other labs have done better than you. oai and a\ are still leading the real frontier, glm can match opus with a <1T model, dpsk has incredibly high inference efficiency, we have very much work to do!
Xinyu Yang@Xinyu2ML

Why can Kimi ship K3? Let me tell my story. Earlier this year, I left academia for industry. I talked to a lot of companies along the way. Here's what I saw: 1⃣Arrogance. They believe the AI war is over, and they won. No hunger for the future, and no hunger for talent. 2⃣Restlessness. Young labs short on foundation, either rushing to catch the frontier or pivoting away from the competition. 3⃣Fear. Strong teams with real experience, but from the second tier, they can't quite bring themselves to aim for #1. 4⃣Misalignment. Everyone is optimizing for their own credit, but nobody really cares whether the company can reach AGI. Kimi was different. Over many conversations with the founders, the same thing came through every time: a raw, genuine hunger for AGI. I joined. The hunger was real. We shipped K3. This is only the beginning.

English
8
13
186
55.9K
Yitong chen retweetledi
Kai
Kai@real_kai42·
憋了太久了,终于能聊 k3! 天天被问,现在可以正面回答,Kimi Code 是用 K3+Kimi Code 构建的,用多了 K3,感觉 claude/gpt有点流口水(菜) 好消息:K3 终于发了,大家可以一起嗨了 坏消息:我得跟大家一起抢 K3 的容量了😭,我已急哭 🛫🛫🛫
中文
171
23
998
180K
李博
李博@leeeboo·
开发 DataLayer Studio 时,GPT‑5.6 Sol 给了我很大帮助:我真的没想到一个模糊的想法经过几百次的commits可以真的变成一个我自己每天用的东西: apps.apple.com/us/app/datalay…还有一个场景,我司的所有AWS的硬件和K8S,现在全由5.6 Sol负责管理,对我们这种没有专业运维的团队帮助很大!
中文
2
1
14
3.9K
Kyle Sargent
Kyle Sargent@KyleSargentAI·
A hot take: the videogen pipeline of the future is an LLM-coded 3D/4D base layer plus a lightweight diffusion renderer. We already know how to build the latter, but the former requires VLMs to be much better verifiers. So if you want to unblock generation, work on perception
Brandon McKinzie@mckbrando

At last, I can share what the Blender video was making: a slop cannon. Here is the final simulation. All made from scratch in Blender by 5.6-Sol.

English
17
10
146
19.9K