Yijie He

182 posts

Yijie He

Yijie He

@YijieHe_zero

Statistics bachelor @Fudan | exchange(learning ML,RL) @Harvard | Master in Management @LBS AI believer

Katılım Ekim 2022
273 Takip Edilen60 Takipçiler
Yijie He
Yijie He@YijieHe_zero·
@gabriel1 I feel it's unrealistic for people to write this kind of updates daily, does that mean model need to develop some kind of capability that it can observe how I clean my inbox for several days and then it knows my preferance. Then we just need to nudge it sometimes
English
1
0
2
114
gabriel
gabriel@gabriel1·
tell the smartest person in the world but who doesn't know you to "clean your inbox", they'll get so many things wrong and it would take you hours of walking through emails to write a full 10 page instruction for what exactly is relevant and what is not, that updates daily
English
15
9
247
18K
gabriel
gabriel@gabriel1·
every job will turn into explaining your intentions to ai explaining what you want to ai is surpringly time consuming, coders already spend 80% of their time doing it, and this will be true for everyone
English
341
234
3.2K
575K
Yijie He
Yijie He@YijieHe_zero·
@YuLin807 AK加入claude基本锤死mythos确实出现self evovle的早期信号了,而且AK很有可能帮助AI 研究问题taste突破,只是不知道距离我们看到需要多久
中文
1
0
1
403
QingYue
QingYue@YuLin807·
这篇采访让我觉着有意思的一个桥段是 陈茜问田渊栋“你怎么看待Andrej Karpathy加入Anthropic这件事” 然后得到的回复是“这件事证明了Auto research / self improve 非常重要” 结合昨天A社发布的那条他们已经开始用Claude来优化训练Claude的帖子来讲,这确实是下一个业界都需要关注都要卷的地方。 我之前对田渊栋也不是很了解,只是知道他好像从 meta 出走这事儿传播度比较高,但是真正让我对他感兴趣的是,采访中他说自己发过一篇和Agent 顿悟相关的研究,看到这个我决定认真追踪他的动态和方向。 我很早之前还在养龙虾的时候说过,我发现这个openclaw 的记忆真的是一塌糊涂,所以我决定不去卷记忆,我要想办法提高他的悟性。所以 会有我给他的soul.md 加入的一条“本自具足”,这样来看,这个方向我们都挺感兴趣的。
QingYue@YuLin807

对于喜欢关注AI前沿的朋友来说,硅谷101是个不错的信源头,对我来讲,一个是足够前沿,另外一个是中文友好,而且你读完是真的会有收获,把前沿的事情制作为大家可以听懂的节目,这是需要我们学习的一个技巧。 另外我去年发现他们居然也有视频号。

中文
43
2
57
24.1K
Yijie He
Yijie He@YijieHe_zero·
DeepSeek V4 is bearish for storage narrative short-term, but the math says otherwise mid-term. V4 achieves Opus 4.6-adjacent capability using 10% of V3.2's KV cache at 1M context. Headline read: KV cache compression kills storage demand. This is wrong on a 12-month horizon. Here's the framework: Per-request KV cache storage requirement ↓ (first-order: bearish storage) Workload expansion = agent task persistence × task scope × user base ↑ Net storage demand = f(compression ratio, workload expansion rate) This is Jevons paradox. We've seen it twice already — early DeepSeek moment and TurboQuant both triggered identical storage sell-offs that reversed within 2 quarters. The muscle memory is now "dip and rip." The non-obvious point: North American labs have structurally different incentives. Anthropic/OAI's optimization function = safety alignment + architectural stability to ship SOTA. Rearchitecting KV cache layout at coarser granularity introduces training instability (this is why V4 slipped from Feb to April). For frontier labs chasing capability, compute efficiency from architectural rewrite is not worth the model stability tax. Translation: even if the open-source recipe is public, industrial-scale replication by North American hyperscalers takes 6+ months minimum. Storage demand from frontier training/inference workloads is insulated on that timeframe. Trading view: SNDK: 3-5% downside on narrative; 1200 target intact post-ER. Position sensitivity is now low given how many rounds of this the holders have survived. The more interesting read-through is Chinese compute stack. V4 did substantial Ascend NPU adaptation — the domestic AI compute ecosystem just got a credible workload anchor. watching: SNDK, domestic compute names
DeepSeek@deepseek_ai

🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at chat.deepseek.com via Expert Mode / Instant Mode. API is updated & available today! 📄 Tech Report: huggingface.co/deepseek-ai/De… 🤗 Open Weights: huggingface.co/collections/de… 1/n

English
1
0
0
188
Yijie He
Yijie He@YijieHe_zero·
@felixrieseberg Thank you, I will try this to see if it can fix my problem. Thank you for your great product!
English
0
0
0
568
Felix Rieseberg
Felix Rieseberg@felixrieseberg·
@YijieHe_zero Thank you! I love Notion, so getting this right is important to me. I’ve worked for a while on Notion’s web tech, so I know that it’ll be hard for the current models to interact with the page like a human - but the Notion team has an MCP connector that should get you pretty far!
English
1
0
8
6.3K
Felix Rieseberg
Felix Rieseberg@felixrieseberg·
👋 Hi, I'm Felix and I work on Claude Cowork, bringing Claude Code closer to all kinds of knowledge work. It's an early and rough preview, please tag me in any feedback - we want to iterate very quickly and make it a little better every day.
English
372
162
5.2K
598.3K
Yijie He
Yijie He@YijieHe_zero·
@polynoamial I have the same feeling about using AI, it can speedup what I can do but cannot help me do what I am not capable. I tried to use Claude and Chatgpt to do the financial model based on the prospectus. But it can hardly really give me accurate number and solid logic.
English
0
0
0
172
Noam Brown
Noam Brown@polynoamial·
I vibecoded an open-source poker river solver over the holiday break. The code is 100% written by Codex, and I also made a version with Claude Code to compare. Overall these tools allowed me to iterate much faster in a domain I know well. But I also felt I couldn't fully trust them. They'd make mistakes and encounter bugs, but rather than acknowledging it they'd often think it wasn't a big deal or, on occasion, just straight up try to gaslight me into thinking nothing is wrong. In one memorable debugging session with Claude Code I asked it, as a sanity check, what the expected value would be of an "always fold" strategy when the player has $100 in the pot. It told me that according to its algorithm, the EV was -$93. When I pointed out how strange that was, hoping it would realize on its own that there's a bug, it reassured me that $93 was close to $100 so it was probably fine. (Once I prompted it to specifically consider blockers as a potential issue, it acknowledged that the algorithm indeed wasn't accounting for them properly.) Codex was not much better on this, and ran into its own set of (interestingly) distinct bugs and algorithmic mistakes that I had to carefully work through. Fortunately, I was able to work through these because I'm an expert on poker solvers, but I don't think there are many other people that could have succeeded at making this solver by using AI coding tools. The most frustrating experience was making a GUI. After a dozen back-and-forths, neither Codex nor Claude Code were able to make the frontend I requested, though Claude Code's was at least prettier. I'm inexperienced at frontend, so perhaps what I was asking for simply wasn't possible, but if that was the case then I wish they would have *told* me it was difficult or impossible instead of repeatedly making broken implementations or things I didn't request. It highlighted to me how there's still a big difference between working with a human teammate and working with an AI. After the initial implementations were complete and debugged, I asked Codex and Claude Code to create optimized C++ versions. On this, Codex did surprisingly well. Its C++ version was 6x faster than Claude Code's (even after multiple iterations of prompting for further optimizations). Codex's optimizations still weren't as good as what I could make, but then again I spent 6 years of PhD making poker bots. Overall, I thought Codex did an impressive job on this. My final request was asking the AIs if they could come up with novel algorithms that could solve NLTH rivers even faster. Neither succeeded at this, which was not surprising. LLMs are getting better quickly, but developing novel algorithms for this sort of thing is a months-long research project for a human expert. LLMs aren't at that level yet.
Noam Brown tweet media
English
122
204
2.8K
422.2K
Yijie He
Yijie He@YijieHe_zero·
what is the architecture behind Nano banana 2, transformer or dit? if it is transformer, the success and massive adoption of nb2 will boost the demand for storage dramatically. If it is dit, then the boost will be trivial. From my opinion, bc it is good at reasoning and text, it should be transformer, but want some validation
English
0
0
0
282
Yijie He
Yijie He@YijieHe_zero·
SBF破产清算时候卖掉了Anthropic的投资 这是2021年他们500M的初始投资,股权15%,后来稀释到7.8% 现在Anthropic估值180B,光这一笔就有12B的回报,足以填补当时8B的亏损 但是已经破产了 所以不要上杠杆,survive是最重要的
中文
0
0
0
257
Yijie He
Yijie He@YijieHe_zero·
Quick Q for the crowd: Did Shopify actually roll out full ChatGPT integration yet? I’ve seen Yipit data showing Shopify is growing 40% YoY, but my impression is the Shopify x OpenAI announcement was made… while the actual integration might not be live yet. Anyone have clarity on this? 👀 #Shopify #ChatGPT #AI #ecommerce #OpenAI
English
0
0
0
201
Yijie He
Yijie He@YijieHe_zero·
Looking for practitioners running large-scale LLM inference (prod or serious staging). I’m estimating accelerator demand for 1T tokens/s and need real-world ranges for prefill/decode TPS & utilization. DMs open, can credit or anonymize. #LLM #AIInfra #MLOps Reply 1 (details, EN): Context & asks (short): • Rule-of-thumb Input:Output ≈ 3:1 • Prefill TPS ≈ PeakFLOPs / FLOPs per token • Decode TPS ≈ PeakBW / bytes per token • Utilization assumed: prefill 0.35–0.65, decode 0.25–0.55 What I need from prod experience: Prefill TPS: typical N_active ranges for dense vs sparsity/mixture (e.g., GPT-class/Claude/Gemini-class). Decode TPS: typical d_model & context T in deployment (p50/p90). Utilization: realistic ranges you’ve seen for prefill/decode under steady load. Happy to trade notes and share final summary back.
English
0
1
1
98
Yijie He
Yijie He@YijieHe_zero·
ChatGPT推出Pulse功能,将根据用户记忆主动推送内容,并通过用户反馈持续调整推送主题 目前AI chatbot只是被动响应,OAI致力于让AI助手能理解目标,主动预测需求以助力目标实现。此举服务于OAI让每个人拥有Personal Assistant的战略,为后续商业化做准备 玩法上类似今日头条,根据全局记忆推送用户可能感兴趣的主题,内容上不依赖UGC而是AI生成 在模型能力不能拉开显著差距的情况下,对用户长期偏好的了解是产品的护城河。AI长期记忆和用户偏好建模上是头部lab重点投入的方向OAI有先发优势,用户规模最大 GOOGL和META有此前积累的用户数据 XAI通过推出虚拟女友来提升互动收集偏好数据 Anthropic在个人偏好理解上投入较少
OpenAI@OpenAI

Now in preview: ChatGPT Pulse This is a new experience where ChatGPT can proactively deliver personalized daily updates from your chats, feedback, and connected apps like your calendar. Rolling out to Pro users on mobile today.

中文
0
0
0
112
Yijie He
Yijie He@YijieHe_zero·
Currently, the user base primarily consists of tech enthusiasts and independent designers, resulting in minimal incremental revenue, with actual income estimated at around $100 million. Looking ahead, Nano Banana is expected to drive a 5% increase in GCP revenue and enhance the effectiveness of Google Ads, with a realization period of approximately 6-12 months. By targeting the $3.5 billion enterprise creative market, industries such as marketing, media & entertainment, and gaming will gradually migrate to the GCP ecosystem after adopting Nano Banana for creative asset generation. This will in turn drive the migration of storage and other workloads, potentially lifting GCP revenue by over 5%. Google Ads is vertically integrating generative AI capabilities like Nano Banana, internalizing the ad creative process within its platform. This strategy aims to generate and optimize a massive volume of creative assets at scale, based on marketing goals and product data. This, in turn, drives performance improvements in automated ad products like Performance Max, ultimately attracting more advertiser budget. Compared to previous text-to-image products, Nano Banana has broadened its application scenarios and lowered the barrier to entry. This is because Google's native multimodal approach endows Nano Banana with a more comprehensive world understanding and more precise adherence to text prompts.
Josh Woodward@joshwoodward

🍌 @GeminiApp just passed 5 billion images in less than a month. What a ride, still going! Latest trend: retro selfies of you holding a baby version of you. Can't make this stuff up!

English
0
0
0
69
Yijie He
Yijie He@YijieHe_zero·
It seems weird that nobody discuss this If OAI cannot consume its compute, who will rent it from openai. It seems weird to me that for some spike demand and hoard this large amount of backup server. It almost take account of 1/3 cost in 2028
Yijie He tweet media
English
0
0
0
34
Yijie He
Yijie He@YijieHe_zero·
AI transforms shopping's supply-demand matching: from demographic → behavioral data → understanding WHY people need things. Advertising evolution: generic ads → personalized generation hitting individual needs. Power shift: TV/print → Google/Meta → potentially OAI for search, AI communities for entertainment. Key question: Brands used to bid for keywords. In AI recommendation era, how do they influence AI suggestions? The battle shifts from buying attention to shaping AI preferences.
English
0
0
0
31
Yijie He
Yijie He@YijieHe_zero·
GOOGL在法院判决无需拆分chrome后,在美国地区推出浏览器AI助手Gemini in Chrome, 并计划在未来几个月推出更多浏览器Agent功能 Gemini in Chrome依托Gemini原生多模态功能和超长上下文,支持跨页面总结信息/模糊搜索历史信息,对youtube视频进行问答,并整合google生态中的日历地图等应用 Gemini目前月活450M,chrome浏览器全球用户约4B,若未来全面开放将显著提升Gemini触达人群,对TPU算力需求带来极大的需求提振 未来推出浏览器 Agent 功能意味着 AI 工具从单纯的问答,迈向“完成任务”Agent功能可根据 Project Mariner 的 Demo 推测,具体功能包括:跨页面信息提取与根据指令执行任务、以及“教—学—复用”的循环(用户先演示一次操作,Agent 学会后可重复执行) 基于上述功能可解锁的应用场景包括机酒预订、购票、查找房源、在线购物、表单数据收集等
中文
0
0
0
97
Yijie He
Yijie He@YijieHe_zero·
OAI的GPT-5和未发布的推理模型在ICPC编程竞赛中成功解答所有题目,击败顶级人类选手(11/12)。GOOGL Gemini 2.5 Deep Think 也达10/12 模型在限时、结构化、可自动验证的环境里已超越顶级人类水准,并被AI研究人员高频使用,加速AI研发 OAI 的下一步突破方向是 L4层 AI Innovators,试图让模型解决前沿的开放性科学问题。这需要更长推理时间,单问题推理仍有约10X扩展空间 L3层能执行行动的 Agent 仍有瓶颈:编程竞赛只要通过 unit test 即算成功,模型可自证正确,但现实世界面临上下文理解和可靠性两大问题 上下文理解 生产环境任务往往涉及跨模块、跨系统的上下文理解,需要处理和记忆大量输入,对模型的理解能力和推理算力要求都更高 可靠性 现实世界更复杂,没有统一的“正确性测试”,代码需在安全性、可维护性、扩展性和大规模流量下的可靠性上接受长期检验 当前业界正在通过用户数据回传、合成数据专家标注等手段提高可用性
中文
0
0
0
82
Yijie He
Yijie He@YijieHe_zero·
@iruletheworldmo I love gpt4.5. When talking to other models, they try to echo me and expand from my thought, but gpt4.5 give something really different and new
English
0
0
0
234
🍓🍓🍓
🍓🍓🍓@iruletheworldmo·
am i the only that thinks gpt 4.5 was a huge success and a remarkable model.
English
56
3
257
17.6K
Yijie He
Yijie He@YijieHe_zero·
6/N In short, permissions are the biggest variable in AI agents replacing humans, and thus the most interesting investment entry point. It’s not just about efficiency — we should ask which parts of our work are irreplaceable. As Demis Hassabis said, top programmers still excel at defining architectures, asking the right questions, guiding code assistants and checking outputs. In finance, knowing what truly moves markets and getting privileged data (often through connections) is a real edge and should be our focus.”
English
0
0
0
31
Yijie He
Yijie He@YijieHe_zero·
5/N Ultimately, there are two possible solutions: (1) An agent platform becomes so dominant that it forces platforms to open up, like WeChat’s mini-programs; (2) platforms agree to share partial access in exchange for revenue or data usage agreements, creating an “Agent API economy.”
English
1
0
0
34
Yijie He
Yijie He@YijieHe_zero·
Why AI Still Can’t Replace Interns In my view, AI agents are limited by access permissions, not by analysis or output skills. The biggest obstacle is that services like Cloudflare block them from gathering data. Access determines an agent’s usefulness; the future will see investment opportunities in managing and bypassing permissions. Data access is like a tax in the AI era. Recently, our interns left, so as the youngest full-time employee I’ve taken on the “dirty work.” I’m trying to let AI handle it, but it’s not yet a full replacement. I think we’re still about six months away from that.
English
1
0
0
48