V_V

55 posts

V_V banner
V_V

V_V

@VVDanielD

硅基生物观察员 👀 | AI enthusiast

Katılım Şubat 2026
99 Takip Edilen1 Takipçiler
沈鹏
沈鹏@shenpengwdh·
我注意到两个细节 DeepSeek V4把每位贡献者写进技术报告,包括已经离开团队的人;Kimi K3公开承认自己与顶尖模型仍有差距,最后那句话,据说是杨植麟亲自写的。 中国AI今天最让我兴奋的,不只是进步的速度,还有这种开放和自信。
沈鹏 tweet media沈鹏 tweet media
中文
52
67
1.2K
152.2K
Max For AI
Max For AI@MaxForAI·
据传钉钉将改名为“千问办公” 同时淘宝将改名千问严选 饿了么改名千问闪购 盒马将改名千问超市😎
中文
76
5
294
149.9K
V_V
V_V@VVDanielD·
@MaxForAI 快来倾销显卡
中文
0
0
4
1.9K
Max For AI
Max For AI@MaxForAI·
黄仁勋宣布美国对于中国的芯片禁令失败了。 他在福克斯的电视直播上表示: 将中国排除在英伟达(客户名单)之外,已不再意味着将中国挡在人工智能大门之外。 华为的崛起便是禁令转化为经济助推力的例证:它创造了市场,并教会本土供应商如何变强、成长并走向出口。 如今,真正的较量在于谁掌控芯片、人才、能源、基础设施、模型、应用以及整个智能技术栈。 将芯片政策视为可随意开关的阀门是错误的。 人类历史又一次证明了试图封锁来让对手落后是不现实的。 真正管用的方法是倾销。
中文
48
42
328
80.4K
V_V
V_V@VVDanielD·
I used ChatGPT Work to turn two spreadsheets and a five-year sales target into an editable quarterly business deck with clear charts, customer analysis, and an action plan. I love how it finds the story in the data and turns rough inputs into work I can actually use.
V_V tweet media
Greg Brockman@gdb

Was very cool to hear about the reasons people love Sol. We're doing the promotion again, except this time for ChatGPT Work: Tweet what you love about ChatGPT Work, claim $100 in free credits, get more work done. First 10k get the free tokens: share-chatgpt-work.openai.chatgpt.site

English
0
0
0
11
V_V retweetledi
Arena.ai
Arena.ai@arena·
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!
Arena.ai tweet media
Kimi.ai@Kimi_Moonshot

Meet Kimi K3

English
1.1K
4K
26.4K
19.2M
V_V retweetledi
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Its intelligence is comparable to Opus 4.8 and GPT-5.5 but remains behind Fable 5 and GPT-5.6 Sol. Moonshot AI has expressed plans to release the 2.8T parameter model's weights, which would make it the leading open weights model Key results: ➤ Strong agentic task performance: @Kimi_Moonshot's Kimi K3 reaches an Elo rating of 1668 on GDPval v2. This is a marked improvement over K2.6’s 1190, surpassing GLM-5.2 (1514), GPT-5.5 (1494), and Claude Opus 4.8 (1600). However, it still lags behind Claude Fable 5 (1760). Kimi K3 also scores an impressive 53% and takes the #1 position on AutomationBench-AA, our implementation of Zapier’s Agentic SaaS workflow evaluation. ➤ Second-highest performance on AA-Briefcase (agentic knowledge work): On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5. It is well-rounded: its rubric scoring and analytical quality almost reach Claude Fable 5’s scores, while GPT-5.6 Sol continues to outperform other leading models on presentation quality. ➤ Set to lead open weights models once weights are released: Moonshot AI has not yet released the weights but expressed plans to do so. Once available, Kimi K3 would clearly lead other open weights models including GLM-5.2 (51) and DeepSeek v4 Pro (44). However, at 2.8T parameters, it is significantly larger than its open weights peers (eg. GLM-5.2 at 753B params and DeepSeek V4 Pro at 1.6T), as well as the Kimi K2 to K2.6 models (1T params). ➤ Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers: Moonshot AI’s pricing for K3 is significantly higher than their K2 pricing (K3’s output token price is $15/1M tokens while K2.6 was $4). This positions the model as cheaper on a cost per task basis than Opus 4.8, similar to GPT-5.6 Sol ($1.04) and more expensive than open weights peers, GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04) ➤ Improved token efficiency alongside higher intelligence: Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6. The new model used approximately 132M output tokens to complete all nine evaluations, compared to approximately 166M for K2.6, while achieving higher scores. ➤ Native multimodal capabilities: Kimi K3, like K2.6, is released with native image and text multimodal input. If weights are released, this will position Kimi K3 as one of the leading open weights models with multimodal input capabilities Other model details: Context window: 1M Size: 2.8T total parameters Pricing: The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens. Modality: Native multimodal input supports text and images, and the model remains text-only for output. Accessibility: Accessible at launch through Moonshot’s first party API. Model weights are not yet released but Moonshot AI has expressed plans to do so.
Artificial Analysis tweet media
English
180
715
6.2K
1.7M
V_V
V_V@VVDanielD·
@Pluvio9yte antigravity内置了这个
中文
0
0
0
108
雪踏乌云
雪踏乌云@Pluvio9yte·
强烈推荐这个skill,通过问答的形式把你不确定的每个点都榨干,最后出一版事无巨细的规划 用了grillme之后,我才发现原来有那么多需求是我跟ai没有聊透的 github.com/mattpocock/ski…
中文
102
77
481
85.5K
V_V
V_V@VVDanielD·
@realNyarime 电信给我们公司免费试用一个月,结果还没我的机场稳定
中文
0
0
0
52
奶昔🥤
奶昔🥤@realNyarime·
找电信的朋友问了一下跨境网络宽带业务,应该是这个叫SDWAN跨境国际专线?说是给跨境电商用的,有营业执照就可以申请 拿到这份跨境线路资费表的时候,我的手还在颤抖,并且美国原生IP还要加钱,能保证你Claude不被封,印象里三大运营商在美的业务不是已经退出了吗? 果然垄断的东西就是贵,广州十三行 为你点赞,拉得多可以打5折
奶昔🥤 tweet media
中文
75
24
348
98.9K
V_V
V_V@VVDanielD·
I built PageVault with GPT-5.6—a self-hosted tool that turns AI-generated HTML, Markdown, and infographics into shareable links. What I like most about GPT-5.6 is how it helped me turn a small personal need into a real, deployable open-source product. github.com/vdeng-ai/PageV…
Tibo@thsottiaux

Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched? Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens! switch-to-codex.openai.chatgpt.site

English
0
0
1
32
E〽️ir
E〽️ir@emirmalte·
J’offre 1000€ à 3 personnes qui trouvent le score exact de ce match se termine dans 10 heures
E〽️ir tweet media
Français
10.6K
419
6.2K
1.4M
V_V
V_V@VVDanielD·
@keenn_eth 在世界杯还没开始的时候,我让ai预测冠军,ai的结果是法国
中文
0
0
0
9
V_V
V_V@VVDanielD·
AI 生成的 HTML、Markdown在聊天工具经常不能预览。于是做了个小工具 :上传文件 → 生成链接 → 丢进聊天窗口,对方点开就能看。 支持链接过期、文件保留时间和私有存储,可部署到 Cloudflare 免费层,也支持 Docker。 代码已经开源,欢迎踩踩、提提意见 github.com/vdeng-ai/PageV…
中文
0
0
0
19
V_V
V_V@VVDanielD·
@youy1qwq 当拍照都不是真实的
中文
0
0
0
71
由依🍥
由依🍥@youy1qwq·
华为长焦拍飞机被识别成鸟了
由依🍥 tweet media
中文
734
2.2K
34.7K
5.2M
V_V
V_V@VVDanielD·
@elonmusk 试了,目前主力模型,因为性价比最高
中文
0
0
0
11
V_V
V_V@VVDanielD·
这个排名更客观,等看看gpt5.6的排名
Artificial Analysis@ArtificialAnlys

SpaceXAI’s Grok 4.5 scores 54 to place fourth on the Artificial Analysis Intelligence Index following only Fable 5, GPT-5.5, and Opus 4.8. It scores on par with GPT-5.5 in Codex on the Artificial Analysis Coding Agent Index in the Grok Build harness, at much lower cost Grok 4.5 improves 16 points over Grok 4.3 on the Intelligence Index, bringing SpaceXAI to the intelligence frontier behind only OpenAI and Anthropic, and outperforming all open weights models and notably Google’s Gemini models. Key standout areas of performance are agentic knowledge work and coding. Grok 4.5 in Grok Build scores 76 on the Artificial Analysis Coding Agent Index, on par with GPT-5.5 (xhigh) in Codex and just below Fable 5 (max) in Claude Code, and at a small fraction of the token usage and price. Congratulations to @SpaceXAI, @cursor_ai, and @elonmusk on the impressive release! Key Takeaways: ➤ Grok 4.5 performs very strongly on agentic tasks. Grok 4.5 ranks #4 on GDPval-AA v2 with an Elo of 1543, between Claude Opus 4.8 (1600) and GLM-5.2 (1513). It achieves the top score on 𝜏³-Banking of 33%, above 31% from GPT-5.5 (xhigh), and sits on the cost vs performance Pareto frontier across all three agentic evaluations in the Intelligence Index ➤ Grok 4.5 is one of the most cost efficient models to run for near-frontier intelligence. It costs $0.31 per task on the Artificial Analysis Intelligence Index and $2.59 per task on the Artificial Analysis Coding Agent Index within Grok Build ➤ Low cost for Grok 4.5 is driven by both low pricing and token efficiency. Grok 4.5 has a headline price over 60% lower than Claude Opus 4.8 and GPT-5.5, and used ~14k output tokens per Intelligence Index Task - over 60% lower than Opus 4.8. On the Coding Agent Index, Grok 4.5 stands out on the Pareto frontier of Coding Agent Index score vs. Total Tokens, using only 1.9M tokens for the Coding Agent Index while scoring 76 ➤ As a coding agent, Grok 4.5 in Grok Build is on par with GPT-5.5 and offers efficiency benefits: In our Artificial Intelligence Coding Agent Index that consists of DeepSWE, Terminal-Bench v2, and SWE-Atlas QnA, Grok 4.5 in Grok Build ranks third, on par with GPT-5.5 (Codex) and below Fable 5 (Claude Code). It is also very efficient in achieving this result: Grok 4.5 in Grok Build cost $2.49 per task while Fable 5 in Claude Code cost $11.80 and GPT-5.5 in Codex $5.07. This is driven by relatively low token pricing and the model using far fewer tokens than comparable models (1.9M average tokens used per task), significantly less than Fable 5 in Claude Code (7.2M) and GPT-5.5 in Codex (6.2M) Other model details: ➤ Context window of 500k tokens - a reduction from Grok 4.3’s 1M token context, but retaining configurable reasoning and vision input ➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits are discounted by 75% to $0.5 per 1M tokens, and costs still double with long (>200k token) inputs ➤ As Elon Musk has disclosed, Grok 4.5 is 3x larger than its predecessor at 1.5T parameters

中文
0
0
0
17
政经鲁社长
政经鲁社长@xzzzjpl·
关于台湾问题,我觉得一劳永逸的最佳方案是这样的:第一,两岸实行邦联制,统一国号,其他一切不动。请注意,不是联邦。第二,台湾承诺永不改变现状,大陆承诺永不动兵,写入双方宪法,并制定相关法律。第三,经济上仿照欧盟,取消绝大部分关税,尽可能简化往来措施。你觉得怎么样?
中文
3.3K
73
2.5K
1.2M
V_V
V_V@VVDanielD·
@ActiveTK5929 可以反过来吗?这样我就有足够的vram跑llm了
中文
0
0
0
202
ActiveTK․
ActiveTK․@ActiveTK5929·
WindowsでGPUのメモリ(VRAM)をRAMディスクとしてマウントできるトンデモなアプリ「VRAMDISK」を開発しました!! さらにVRAM上に置いたファイルをそのままGPGPUで超高速に圧縮したり、ハッシュ値計算をしたりする機能もあります。ぜひお試しください! github.com/ActiveTK/VRAMD…
ActiveTK․ tweet media
日本語
60
971
3.5K
373.7K