LannTheSmart

76 posts

LannTheSmart

LannTheSmart

@LannyTheClever

Katılım Kasım 2016
148 Takip Edilen5 Takipçiler
LannTheSmart
LannTheSmart@LannyTheClever·
@BosonJoe Can we fix 4 concurrent requests crash with mtp?
English
0
0
0
36
Joe Muller
Joe Muller@BosonJoe·
Qwen3.6 27B NVFP4 Single Spark Recipe ⚡️ 🚀 vLLM 0.26.0, no patch 🪣 262K context 👾 31 tok/sec decode You could probably run this and the 35B MoE Qwen model on the same Spark if you wanted the best of both howtospark.com/recipes/qwen3-…
English
7
8
155
8.8K
Joe Muller
Joe Muller@BosonJoe·
🌞🎚️🌚
Joe Muller tweet mediaJoe Muller tweet media
QME
2
0
7
924
ascdas
ascdas@tuboshu321·
是的,gpt pro 会因为网络环境被路由 到5.5 mini。前端时间找了很多处理办法。 1)先把手机app 删掉,再安装。关闭vpn 2)重新下载,打开vpn,登录,手机测试gpt pro是否正常 3)这块主要是网页端 chrome存在指纹并且上传很多信息,cloudflare 识别给到openai 不可信,导致网页自动路由mini 4)ios 走另外的线路。
中文
1
1
2
340
余温
余温@gkxspace·
学到真东西了,GPT网页版的 5.6 Pro Review 居然也那么牛逼 我平时调研、深度对话都会用 Pro 模式,思考强度和效果比 Codex 和其它产品都强,看到卡神居然还能这样玩 顺手安利一个很多人不知道的用法: Pro 其实可以在 Codex 中打开,Pro 的结果也可以一键导进 Codex 的对话记录,接着让 Codex 干活,很方便 都给我用起来!!!
余温 tweet media
数字生命卡兹克@Khazix0918

可能很多人不知道,其实现在做代码Review最牛逼的,是ChatGPT网页版上的GPT-5.6 Sol Pro模式。 其Review水平和深度,让Fable 5都甘拜下风。。。

中文
75
22
233
79.6K
Unsloth AI
Unsloth AI@UnslothAI·
@Alibaba_Qwen Second to only Fable 5? We're super excited for it! 🤯 Please release smaller/medium sized models as well so we can run it locally! 🙏
English
55
65
2.2K
157.4K
Qwen
Qwen@Alibaba_Qwen·
Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out. Can't wait to hear what you build. Stay tuned! 🚀  Token Plan international:qwencloud.com/pricing/token-… China:platform.qianwenai.com/pricing/token-…
Qwen tweet media
English
1.3K
3.4K
24.4K
8.3M
LannTheSmart
LannTheSmart@LannyTheClever·
@I_am_oil_oil 为什么不考虑把 5.3spark 也加入进来?这个额度消耗更少,并且更快
中文
0
0
0
19
oil-oil
oil-oil@I_am_oil_oil·
我把这段时间使用 Codex 子 Agent 的方法整理成了一个开源 Skill:github.com/oil-oil/codex-… 主 Agent 负责理解需求、控制范围和最终验收。很小的修改不启动子 Agent,直接自己做;只有需要隔离上下文、并行或者独立复审的时候,才把任务分出去。
中文
4
5
69
8.1K
Mia
Mia@MiaAI_lab·
Run the new @UnslothAI Qwen3.6-35b-NVFP4 on your @NVIDIAAI DGX Spark ease! ✨ 256K context • 24 concurrency ~81 tok/s single session ~350+ tok/s at 24 sessions This is now the recommended Qwen3.6-35B terms of accuracy/performance. Get it here 👇 github.com/MiaAI-Lab/Unsl…
English
23
24
246
14.8K
Mia
Mia@MiaAI_lab·
This is a big one 🔥 Run DeepSeek v4 Flash Abliterated & Uncensored on your 2x DGX Sparks with easy ✨ 1M context • DSpark • 57 tok/s ✨Special thanks for @u1tra_instinct for the hard work on this! github.com/MiaAI-Lab/Deep…
English
17
15
213
16.6K
LannTheSmart
LannTheSmart@LannyTheClever·
@MiaAI_lab @NVIDIAAI What is your benchmark? Which scenarios did you take into account? I would really like to run some tests on my own machine.
English
2
0
1
226
Mia
Mia@MiaAI_lab·
Which local model is the best for Agentic Workflows for a single @NVIDIAAI DGX Spark? (or any other 96-128gb VRAM rig) After running 84 scenarios, 16 categories, 8 trials each, on Hermes-Agent style multi-turn tool orchestration, there is a very clear winner. 🏆 Qwen 3.6 35B A3B Q8_K_XL is #1 It’s the only model that hit perfect scores across the board with zero catastrophic failures. The Full Ranking: Qwen 3.6 35B A3B UD Q8_K_XL — 91.0 Qwen 3.6 27B NVFP4 — 89.0 Qwopus 3.6 27B Coder MTP — 85.2 DeepSeek V4 Flash Q2 — 86.5 Agents-A1 Q8_0 — 83.4 Gemma 4 26B — 81.4 Nemotron 3 Nano Omni 30B — 79.0 Bottom line: If you’re running agents locally on a DGX Spark or any 96-128GB rig in 2026, Qwen 3.6 35B Q8_K_XL is currently the move. Full report + deep dive 👇 github.com/MiaAI-Lab/Best…
Mia tweet media
English
111
45
558
37.2K
LannTheSmart
LannTheSmart@LannyTheClever·
@jxnlco GPT's search is simpler and more convenient. One advantage of the web version is the ability to use the Pro model; and once MCP is supported, I will be able to sync my Obsidian vault directly via conversation to discuss my journals and daily work.
English
0
0
0
10
jason
jason@jxnlco·
If you use Codex, is there any reason you still use ChatGPT? what do you use it for? how has it been better or critical for you?
English
1K
13
1.1K
256.7K
LannTheSmart
LannTheSmart@LannyTheClever·
@SpaceTimeViking I'm quite concerned that uncensored models generally perform worse than the original versions. Are there any versions that haven't been censored?
English
0
0
0
2
ÆON FORGE ✨
ÆON FORGE ✨@SpaceTimeViking·
Go get it! Validated and fully uploaded ready to run on Apple Silicon! github.com/AEON-7/Qwen3.6… Follow the QuickStart guide, or point your AI agent of choice at the Agents.md to set it up for you. MLX 8-bit for highest fidelity MLX FP4 for highest speed and good fidelity (Runs on as little as 24GB unified memory systems consuming ~17GB of memory when running)
ÆON FORGE ✨@SpaceTimeViking

Apple Silicon owners rejoice! MLX Qwen3.6-27B-AEON-ULTIMATE Models are uploading NOW! Unlock your local AI powers with a truly brilliant and uncensored model that will never refuse. Removal of censorship also unlocks some capabilities by removing censorship overhead. It's in a class of it's own and great at writing code but also handles multimodal input. github.com/AEON-7/Qwen3.6…

English
6
4
48
6.2K
向阳乔木
向阳乔木@vista8·
一个轻量快速的RSS免费客户端,还支持用自己的API key做 AI总结、问答。 感觉Papr是个不错的项目,地址和安装见评论区。
向阳乔木 tweet media
中文
8
5
51
12K
LannTheSmart
LannTheSmart@LannyTheClever·
@Saccc_c 这个字幕是 hyperframes 做的吗?有没有教程捏
中文
0
0
1
1.1K
Sac
Sac@Saccc_c·
Claude Code创始人Boris和团队成员的20分钟内部对话,信息密度极高 “不跟agent对话,跟loop对话,让loop替我来prompt” “从 4.6 起就不用 plan mode 了,全开 auto mode” “不需要prompt engineering和context engineering。给模型最精简的必要上下文和工具,让模型自己探索方案” “一半的工程在手机上完成,走路时和agent语音对话” 建议收藏看完,价值远超普通vibe coding教程
中文
116
146
804
109.7K
LannTheSmart
LannTheSmart@LannyTheClever·
@LufzzLiz 这个能接入到 本地Agent 吗?我想让它来操作公司的相关事务,不知道能不能做到。
中文
1
0
1
91
岚叔
岚叔@LufzzLiz·
这个开源项目有可以:Open-LLM-VTuber,一个 8.9k stars、1.1k forks 的开源本地 AI VTuber 框架,最新发到 v1.2.1,v2 也在规划重写。 它把 LLM、语音识别、TTS、Live2D 串成一套:你说一段,它听懂、回复、合成语音,再让角色开口。 也支持“可打断”:AI 说话时你直接插话,前端 VAD 检测到人声就停播,后端取消当前任务,把“被打断”写回上下文。 本地跑推荐有点硬件基础:M 系列 Mac、Nvidia GPU、比较新的 AMD GPU,支持 ROCm 更好;其他 GPU 也能折腾,或者用一个足够强的 CPU 顶上。 噪音处理比较现实,靠 VAD + 停麦,不吹电话级实时降噪。 适合折腾本地桌宠和 VTuber Agent 原型,公网/商用还得自己补鉴权、限流和测试。
岚叔 tweet media
中文
1
0
8
1.4K
LannTheSmart
LannTheSmart@LannyTheClever·
@mylifcc 是不是越复杂的项目收益越大呢?能不能通过 prompt 的方式,让agent生成更加结构化的、带上下文的 Handout 文档?
中文
0
0
0
14
lifcc
lifcc@mylifcc·
Handoff Debt——AI Coding Agent 的「交接债务」问题 当我们 benchmark Coding Agent 时,总是假设一个 agent 从头到尾 uninterrupted 搞定一个 SWE-bench 任务。但现实中的软件工程根本不是这样:任务会被打断、被 reassigned、被 review,然后由另一个人(或另一个 agent)接手。 这篇 arXiv:2606.02875 正式把这个现实痛点命名为 Handoff Debt(交接债务):前任 agent 留下的工作如果不够透明或不够结构化,后任 agent 就需要付出大量「重新发现」(rediscovery)成本。
lifcc tweet media
中文
5
3
9
948
辉小志
辉小志@imvihv·
codex的对话压缩功能做的很好,压缩后还能记得之前的大部分信息,问了下GPT,它的压缩其实并不是对话总结,压缩后的内容是人类不可读的、不透明的。所以我们确实不知道openai是怎么做到的,可能是训练了一个小模型专门做压缩。
中文
25
2
172
46.5K
逸尘
逸尘@gengdaJ·
太夸张了,阶跃星辰的stepaudio-2.5-asr和豆包 ASR(语音转文字)质量持平,但是价格要便宜整整五倍!我那动不动就几个小时的会议转写终于有救了!​ ​ 我之前抓取对标视频口播文案、总结视频内容和会议纪要,一直用的都是豆包的ASR,我以为豆包一块钱一个小时可能已经很便宜了,但是Step的这个ASR更夸张,一个小时就一毛钱!至于效果,如图所示,中英文识别能力完全相当!​ ​ 更牛逼的是,Step AI出了一个plan套餐,涵盖了TTS(语音转文字)、ASR、生图和平常Coding的大模型。所以,我现在日常工作就用100刀的Codex,然后再搭配一个Step最便宜的套餐,就能在特殊情况下操控Codex调用Step的API来实现语音转文字、文字转语音的效果,满足更多工作需求。​ ​ 我这里做了几个demo来演示下@StepFun_ai的这些模型用法:👇​ 1.“从爱到清醒的决裂”的原创宫廷男女对​ 2.抖音个人成长赛道百万爆款转写​ 3.Claude Code接入StepFun
逸尘 tweet media
中文
25
11
77
46.9K