Yi Meng

267 posts

Yi Meng banner
Yi Meng

Yi Meng

@yimengong

I like rust 🦀

Singapore Katılım Ocak 2025
238 Takip Edilen16 Takipçiler
Yi Meng retweetledi
Glauber Costa
Glauber Costa@glcst·
A new release of Turso is here. And the most important thing: this is our first release that no longer carries the beta label. Pekka goes into detail in the blog post below on *everything* that is included, but some of my favorite: * Concurrent writes can now be synced to the Cloud. Previously, syncing to the cloud only worked with WAL mode (serialized writes). We will be enabling this to our beta cloud users shortly. * Still on how we go beyond SQLite, you can now create arbitrary sequences, like you can do in Postgres. * There is now a .NET SDK, contributed by the great @awakecoding
Pekka Enberg@penberg

Turso 0.7 is now released 🚀 Faster concurrent writes, improved SDKs, and more! turso.tech/blog/turso-0.7…

English
4
10
175
16.5K
ThePrimeagen
ThePrimeagen@ThePrimeagen·
20% of my time with an 5.5 is saying: "Did you do it?" ... "Ok then do it"
English
134
22
1.3K
84.4K
Yi Meng retweetledi
Rodrigo
Rodrigo@rodro_pino·
#Juno is now more sexy for #starknet validators. Reduce your Mainnet DB size to under 100 GBs with pruning :) Less storage means more access to cheaper Validator bundles, make sure you are getting the best bang for your buck ;) Check it out! github.com/NethermindEth/…
Rodrigo tweet media
English
5
6
24
2.4K
Yi Meng
Yi Meng@yimengong·
@glcst you did him so dirty 😭
English
0
0
0
20
Glauber Costa
Glauber Costa@glcst·
In a functioning company, the role of the CTO is to define the product and technical direction. The goal of the CEO is to be able to articulate the company and its mission. As part of my job, I have to write sales deck, pitch decks, board decks, and always present our company in a truthful and honest manner. I love my job
Glauber Costa tweet media
English
8
0
76
5.4K
Yi Meng
Yi Meng@yimengong·
@cline Headlines are useful, but more context on the bug fix would be helpful. It would give better insight into where GLM 5.2 might perform well in.
English
0
0
0
167
Cline
Cline@cline·
We've kept hearing how GLM-5.2 beats Opus 4.8, and are skeptical of benchmarks - so we tested them on a real bug from the Cline repo. While both models fixed the issue, GLM was the winner in terms of cost and code quality: - GLM used twice as many tokens (GLM 1.1m vs Opus 660K) but cost half as much (GLM $0.41 vs Opus $0.81) - Opus finished quicker - 1.6 min and 12 tool calls vs GLM 4.7 min and 28 tool calls - GLM cleaned up dead code and verified the build compiled before completing. Opus didn't - it left type errors that passed tests but broke the production build. Both runs used the same Cline harness prompting and tools, so it seems GLM is RL trained to spend more tokens verifying its work before completing. Impressive work by the @Zai_org team!
Cline tweet media
English
228
610
7.7K
894.7K
Yi Meng
Yi Meng@yimengong·
deepseek chat instant response is pretty rapid
English
0
0
0
24
Yi Meng
Yi Meng@yimengong·
@TheAhmadOsman curious about why you ranked kimi 2.7 higher than glm 5.2
English
1
0
1
639
Ahmad Awais
Ahmad Awais@MrAhmadAwais·
ok kimi k2.7 code seems like a real deal. it's already dethroned k2.6 in every test i've thrown at it. still benchmarking, but my early money is on this being the new top open model.
Command Code@CommandCodeAI

Kimi K2.7 Code is now in available in Command Code. 10x free credits in Go. Our new #1 open mode in internal benchmarks. cmd update to v0.37.0 select via /model • 256K context 🍃 • 30% lower reasoning tokens than K2.6 ✅ • Open weights 1T-parameter MoE - 32B active ⚡

English
8
3
154
8.4K
Yi Meng
Yi Meng@yimengong·
@banteg he was so optimistic that I was actually interested in trying 🤣
English
0
0
0
859
banteg
banteg@banteg·
the only guy on the timeline happy with fable
Taelin@VictorTaelin

this is my personal singularity moment this post may sound like a paid ad. I only wish. I'm concerned, more so than happy. the world is changing, and, among the scenarios where AI goes terribly wrong, inequality is the most realistic, yet, the one Anthropic seems to be the least concerned about. I'm glad OpenAI is taking the opposite stance: *personal AGI for everyone*. I think this is a commendable position in the times we live. but who am I in the queue of the bread? anyway, Fable is here, so I'll just report my first-hour experience first of all, all my pet prompts are solved. → λ-calculus puzzles → bug questions → one-shot apps all are trivial to it. I don't have anything harder other than my ongoing work so, in the last several days, I've been toying with HVM5, a new interaction net evaluator with a faster loop. after writing the first version, I left 32 GPT-5 agents working for ~20 hours each. this resulted in up to 2x speedups, but the file size increased by 2-fold and quality decreased significantly. I then simplified the whole thing into an even simpler core, and left Opus 4.8 and GPT 5.5 optimizing it for 8 hours. Opus got a legit 6% - 34% speedup in most benches. GPT got better results, but, sadly, an unusable file. I then asked Fable to optimize it. 2 hours later, it landed a 1770% speedup in one case, 100%+ in other 4, and 22% in average. yes, in 2 hours it outperformed me, opus 4.8 and a swarm of gpt 5.5 agents, by one order of magnitude. that could not possibly be legit. "it must be hardcoding the benchmarks" (GPT trauma). so I read its explanation and what it did was, indeed, the most high impact optimization one could try first. seems like HVM5 was wasting a lot of time garbage-collecting unused branches of pattern-match nodes. I had optimized that for static mats, but not for dynamic mats. skill issue. Fable figured how to do it for these, resulting in a massive speedup in some benches but wait, is that *correct*? I'm not sure yet, it is credible, but this is the kind of thing that is very easy to get wrong on interaction nets. the problem is, when I was ready to start auditing Fable's solution so I could tell whether it was buggy or legit, it interrupted me to tell me it had found a massive bug on the code *I* had written. ... wait, what? so... for garbage collection purposes, I stored a bit on lambda term pointers that meant "the variable bound by this lambda has been freed, so, its lambda must free whatever argument it is applied to". that's fine. yet, on duplicator nodes, I also used the same bit to mean "one of the duplicated variables was freed, so, treat this dup as a passthrough no-op". so, if a lambda entered a duplicator, it would mistake the lambda's collection bit for its own, resulting in corrupted interaction! that's a mouthful, why I'm writing this? just so you can appreciate the sheer absurdity of what just happened. I didn't ask it to find bugs. I asked it for an optimization. and even if I did ask it to find bugs, this bug is so astonishingly subtle and specific, identifying it takes mastering the domain to an extent that it beyond even me. I'd easily need hours or days to fix it, *if* I ever came across it. chances are it would just go unnoticed. and Fable found it and fixed it like it was nothing, while it was busy adding a 17x speedup to a file that neither I, nor Opus 4.8, nor a fleet of GPT 5.5 managed to barely make 2x faster. oh and there is also another tab where it is also ripping through Bend's codebase and finishing everything I had to do I don't know what to say anymore this isn't about Anthropic or OpenAI, this is about our collective future as a species. the world is changing, and we need to be aware of it, and discuss how to handle this change. receipt below . . .

English
51
7
676
65.3K
Yi Meng
Yi Meng@yimengong·
ChatGPT downtime got me using Deepseek Chat instead 🤣
English
1
0
2
112
Yi Meng
Yi Meng@yimengong·
There seems to be more noticeable downtime relating to ChatGPT recently. Please fix @thsottiaux
English
0
0
0
103
Yi Meng
Yi Meng@yimengong·
@Marcia_Ong are you doing systems programming? why are you learning about bloom filters haha
English
1
0
1
40
Yi Meng
Yi Meng@yimengong·
@EnoReyes @matanSF Love that during a time when tokens are getting more expensive and limited, you folks are shipping features that help keep costs low 🙏🏻
English
0
0
0
113
Eno Reyes
Eno Reyes@EnoReyes·
Update, in the last 9 days since launch this has saved our customers over a hundred billion tokens already. In addition to the context savings this provides, this helps prevent unnecessary context bloat which lets Droid stay focused on its current task for longer.
Factory@FactoryAI

Introducing Deferred Context Engine in Droid.

English
9
12
265
42.2K
Yi Meng
Yi Meng@yimengong·
@thsottiaux I would love to see more specialised agentic coding benchmarks for real-world tasks like frontend, backend, devops, and systems programming. Those would feel much more useful when deciding whether to try something than a generic benchmark.
English
0
0
0
343
Tibo
Tibo@thsottiaux·
Do you still trust benchmarks or do you just listen to your friends? What makes you try a new model?
English
959
36
2.1K
236.4K
Tibo
Tibo@thsottiaux·
Codex Thursday has exceptionally moved to another day. Friday it is.
English
477
128
4.8K
476.9K