

the watcher
49 posts

@wwtchr
Gifted Watcher. Festina lente. Observe, dissect, think twice — then act. Spot what haste blinds most people to. Building systems that actually endure







VCs used to worry frontier labs will kill applications. Luckily, with open source models, suddenly those application companies will be just fine, with their abilities to collect context and route models. Companies like replit, lovable, genspark used to have negative margins and terrible unit economics can now become profitable and competitive again. Expect to see more investments in applications, previously known as "wrapper" companies.




Been thinking about this: 1. Spotify have not publicly talked about engineering in quite some time 2. Many (most?) software engineers are paid Spotify customers, mildly annoyed at the product 3. So when they saw Claude + Spotify they immediately started to air grievances



Price per token != cost per task







Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests. We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort. Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research. Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again. Read our full blog: anthropic.com/news/redeployi…



Sonnet 5 initial impressions: >much more optimized for autonomous execution >better at planning, tool use and implementing complete tasks >more likely to do too much when you only want advice or a small change >sometimes ignores narrow constraints and jumps straight into implementation >stronger as an autonomous agent, but potentially worse as a collaborative coding assistant The paradox: ❌Optimizing models to replace the developer may make them worse at working alongside one. that's why you should stick to GLM-5.2, Kimi or smaller models for daily assistance... they may be less capable overall, but they’re more predictable and more likely to do exactly what you asked. Tip: Use Sonnet 5 for clearly scoped end-to-end tasks. For pair programming, precise edits and human-in-the-loop workflows, a more obedient model may actually be the better choice.

การมาล่าเงินพ่ออีลอน หลักๆไม่มีอะไรมาก แค่ฟอลติ๊กฟ้า ให้เยอะเข้าไว้ อย่าไปอัลฟอล อย่าไปมิ้ว มันต้องเปิดใจรับในความต่างของกันและกัน ฟอลแล้วมิ้วทำไปเพื่ออะไรก็ในเมื่อเราต้องฟอลติ๊กฟ้าเยอะๆ เพื่อให้มีปฏิสัมพันธ์และการมองเห็นกันมากยิ่งขึ้นไม่ใช่หรือ ค่อนข้างจะงงเวลาที่ฟอลกันแล้วมิ้ว ** ต้องทวิตเอง เป็นโพสต์บ้างคลิปบ้างรูปบ้าง อย่าแค่บ่นว่าทวีตอะไรไปก็ไม่มีใครเห็น เอาแต่ทวิตบ่น เอาแต่นอยด์ ** ไปคอมเม้นท์ให้หลากหลาย ไม่ใช่เกาะแต่แอคใหญ่อย่างเดียว ** รีไปเถอะ แล้วกดใจด้วย ไม่ว่าจะของใครแอคที่เราฟอล หรือแม้แต่ไม่ได้ฟอลกันก็ทำได้ ** ทั้งรี ทั้งกดใจ ทั้งคอมเม้นท์ ถ้าทำได้ยิ่งดี แต่ห้ามทำเหมือนบอท 📌 มีหลายแอคมาเล่นกับแอคเรา เราเลยอยากเข้าไปรีไปดันให้ แต่เข้าไปกลับพบว่า เค้าไม่ทวิตอะไรเลย นอกจากรีทวิต กับคอมเม้นท์ในทวิตคนอื่น อย่าลืมนะเราไม่มีทางรู้หรอกว่าทวิตไหนของเรามันจะแมส เพราะฉะนั้นทวิตไปเลย อย่าไปกังวลว่าจะมีใครมาตอบโต้รีทวิตให้เราหรือไม่ 🤟🏻 📌 สุดท้าย อย่าลืมว่า ถ้าเราซัพพอร์ตเขา เขาก็จะซัพพอร์ตเรากลับเช่นกัน 📌สุดท้ายนี้ ขอให้ทุกคนที่กำลังสร้างรายได้ในx สำเร็จตามเป้าหมายที่ตั้งไว้ 🙌🏻✌🏻🍀 #ติ๊กฟ้า





we launched code interpreters for deep agents last month. Basic idea is to let agents plan, delegate, and organize context using code instead of chained tool calls Code interpreters don't need a sandbox, but we still need a way to securely run that code! (and running untrusted code is a famously hard problem) Here's the writeup on how we're looking to do just that:





We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2. Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication. We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating. 🧵👇