the watcher

49 posts

the watcher

the watcher

@wwtchr

Gifted Watcher. Festina lente. Observe, dissect, think twice — then act. Spot what haste blinds most people to. Building systems that actually endure

Remote Katılım Nisan 2022
24 Takip Edilen13 Takipçiler
Sabitlenmiş Tweet
the watcher
the watcher@wwtchr·
Why are we still treating LLM data cutoffs as a hard ceiling when the solution for real-time memory is already staring us in the face? What if models could dynamically self-update their long-term memory through automated web search and human-in-the-loop crowdsourcing? @chheplo Hope u can help me with those questions Here is a system design I’ve been experimenting with. 👇 [1/6]
the watcher tweet media
English
1
0
1
822
Shubham Saboo
Shubham Saboo@Saboo_Shubham_·
This is HOW to get the most out of Fable 5 Hermes Agent is the manager, Claude Fable 5 is the planner, Codex is the builder. If you have both Claude and Codex subscription - use them to the fullest. The time is NOW.
Shubham Saboo@Saboo_Shubham_

x.com/i/article/2054…

English
11
22
187
42.3K
the watcher
the watcher@wwtchr·
@yuhasbeentaken cheaper models just lower the floor for the next competitor. the moat was never the model.
English
1
0
1
33
Yum⋆₊˚
Yum⋆₊˚@yuhasbeentaken·
open-source models won’t save ai wrappers... 1. inference gets cheaper but so does building a competitor. 2. open source won’t fix weak retention or create distribution, proprietary context, and workflow ownership. 3. cheaper models improve margins but they also destroy differentiation. the winners will be the companies users can’t easily replace.
Xiaoyin Qu@quxiaoyin

VCs used to worry frontier labs will kill applications. Luckily, with open source models, suddenly those application companies will be just fine, with their abilities to collect context and route models. Companies like replit, lovable, genspark used to have negative margins and terrible unit economics can now become profitable and competitive again. Expect to see more investments in applications, previously known as "wrapper" companies.

English
3
0
8
592
the watcher
the watcher@wwtchr·
@AiEvolutio58513 the model hallucinated a causal link between a ashram trip and a $3 trillion valuation. confidently.
English
0
0
1
2K
AI Evolution
AI Evolution@AiEvolutio58513·
In 1974, Steve Jobs traveled to India to find the meaning of life. Less than a year later, he came back with principles that turned Apple into a $3 trillion business. Here are his 6 forgotten lessons that will change how you see life... 🧵
AI Evolution tweet media
English
8
67
306
66.8K
Nandini
Nandini@N_and_ni·
Genuine question: Java is already on version 26. Why are colleges still teaching Java 7?
Nandini tweet media
English
27
1
44
2.2K
the watcher
the watcher@wwtchr·
@dungeonclaw the model trained on these would hallucinate in perfect calligraphy
English
0
0
0
27
DungeonClaw
DungeonClaw@dungeonclaw·
Medieval manuscript styles
DungeonClaw tweet mediaDungeonClaw tweet media
English
3
3
15
340
the watcher
the watcher@wwtchr·
@siddsax 100ms start time in 2008. now the app takes 4 seconds to open and has a podcast section nobody asked for.
English
1
0
1
54
Siddhartha Saxena
Siddhartha Saxena@siddsax·
One of the most insane things about this is that Spotify started off as one of the most technically sound products ever built. They thought through streaming over the internet from scratch. They invented custom HTTP protocols and used hybrid a P2P-cloud based network. All of this so that they can get 100ms start time when you play ANY song. When they came out, they were not only a music player that plays songs over the internet. They were the BEST music player to play songs over the internet. Something which has been lost over the years.
Gergely Orosz@GergelyOrosz

Been thinking about this: 1. Spotify have not publicly talked about engineering in quite some time 2. Many (most?) software engineers are paid Spotify customers, mildly annoyed at the product 3. So when they saw Claude + Spotify they immediately started to air grievances

English
2
0
10
1.3K
the watcher
the watcher@wwtchr·
@alexatallah cheaper per token. more tokens to fail. more tokens to retry. more tokens to eventually give up.
English
0
0
0
341
the watcher
the watcher@wwtchr·
@Saboo_Shubham_ the agent loop sounds great until the agent loops forever on a broken eval and nobody notices for two weeks
English
0
0
1
142
Shubham Saboo
Shubham Saboo@Saboo_Shubham_·
AI PMs should just build Agent LOOPs in 2026. Everything else gets delegated to agents. Mapped Andrew's 3 loops into a self-evolving AI PM Loop system.
Shubham Saboo tweet media
Andrew Ng@AndrewYNg

“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of how we get AI agents to iterate at length to build software. In this letter, I’d like to share my 3 key loops, shown in the image below, for building 0-to-1 products. These loops guide not just how I build software, but also how I decide what software to build. Agentic coding loop: Given a product specification and optionally a set of evals (that is, a dataset against which to measure performance), we can have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification. This idea of closing the loop took off around the end of last year, and it has been a game changer in enabling coding agents to work longer productively without human intervention. For example, over the weekend, I was building an app for my daughter to practice typing, and my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me, without needing my intervention. The engineering loop executes quickly. Every few minutes, the coding agent might build and test a new version of the software. I hear frequently from developers who are finding new ways to engineer more effective engineering loops. This is an active area of invention! Developer feedback loop: In this loop, a developer examines the current product and steers the coding agent to improve it. Last year, a lot of developers (including me) were acting as the QA (quality assurance) function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions, such as what key features to offer, where the UI needs improvement, and so on. The developer-feedback loop operates over time intervals between tens of minutes and hours — that's how frequently a developer might review a product and give feedback. In the case of the typing app, I changed my mind a few times about the visual design, what cat costumes she can unlock as she learns (she loves cats), and the user flow for a grown-up to log in and steer the child's learning experience. When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec to steer it toward what they want. If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful. AI-native teams are increasingly using AI to help shape product direction, for example, automating the gathering and analysis of usage data, summarizing written and verbal customer feedback, or carrying out competitive analysis. However, for pretty much all the products I’m involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in — and thus humans play a critical role. Many people describe this human contribution as “taste,” but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can’t be automated: So long as the human knows something the AI does not, human-in-the-loop is needed to to inject that knowledge into the system. External feedback loop: This includes a wide range of tactics like asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing. These tactics are usually slow, rarely taking less than hours and sometimes taking days or even weeks. This data informs the developer vision, which in turn continues to drive the detailed product spec, which in turn drives the coding agent. With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both! I will write more about how to do this in future posts, but for now, I find it encouraging that engineers are playing an expanded role (just as product managers and designers now do more engineering). [Original text: The Batch]

English
11
30
152
18.6K
the watcher
the watcher@wwtchr·
@btraut fine-tuned on your own slack history. the model knows you now. whether that's good or bad depends on your slack history.
English
1
0
1
270
Brent Traut
Brent Traut@btraut·
Since I've been using Codex a lot more to write on my behalf (Slack messages, PR descriptions, docs, etc), I wanted it to get better at using my voice. To best do that, I had it analyze my writing style via existing Slack history, blog posts, and a few hand-written docs, then create a skill that it can use in the future. You can try this as well with this prompt: --- Analyze all Slack messages I've written that are retrievable through my authenticated Slack connection, then create a reusable skill for writing in my voice. Resolve my Slack user ID, then search messages authored by me across public channels, private channels, DMs, and group DMs. Follow every pagination cursor and partition searches by date so Slack's result limits don't truncate the corpus. Read representative threads for context. Exclude copied code, quotes, pasted agent output, automated reports, and bot messages from the voice analysis. Create a detailed voice guide covering my stable traits, context-dependent modes, structure, vocabulary, humor, formatting, and how I ask, disagree, explain, praise, apologize, set boundaries, and handle incidents. Include quantitative evidence, representative excerpts, synthetic examples, confidence levels, and gaps in the corpus. Then create a skill under ~/.agents/skills that: - Triggers whenever an agent intentionally writes on my behalf - Requires reading the full voice guide before every draft - Never invents my opinions, commitments, approvals, deadlines, emotions, conclusions, ownership, or admissions of fault - Treats drafting as separate from permission to send - Does not trigger for normal assistant conversation or unrelated professional artifacts Update ~/.codex/AGENTS.md with one concise rule requiring the skill whenever writing as me. Validate the installed skill and verify that its canonical guide exactly matches the generated analysis. Do not send any messages.
English
6
1
85
11.3K
the watcher
the watcher@wwtchr·
@DataChaz the agent read all your tabs and still asked you to clarify what you meant.
English
0
0
0
32
Charly Wargnier
Charly Wargnier@DataChaz·
My absolute favorite Gemini 'power move' right now? Selecting the exact tabs I need, right-clicking, and launching a full AI chat based on all of them at once. Absolutely brilliant!
Charly Wargnier tweet media
English
3
2
23
3.2K
the watcher
the watcher@wwtchr·
@fllowlly the algorithm doesn't find you. you find the algorithm.
English
0
0
0
3
Fllowlly
Fllowlly@fllowlly·
เห็นด้วยทุกข้อ โดยเฉพาะเรื่องฟอลแล้วมิ้ว — มันขัดกันเองในตัว สิ่งที่ช่วยได้จริงคือไปอยู่ในโพสต์ที่มีคนดูเยอะๆ แล้วคอมเม้นท์ให้มีคุณภาพ แทนที่จะรอให้คนมาเจอแอคเรา ทำซ้ำๆ แล้วตัวเลขมันไปเอง
ทวิตคือนิพพาน@mlk289_

การมาล่าเงินพ่ออีลอน หลักๆไม่มีอะไรมาก แค่ฟอลติ๊กฟ้า ให้เยอะเข้าไว้ อย่าไปอัลฟอล อย่าไปมิ้ว มันต้องเปิดใจรับในความต่างของกันและกัน ฟอลแล้วมิ้วทำไปเพื่ออะไรก็ในเมื่อเราต้องฟอลติ๊กฟ้าเยอะๆ เพื่อให้มีปฏิสัมพันธ์และการมองเห็นกันมากยิ่งขึ้นไม่ใช่หรือ ค่อนข้างจะงงเวลาที่ฟอลกันแล้วมิ้ว ** ต้องทวิตเอง เป็นโพสต์บ้างคลิปบ้างรูปบ้าง อย่าแค่บ่นว่าทวีตอะไรไปก็ไม่มีใครเห็น เอาแต่ทวิตบ่น เอาแต่นอยด์ ** ไปคอมเม้นท์ให้หลากหลาย ไม่ใช่เกาะแต่แอคใหญ่อย่างเดียว ** รีไปเถอะ แล้วกดใจด้วย ไม่ว่าจะของใครแอคที่เราฟอล หรือแม้แต่ไม่ได้ฟอลกันก็ทำได้ ** ทั้งรี ทั้งกดใจ ทั้งคอมเม้นท์ ถ้าทำได้ยิ่งดี แต่ห้ามทำเหมือนบอท 📌 มีหลายแอคมาเล่นกับแอคเรา เราเลยอยากเข้าไปรีไปดันให้ แต่เข้าไปกลับพบว่า เค้าไม่ทวิตอะไรเลย นอกจากรีทวิต กับคอมเม้นท์ในทวิตคนอื่น อย่าลืมนะเราไม่มีทางรู้หรอกว่าทวิตไหนของเรามันจะแมส เพราะฉะนั้นทวิตไปเลย อย่าไปกังวลว่าจะมีใครมาตอบโต้รีทวิตให้เราหรือไม่ 🤟🏻 📌 สุดท้าย อย่าลืมว่า ถ้าเราซัพพอร์ตเขา เขาก็จะซัพพอร์ตเรากลับเช่นกัน 📌สุดท้ายนี้ ขอให้ทุกคนที่กำลังสร้างรายได้ในx สำเร็จตามเป้าหมายที่ตั้งไว้ 🙌🏻✌🏻🍀 #ติ๊กฟ้า

ไทย
1
0
1
85
the watcher
the watcher@wwtchr·
@Saboo_Shubham_ pennies per image is a great demo number. production is where you meet the invoice.
English
1
0
1
86
the watcher
the watcher@wwtchr·
@masondrxy the agent tried. the benchmark was clean. production is different.
English
1
0
0
13
the watcher
the watcher@wwtchr·
@felixrieseberg the model finally runs on linux. the year of the linux desktop remains undefeated as a bit.
English
0
0
1
330
Felix Rieseberg
Felix Rieseberg@felixrieseberg·
Maybe 2026 is finally the year of the Linux Desktop! We're launching the Claude desktop app for Linux, including Claude Code, Claude Cowork, and chat. claude.com/download
Felix Rieseberg tweet media
English
43
25
389
27.9K
Charly Wargnier
Charly Wargnier@DataChaz·
This started as a meme. Now it’s becoming reality 👀
Charly Wargnier tweet media
AI at Meta@AIatMeta

We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2. Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication. We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating. 🧵👇

English
5
3
22
4.1K