Chris Zehner

2K posts

Chris Zehner banner
Chris Zehner

Chris Zehner

@cbzehner

Your friendly neighborhood engineer. Functionally literate, inveterate dabbler, impatient debugger. pronouns: he/him/y’all https://t.co/YVb2i0bj2i

San Francisco, CA Katılım Temmuz 2008
1.4K Takip Edilen329 Takipçiler
Chris Zehner retweetledi
Merlin
Merlin@TheWizardTower·
One enormous advantage of LLMs is they're putting coders in the role of Engineering Managers, which gives them enough detachment to observe behavior and output neutrally, from a detached perspective, without their own ego in the mix. Which means all of the obvious advice (strong types, lots of tests, clean formatting, single concern function, etc) are becoming obviously true for a lot of people. They've always been true, it's just now obvious.
💻🐴Ngnghm@Ngnghm

Claude is surprisingly good at finding bugs in Scheme macros, yet in extremely bad taste at proposing fixes. Also, parentheses as such are no big deal to it, but deep nesting is, which can be alleviated by suitable macros and coding style—same as with humans, really.

English
8
9
220
21.9K
Chris Zehner
Chris Zehner@cbzehner·
@kunchenguid Have you experimented with lowering the effort levels further to extend quota usage? I’m seeing fairly good results with Fable/Sol on low effort and escalating up to medium in the same way that I was using medium -> xhigh on GPT-5.5.
English
1
0
2
314
Kun Chen
Kun Chen@kunchenguid·
a bit confused how to choose between fable, gpt 5.6 and grok 4.5? my updated agent org chart is here! this is from my past 3 days tinkering various permutations of these models and i finally arrived at a really smooth setup firstmate: GPT-5.6-Sol xhigh in Pi or Grok 4.5 high in Grok Build both are fast enough to make the orchestration loop feel genuinely interactive. these two harnesses also work very well with firstmate, while codex has the foreground polling limitation secondmates: fable 5 in Claude Code. it’s a great fit for complex product and technical design, where depth matters more than latency crewmates: Grok 4.5 high in Grok Build handles bug fixes. GPT-5.6-Sol high in codex handles feature development, with fallback being opus 4.8 in claude code dynamically decided based on quota-axi things that require real-time information from X always uses grok, while image generation tasks always use codex. harness matters everything then passes through /no-mistakes on GPT-5.6-Sol medium for consistent, cost-efficient adversarial review and fixes smooth sails!
Kun Chen tweet media
English
36
5
129
8.3K
jon allie
jon allie@jonallie·
I had the weirdest experience today where GPT 5.5 was producing much worse code than I'm used to (complex and over engineered..full of useless helper functions replacing 3 lines of straightforward code). After getting increasingly frustrated, I finally noticed that I'd somehow ended up on "medium" reasoning instead of "low", and switching it back to low, and asking it to review the code it just wrote yielded a much better result.
English
8
0
58
21.1K
Chris Zehner
Chris Zehner@cbzehner·
And GPT 5.5 medium through xhigh.
English
0
0
1
33
Chris Zehner
Chris Zehner@cbzehner·
Sonnet medium + Fable low are all you need. Revisit weekly.
English
1
0
2
56
Chris Zehner
Chris Zehner@cbzehner·
What are y'all using Haiku for? I'm currently only using it for programmatic data transformation (unstructured -> json) Every time I push for delegation, I end up with Sonnet subagents.
English
0
0
0
37
Chris Zehner
Chris Zehner@cbzehner·
Good agent prompts are more like emails than chats. Write out a detailed message. Think about it. Think about it again. Hit send. ...realize you didn't think about it long enough
English
1
0
2
55
Chris Zehner
Chris Zehner@cbzehner·
@thosiawa I'm using the json log files from the sessions as a heuristic. The status line reads the mtime of the session transcript (which Claude Code appends to on every turn) and diffs it against the system clock to estimate whether the prompt cache is still warm. github.com/cbzehner/cship…
English
0
0
1
50
Tom Hosiawa
Tom Hosiawa@thosiawa·
@cbzehner Oh, nice. Were you able to get it to check the cache time? Or do you use something else?
English
1
0
0
21
Chris Zehner
Chris Zehner@cbzehner·
After spending a day pairing models to effort levels, Fable ages like a fine wine - still going days after uncorking thanks to low/medium effort usage instead of my prev (Opus, high)
English
0
0
0
41
Chris Zehner retweetledi
Kun Chen
Kun Chen@kunchenguid·
my hot take on how much AI code we should review - you should review as much code from AI as your engineering director reviewed your code before AI here’s the chain of thought: - why do we even use AI to code? it’s to allow us to ship more - how much more should a single developer be able to ship now, compared to pre-AI? i see us going from 1-10x in the past 3 years, and on a trajectory to hit the 100x magnitude soon - that means every developer is going to own as much scope as a pre-AI director of engineering - i haven’t met a single eng director who said their team’s codebases were perfect and exactly how they would like it to be. why? because people who try to achieve that will fail to become a director - how do directors handle that level of complexity? it’s absolutely not by reviewing and micro-managing every engineer’s code. it’s through managing the culture, workflows, resource allocation, guardrails and measurable outcomes - when a director sees the team struggle on productivity or quality, they might lean in and try to understand the state of the codebase to develop some intuition for how to improve things systematically. even this is often done with the help from their principal engineers - i believe this is the right balance for how we should manage AI so, if we want to get a massive boost from AI, we must be prepared to operate in a way that allows us to manage much higher complexity, which requires that we remove ourselves as a bottleneck and manage the outcome at a different level shape your AI agents’ workflows - are they doing adversarial review? are there good automated tests? are they presenting evidence before shipping? are they doing phased rollout? are there good metrics to catch problems? survey your agents for feedback - ask them to reflect on their past sessions and report biggest problems causing them to struggle, and allocate enough tokens to get those problems fixed focus on outcomes - are your agents doing busy work? do you truly understand customer requirements and what work is worth doing? are your agents’ work generating the business outcome you expect? that’s how we truly scale
English
35
29
267
16.7K
Chris Zehner retweetledi
Wade Foster
Wade Foster@wadefoster·
We're killing the DM at @Zapier. Starting with the executive team. We've long held Default to Transparency as a value. That value has largely encouraged communication in public channels. But as the company grew, DMs are a hard habit to resist and break. But every DM is a gap in our Shared Brain. It's context that is lost for humans and AIs. As a result the cost of DMs keeps going up. So earlier this year I posted about our exec transparency leaderboard. The leaderboard has become quite the competition internally… I'm 3rd today. My co-founder @BryanHelmig has held the top spot as long as I can remember… It sets a standard for the rest of the company. In fact, since last year we’ve seen the % of Slack messages in public channels go from 33% to 46%. What the leaderboard measures Transparency is a team sport, and a disinfectant. Every month we track what percentage of our execs' Slack messages happen in public channels versus private DMs. When your CEO debates strategy in a DM, that decision is invisible to every agent and every team that needs to know what was decided and why. The decision happens but the reasoning vanishes. When that conversation happens in a channel, it stays. New hires can search it, agents can read and verify it, etc. Your Shared Brain knows what's true now: ask it a question and the answer reflects the latest reality. Taking It to the Next Level Reducing DMs are one way to increase transparency and open up context for humans and AI, but there are other mechanisms that help too. Three things beyond the leaderboard: 1. Meetings get recorded, transcribed, and become queryable 2. We run a shared skills library. Anyone on the team can encode a workflow they've figured out into a skill and share with the team 3. And we keep score. It's a silly scoreboard, but it subtly drives positive behaviors Raising Your Ambition In order to get the most of AI in your company, the AIs need context. So making your context queryable is one of the most practical moves you can make to improve the effectiveness of your AI agents. P.S. I’m coming for #1, Bryan...
Wade Foster tweet media
English
94
54
889
356.3K
Chris Zehner
Chris Zehner@cbzehner·
@GergelyOrosz If you’re trying to fit in a quick swim in SF, there’s SwimFrancisco.com You can check the public pool schedule without needing to navigate to all the individual rec websites
English
0
0
2
314
Gergely Orosz
Gergely Orosz@GergelyOrosz·
Somewhat surprised to learn the that neither San Francisco, nor the Bay Area has a publicly accessible, Olympic-sized (50m) pool. Such pools of course are expensive to build and maintain, but I can’t think of any other major metro area that doesn’t have at least one.
English
88
6
554
95.4K
Chris Zehner retweetledi
Anish Moonka
Anish Moonka@anishmoonka·
A sword has been stuck in a rock in Tuscany since the 1180s. In 2001, a chemist from Italy's professional skeptics society, the people whose whole job is exposing fakes, drilled into the stone to test it. The metal came back medieval. The knight was Galgano Guidotti, from Chiusdino, a hill town near Siena. Around 1180 he quit fighting. He rode up a nearby hill, drove his sword into the bedrock up to the hilt, and lived next to it until he died in 1181, at 33. Locals built a round chapel over the rock, finished by 1185, and the sword stayed put. For centuries the obvious explanation was fraud. An iron prop cemented in to impress medieval travelers, the kind of thing you'd expect from a town with a relic to sell. Then Luigi Garlaschelli went in to check. He is a chemist at the University of Pavia and a member of CICAP, the Italian committee that takes apart exactly these claims for a living. His team drilled a small hole into the rock and pulled samples from the buried half of the blade, metal nobody had touched in 800 years. The samples went through atomic absorption spectroscopy and neutron activation analysis, a nuclear technique that reads a metal's exact elemental fingerprint. No modern alloys anywhere. The trace elements matched iron slag from medieval foundries whose waste still litters the ground around the site, and the blade shape matched swords catalogued from the late 1100s. The hilt above the rock and the section buried inside it are one continuous piece, so nobody jammed a fake handle into a boulder. My favorite detail is the one they left alone. Radar picked up a cavity directly under the sword, about six and a half feet by three, the dimensions of a grave. Nobody has ever opened it. One caveat, though: the sword snapped during a theft attempt in 1960 and got repaired and re-set with cement, so the thing under the case today is patched, even if the metal is medieval. The timeline bothers me in a good way. The earliest sword in the stone in the King Arthur legend shows up in a French poem from around 1200. Galgano's sword was in its rock by 1181, chapel and all by 1185. Scholars still argue over which story fed which, but the object beat the legend by roughly two decades, and it was planted by a man walking away from knighthood rather than claiming a throne.
Historic Vids@historyinmemes

According to tradition, a Catholic saint from Tuscany thrust his sword into a stone after experiencing a vision of the Archangel Michael in the 12th century. A scientific analysis conducted in 2001 found that the visible portion of the sword and the section embedded within the stone are genuine and form a single, continuous artifact.

English
7
72
526
80K
Chris Zehner
Chris Zehner@cbzehner·
How it started: Woo! Fable's back 🙌 How it's going: You've hit your session limit · /upgrade to increase your usage limit.
English
1
0
2
34
Chris Zehner retweetledi
Feifei Qiu
Feifei Qiu@feifei_qiu·
Sev A incident today: We lost a child's hard work story today😭. Here's what happened and the hard lesson I learned about building software. At our kids' creative writing workshop, a child forgot his login, so we let him use our shared account. Another kid ended up on the same account on a different device. The first kid spent 40 min writing his chapterbook script, made all the illustrations, hit save, and moved on. He thought it was saved. It wasn't. At some point his draft stopped saving. Every save came back "draft not found." (while the error didn't surface all the way to client). I still can't tell you exactly why, or what happened to it, because we didn't log enough to trace it, and that's part of the problem. His story only ever lived in that one browser tab's memory. We never wrote it to the device or saved it on the server. So when the tab refreshed, that memory was wiped, and there was nothing left to recover. I spent 1.5 hrs trying to get it back. Got lucky with the illustrations, they were still sitting in temporary storage. But his actual script, the words he spent 40 mins wrote himself, was gone for good😭😭😭. What's the learnings? 1, Stop the loss at the source. His script only ever lived in that browser tab, never saved to the device, never to our server, not even in a log I could pull up later. So when I went digging, there was nothing to find, there were stateless for requests using the raw script. The moment a kid writes something, the raw text should be captured somewhere durable, so there's always a copy that can't just vanish. 2, Make loss recoverable. We hard-delete drafts. It should be soft delete, so a "lost" draft is one click from coming back instead of gone forever. 3, Make everything traceable. I couldn't tell what went wrong, or who or what removed his draft accidentally, because we don't log the right things. 4, Fail loudly. The worst part wasn't losing the data. It was telling a kid "saved" when it wasn't. Silent failure is how a small bug turns into an unrecoverable loss. Building the happy path for the app is easy, building all the edge cases and make the app really robust and reliable are the real work
English
2
1
10
1.3K
Chris Zehner retweetledi
AYI
AYI@ayi_sanchez·
Few weeks ago, I found something mislabeled in a second hand shop. Turned out to be a box full of slides from Vietnam war (1968-1970). I managed to scan around 300. Here its a small selection of somebody's memories I bought for 5,99 🧵
AYI tweet mediaAYI tweet media
English
406
3.3K
40.8K
1.3M