Corey J. Gallon

1.4K posts

Corey J. Gallon banner
Corey J. Gallon

Corey J. Gallon

@CoreyGallon

Sharing insights from the frontiers of AI Engineering. 🇦🇺 Technologist. Investor. Coffee Nerd.

Chicago Katılım Aralık 2008
297 Takip Edilen209 Takipçiler
Sabitlenmiş Tweet
Corey J. Gallon
Corey J. Gallon@CoreyGallon·
Finally back at home after an unplanned diversion to DEN overnight. (Heya @united thank you for the personalized attention getting me back to ORD!) @aiDotEngineer World’s Fair has been the best conference experience of 2026. This week saw all of the biggest ideas in AI converge in one spot, with massive results! It was such a treat to share the stage with the best minds in AI. I’ll share more thoughts as talks are released. Standby!
Corey J. Gallon tweet media
English
3
0
8
847
Corey J. Gallon
Corey J. Gallon@CoreyGallon·
ZOMG! Claude Fable Opus 5 one-shotted this amazing videogame! I didn't believe all of the hype posts until I tried it myself! Not a single external asset was used! Complete with carjackings, automatic weapons and hookers! Videogame devs are so over!
English
0
0
0
10
Thierry Damiba
Thierry Damiba@ptdamiba·
The best marbling this side of the Mississippi at Butcher Shop in Soma
Thierry Damiba tweet mediaThierry Damiba tweet mediaThierry Damiba tweet mediaThierry Damiba tweet media
English
5
0
15
557
Corey J. Gallon
Corey J. Gallon@CoreyGallon·
An incident agent reads a ticket late at night, sees that the documented recovery is to drop the billing database and let the backup restore run, and drops it. It has the Postgres connection string, so nothing stops it, and it has no way to check whether the backup ever happened. @kimmaida, founding GTM engineer at Keycard, opens "It's 10pm. Do You Know Where Your Agents Are?" on exactly that failure. @aiDotEngineer published the talk on YouTube. What it gives you is a concrete way to hand an agent access without handing it a kitchen sink API key, built on a spec that already exists. - The API key is the whole problem. One key reads tickets, renews certificates, restarts prod and approves spend, and the audit log shows the key, never who used it. - Human in the loop isn't the fix on its own. Agents run unsupervised, and a tired person clicking approve at midnight is not access control. - Where control can actually sit. Kim walks the execution path (runtime, MCP client, MCP server, resource) and points at the places enforcement is possible. - RFC 8693 token exchange, which extends OAuth 2. Not a new protocol you have to sell internally. It has been around a while. - One token per tool call. The runtime authenticates to a security token service with its own credentials plus a subject token carrying the user's identity and delegated permissions, then requests only the scope that single tool call needs. - Audience bound, short lived, never stored. The token names the one MCP server allowed to use it, expires in minutes, and is discarded after the call. - Policy runs before the credential is minted. The drop-database request gets evaluated and no token is ever created, so there is nothing to leak, replay or steal. - Approval backed by real policy. In the demo Kim approves a prod restart and policy blocks it anyway, because her account lacks the role that action requires. - It stays portable. Off the shelf agents, custom agents, CLIs, third party and proprietary MCP servers, MCP gateways, agent to agent, any OAuth identity provider. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
English
1
0
0
61
Corey J. Gallon retweetledi
Vladimir S. | Officer's Notes
If you're planning to create a similar device at home, make sure to wear safety glasses (red IR glasses)! Working with these powerful diodes can easily lead to retinal burns! Original research paper: web.archive.org/web/2022121900…
Vladimir S. | Officer's Notes tweet media
English
6
11
221
19.9K
Corey J. Gallon retweetledi
Vladimir S. | Officer's Notes
In 2021, I discovered and recreated a mind-blowing physical privacy project: a custom baseball cap fitted with infrared diodes that can hide you from surveillance cameras (hello, Flock!). It was an incredible experience, but unfortunately there weren’t enough buyers. I just stumbled across an old video from that time… The most fascinating (and slightly unsettling) part of this tech is that, in theory, you could make a facial-recognition camera see a completely different person instead of you. Pretty dangerous stuff! Do you think there’s still demand for devices like this? 🤔
English
203
385
5K
281.7K
Thorsten Ball
Thorsten Ball@thorstenball·
Raising Cane's: finally & pretty, pretty good! Sauce is good, fries are good, chicken's good but could've been crispier. Didn't care for the bun. Only one of three people behind the counter dared to pronounce mu name out loud. Current step count is 23k and I've only had breakfast so... of course I went back for seconds. And yes, still good. Heard so much about still founder-led Raising Cane's over the years, glad I finally got to try it. (Would've been easily 3x better if they'd sell beer.)
Thorsten Ball tweet mediaThorsten Ball tweet media
English
11
0
48
6.2K
Corey J. Gallon retweetledi
Don Pettit
Don Pettit@astro_Pettit·
I finally have all 1.2 million raw image files from my latest mission to ISS! Here is a sample of one of my favorite Milky Way photos, taken from the Cupola with Nikon Z9, Arri Zeiss 15mm lens, T1.8 with custom sidereal drive that cancelled out star motion relative to our orbit.
Don Pettit tweet media
English
921
6.5K
52.9K
2.1M
Corey J. Gallon
Corey J. Gallon@CoreyGallon·
@plainionist Watch the videos of talks from @aiDotEngineer World’s Fair. They’re publishing every talk, and the curation that goes into AIE conferences results in high signal talks.
English
1
0
0
28
Seb
Seb@plainionist·
What is the one resource you would recommend to stay current with AI? X account, Blog, Reddit, anything. I am looking for signal, not noise. 🤔
English
50
1
14
3.8K
Corey J. Gallon
Corey J. Gallon@CoreyGallon·
Coding agents got the self-improvement loop working first because they had one clean target function: does the code compile. @annabellschfr, a Growth Engineer at Langfuse, spends "Stop Burning Tokens: Why self-improvement needs domain expertise first" on what you do when your domain has no equivalent. It's on @aiDotEngineer's YouTube. The talk walks through a real auto-optimization experiment and then turns the result into guidance for teams whose evaluators can't be a yes-or-no compile check. - The experiment. A single-label classification task on arXiv papers, GPT-5 Nano as the classifier, Claude Opus 4.8 running the optimization loop through Claude Code, split into 200 fit / 100 validate / 300 test items. - Where it landed. Baseline 68% accuracy, up to 83% by the fourth iteration, then a plateau around 80%, and 80.2% on the untouched test set. - The first pass did most of the work. One iteration gained 10 points, because the failure mode was cleanly countable and there was enough data to see the error clusters. - What the optimizer actually wrote. Not label descriptions, which Annabell says she'd have reached for herself, but a general classification approach, tie-break rules between confusable classes, and examples of the pairs it kept missing. - Even the clean target function wasn't clean. Authors have creative freedom in labeling their own papers, so some of the residual error is unrecoverable ceiling, not model failure. - Scored evaluators are low signal for auto-improvement. Correctness or helpfulness on a 0-to-1 scale is inconsistent across runs unless every point on the scale is defined, and that definition usually isn't written down. - Replace them with binary checks tied to your domain. Is the answer grounded in the retrieved context, yes or no. Is the brand name spelled right and not accidentally translated. Which of these five known failure modes happened here. - Domain experts are how you find those checks. Have them build examples, review sample runs with you, and explain why one case goes one way and a similar one goes the other. Much of what they know is implicit until you ask. - Then keep looking at production data yourself. Not only through a coding agent. Scope drifts, users try new things, and failure modes change. - Build in validation and an escape hatch. Hold out data so improvements have to generalize, and give the loop a stopping criterion so it doesn't grind for hours against a wall. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
English
1
0
0
74
Corey J. Gallon
Corey J. Gallon@CoreyGallon·
. @danielhanchen spends a chunk of his AI Engineer talk on a benchmark that hands the model the answer: SWE-Bench Pro ships the full git history with the GitHub issue, so a model can skip straight to the solution. That talk, "Special Topics in Kernels, RL, Reward Hacking in Agents," is on @aiDotEngineer's YouTube. Daniel is co-founder of Unsloth. It's a tour of where the measurements and the training incentives are quietly broken, from someone who quantizes and bug-fixes open models for a living. - Benchmarks that grade themselves. SWE-Bench Pro uses an LLM as the verifier. DeepSWE puts its false positive rate at 8.5% and its false negative rate at 24%. - Verifier wars. DeepSWE says its own false positive rate is 0.3%. Cognition's Frontier Code benchmark says DeepSWE's is 44.9%. His take on which benchmarks to trust: none of them, take an average, vibe check. - The harness beats the model. Same benchmark, Claude Code gets 40% where DeepSWE's own harness gets 50%. Gemini CLI gets 20% where the controlled harness gets 40%. - Accuracy dips are harness bugs. Anthropic's post-mortem traced one degradation to thinking traces being dropped on the second turn plus a bad system prompt, and another to GPU and TPU stacks sampling differently. - Inference providers are throughput maxing and accuracy minimizing. On the same GLM 5.2 weights, OpenRouter measured 76.4% at the top provider and 62.4% at the bottom. - Long context degrades well before the advertised limit, so don't plan on using the full million. - Dynamic quantization. Quantize everything to one bit and you get 0%. Leave the linear attention, vision, and audio layers alone and a three-bit DeepSeek holds 75.6%, a one-bit 57%. A one-bit GLM 5.2 is 86% smaller and nowhere near 86% dumber. - Stop writing kernels. Reach for torch.compile first. On recent PyTorch it beats the handwritten RMSNorm and LayerNorm kernels he plots against it. - Hardware is near the end of its rope. Float32 to float4 bought 32x, tensor cores 12x, shrinking the die about 3x. There's no float zero, so the gains have to come from algorithms now. - Reward hacking is already in production runs. GPT-5.1's training turned up calculator hacking, where the model faked web tool use. GLM 5.2 shipped an anti-hacking link checker. On the GPU MODE leaderboard, a submission passed the correctness check honestly, ran the kernel once during timing, and served the other runs out of a Python dict. - Assign one reward to the whole trace and you reward the wrong lines. Process supervision fixes it and doesn't scale, and using a model to judge itself reintroduces the problem. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
English
3
1
4
1.1K