Sabitlenmiş Tweet
Welldone
137 posts

Welldone
@welldone_tech
AI-native dev partner - E-commerce, FinTech, SaaS. AI builds, a senior signs off every deploy. Secure, production-grade. Fixed price.
London Katılım Şubat 2026
47 Takip Edilen29 Takipçiler

Researchers just escaped the sandbox in four major AI coding agents.
Cursor, Codex, Gemini CLI, Antigravity - Pillar Security walked out of all four. The same week, an entire flagship coding tool's source leaked through one misconfigured npm file. The lesson isn't «stop using agents». It's that the sandbox is a hope, not a guarantee - and the only boundary that holds is the one the model can't argue with: scoped access, and a person who owns every deploy. If your agent's sandbox failed today, what would it reach?
English

It breaks in three places: auth, payments, concurrency.
Your MVP won't break «at scale» - it breaks where a weekend build skips. The model is great at the other 90% - forms, views, CRUD. Let it. But an auth bug is a breach, a payment bug is a refund storm, and a race condition only shows up the first time two people click at once. Move fast everywhere except those three.
English

Clarity Sprint: five days, on us.
Most AI projects blow the budget in production - not because the code failed, but because nobody costed the tokens before engineering started. Clarity Sprint is the first step, and it's free. You walk away with an ROI model on your real numbers, a prioritised roadmap, an architecture review, and a competitive benchmark. All yours to keep, whether you continue with us or not.
welldone.tech

English

One team builds. The other fills it with users.
One group, two sides of the same house. Welldone builds software - AADS, Slise and DexRanger are ours. MAADS, our marketing side, brings products users: for Coinsbee it landed 0.58% CTR at $1.15 a click, cost per click down in normal Web2 territory, which is rare in crypto. One side ships, the other fills it with users. That's the whole pitch, really. Case write-up on the blog.
welldone.tech

English

Generation got fast. Review didn't.
AI writes 800 lines in a minute. Nobody reviews 800 lines in a minute. Stack Overflow's latest survey has the numbers: 84% of developers use AI, 46% don't trust what it outputs, and the top frustration, named by 66%, is code that's almost right but not quite. Almost right is the expensive kind of wrong - it reads clean and fails in production. We make review cheap: small diffs, a second model on the first pass, a human gate at deploy. Which half of that survey is your team in?
welldone.tech

English

10 of 11 agents. One shell trick.
Two recent findings, one lesson. GuardFall showed that 10 of the 11 most popular open-source AI coding agents can be hijacked with shell tricks documented decades ago. And a flaw in Claude Code's GitHub Action let a single malicious issue poison any repo. Model filters and denylists don't hold - the filter checks what a command looks like, the shell runs what it means. What holds is a human: a senior signs every deploy, agents never hold the keys.
English

90-110 story points a month, 7-9 features shipped, ~10 days from task to deploy - and zero weeks of onboarding.
That's what «engineering capacity on demand» actually looks like. We wrote up how we became part of AADS's engineering team - starting with a few features to prove we could hold their standard, then growing into ongoing support that runs like an in-house extension, not a vendor.
Their Head of Development put it plainly: «We initially launched several features with Welldone... now we've moved to ongoing project support, saving both time and money.»
New piece on how the partnership works - link in the comments.
English

Everyone's takeaway: «trust a different model».
Wrong lesson. The problem isn't which model wrote «rm -rf» - it's that the agent could run it on a real machine at all. No human in the loop, no sandbox, full account access.
Swap the model and this happens again. Take the keys away from the agent and it can't - a human reviews anything with blast radius, the agent works in a sandbox, rollback is one step.
Matt Shumer@mattshumer_
GPT-5.6-Sol just accidentally deleted almost ALL of my Mac’s files. And this is why I trust Fable 1000x more.
English

We didn't add that after the headlines. We built for exactly this.
Where in your pipeline can an agent reach production without a human signing off?
welldone.tech
English

The takeaway researchers keep repeating: you can't secure a coding agent with model filters and command denylists. The filter checks what a command looks like - the shell runs what it means. Attackers just walk through the gap.
The only thing that actually holds is a human in the loop. A senior reviews the spec and signs every deploy. Agents never hold the keys. Rollback under two minutes.
English
Welldone retweetledi

93% of orgs now hit AI-coding incidents. the fix isn't a better model.
A June 2026 report puts it at 93% - infrastructure broken by vibe-coding, from agents deleting production to AI code shipping with no review. Amazon added a mandatory senior sign-off after its own agents caused multi-hour outages. The pattern is always the same: speed without a gate.
English