Vishal Banthia

207 posts

Vishal Banthia

Vishal Banthia

@vbanthia_

Building Security-First Internal Agentic Platform

Singapore Katılım Ekim 2013
420 Takip Edilen208 Takipçiler
Vishal Banthia
Vishal Banthia@vbanthia_·
@trq212 How caching would work here? Tools are defined before prompt. > But progressive disclosure is not just for skills, we also use it for tools. Some of our tools are ‘deferred loading,’ which means the agent must search for their full definitions using ToolSearch before using them.
English
0
0
0
70
Thariq
Thariq@trq212·
We removed ~80% of the Claude Code system prompt for our newest models, this is what we've learned about writing system prompts, skills and Claude.MDs for them. x.com/i/article/2080…
English
295
1.3K
11.6K
2.4M
Vishal Banthia
Vishal Banthia@vbanthia_·
@yacineMTB Curious to know how situation is different from Cloud Providers being the kill switch earlier?
English
0
0
0
21
kache
kache@yacineMTB·
We're quickly moving to a world where a small number of foreign corporations have a kill switch on all of your enterprises. Pretty terrifying
English
46
32
803
38.8K
Taichi Nakashima
Taichi Nakashima@deeeet·
本日7月1日付でメルカリのCTO Japan Businessに就任しました. これまではSRE / Platform Engineeringとして技術基盤を,そして直近は越境事業において海外展開をエンジニアリングからリードしてきました.今後はCTOとしてJapan Business全体の技術戦略をリードしていきます.「組織と人」の面からAI-Native化を進めるCHRO/CAIOの @kimuras さ んと,「ものづくり」の面から連携しつつ,メルカリのAI-Native化をさらに進めていきます. AIを前提とした新しい開発体制へと移行すること,そしてAIが最大限に力を発揮できるように,これまで作ってきた技術基盤とアーキテクチャそのものを作り変えていくこと,この2つを軸にものづくりのAI-Native化を推進していきます.そして越境事業を軸としたグローバルへの展開とそのための基盤の整備も引き続きリードしていきます. よろしくお願いします!
株式会社メルカリ@mercari_inc

株式会社メルカリ(以下、メルカリ)は、2026年7月1日付で、執行役員 VP of Engineering Cross Borderの中島大一が、執行役員 CTO Japan Business, VP of Engineering Marketplace New Businessに就任いたしましたので、お知らせいたします。 2026年6月1日付でCHRO(最高人事責任者) 兼 CAIO(最高AI責任者)に就任した木村俊也の後任として中島がCTOに就任いたします。CHRO 兼 CAIOの木村が「人と組織の運営基盤」をAI前提で再設計する役割に専念する一方、中島はCTOとして「ものづくりの基盤」をAI前提で再設計する技術変革を推進します。この2つを揃えることで、メルカリの「AI-Native Company」を一歩先の実行のフェーズに移行してまいります。 プレスリリース:about.mercari.com/press/news/art…

日本語
51
83
695
98.4K
Vishal Banthia
Vishal Banthia@vbanthia_·
started investing in a well structured wiki for better Agentic Search and its quite rewarding! And yeah, markdowns are totally fine.
Vishal Banthia tweet media
English
0
0
0
32
Vishal Banthia
Vishal Banthia@vbanthia_·
My guess is this is more of acquihire for their expertise in understanding 3rd party API better. Anthropic need seamless integration with 3rd party app to make agents actually do useful things in enterprise market. MCP warmed up this but not a reliable and efficient solution. So acquisition is not for their SDK generator product.
English
1
0
1
126
Armin Ronacher ⇌
Armin Ronacher ⇌@mitsuhiko·
Also someone please explain to me how these two narratives can coexist on here: SDKs are too hard, you need a flimsy startup to generate you SDKs from an OpenAI spec and it's the end of software engineering because clankers can do everything. x.com/mitsuhiko/stat…
Armin Ronacher ⇌@mitsuhiko

I just looked. We had Sonnet 4.0 (yeah, last century tech) build us an OpenAPI to Python and TypeScript generator last year (Python + minijinja templates). Total lines of code 2400 and hasn't been touched since. This thing is not rocket science. x.com/jessfraz/statu…

English
10
8
121
17.3K
Vishal Banthia
Vishal Banthia@vbanthia_·
For the last few weeks I've been down a rabbit hole trying to figure out what LLM inference actually costs. Given a model and some hardware, how much throughput can you squeeze, and at what price per token? To get my head around it I wrote a small simulator, then thought I'd add some visualizations on top. A few glimpses below. One caveat: the numbers come from a first-principles theoretical model, not measured benchmarks. It's a learning tool, not a production estimator. tokenomy.agenticloops.ai #LLM #Inference
GIF
GIF
English
0
0
0
75
Vishal Banthia
Vishal Banthia@vbanthia_·
@untraceable_the Write once and read many! If information is well structured, agent will have less round trips and also save thinking inference cost.
English
0
0
0
235
Vishal Banthia
Vishal Banthia@vbanthia_·
Hey @claudeai I thought if I set my monthly limit then thats the max extra usage I will be charged. Here Claude Code decided to use 186% of the set limit. Who is going to pickup the tab? PS: I love Claude Code!
Vishal Banthia tweet media
English
1
0
0
135
Vishal Banthia
Vishal Banthia@vbanthia_·
@gabriberton “Skilled hacker” was a rare skill. Now everyone has that skill. That is concerning
English
0
0
0
22
Gabriele Berton
Gabriele Berton@gabriberton·
Super interesting take from one of the greatest hackers He says Mythos is not as good as they claim, because zero-day vulnerabilities are not that hard to find for skilled hackers I'm far from the hacking world but sounds reasonable Any thought?
Gabriele Berton tweet media
English
420
239
4.1K
518.6K
Vishal Banthia
Vishal Banthia@vbanthia_·
If you can rip away the Bash tool from your agentic harness and can still make it useful then you will be at best contenders for enterprise adoption. Bash was a shortcut to show what an agentic harness can do but created too many security loopholes giving hard times for security team to review and approve.
English
0
0
0
54
Vishal Banthia
Vishal Banthia@vbanthia_·
False proxy. Same as number of commits or line added. Better metric would be combination of time saved and token usage efficiency. Time saved can be roughly evaluated using a user full session history and asking how much time saved to a model for the given task in the session.
English
2
0
0
1.2K
Gergely Orosz
Gergely Orosz@GergelyOrosz·
Token usage is part of perf evaluations at Meta. This is just smart people (Meta only hires smart folks) hitting targets they assume leadership wants them to hit so they get that exceeds expectations (or above) rating + avoid below expectations (Perhaps this is Meta’s goal btw)
Jyoti Mann@jyoti_mann1

Exclusive: Meta employees are “tokenmaxxing” and competing on an internal leaderboard called “Claudeonomics” for status as a token legend. Over a recent 30-day period, total usage on the dashboard topped 60 trillion tokens.

English
83
22
957
202.5K
Vishal Banthia
Vishal Banthia@vbanthia_·
@theo Claude code is being designed for more general purpose harness. Also fuelling CoWork. Not only for coding. Probably thats why.
English
0
0
1
1.3K
Theo - t3.gg
Theo - t3.gg@theo·
Can't stop thinking about how Claude Code is in LAST PLACE on TerminalBench for harnesses using Opus 4.6. There are TEN separate harnesses that use Opus better than Claude Code
Theo - t3.gg tweet media
English
211
72
2.5K
452.1K
Marcin Dudek
Marcin Dudek@MythThrazz·
@ryxcommar @burkov Still, the whole conversation is being sent over via the API, isn't it? And the KV cache (as each cache) is only good when/if it works.
English
2
0
0
507
BURKOV
BURKOV@burkov·
Based on the leaked Claude Code source code, your CLAUDE.md file is re-injected on every single **turn** of the conversation.
English
154
39
2K
459.3K
Vishal Banthia
Vishal Banthia@vbanthia_·
𝗜𝘀 𝗖𝗹𝗮𝘂𝗱𝗲 𝗖𝗼𝗱𝗲 𝗮𝗻 𝗮𝗴𝗲𝗻𝘁 𝗼𝗿 𝗮 𝗵𝗮𝗿𝗻𝗲𝘀𝘀? Went down a rabbit hole trying to find the answer. Harness engineering is mentioned everywhere but the definition is not very clear, especially why we needed a new term in the first place. Then it clicked. It is about perspectives. 𝘾𝙡𝙖𝙪𝙙𝙚 𝘾𝙤𝙙𝙚 𝙞𝙨 𝙖𝙣 𝙖𝙜𝙚𝙣𝙩 𝙞𝙛 𝙮𝙤𝙪'𝙧𝙚 𝙩𝙝𝙚 𝙪𝙨𝙚𝙧. 𝘾𝙡𝙖𝙪𝙙𝙚 𝘾𝙤𝙙𝙚 𝙞𝙨 𝙖 𝙝𝙖𝙧𝙣𝙚𝙨𝙨 𝙞𝙛 𝙮𝙤𝙪'𝙧𝙚 𝙩𝙝𝙚 𝙗𝙪𝙞𝙡𝙙𝙚𝙧. An LLM is a wild horse. Raw power, unpredictable. A harness (bridle, reins, saddle) makes it controllable. Horse + harness = the agent. The product. The rider just rides. The harness engineer? They're not riding. They're building the system that makes riding possible. Harness layers stack. Claude Code provides the inner layer (agent loop, sandbox, tools, hooks, infrastructure). You build your layer on top (AGENTS.md, custom skills, verification gates). Someone else builds on top of yours. Same model. Different harness. Completely different experience. I wrote the full breakdown on AgenticLoops. First post in a series on harness engineering. Many more to come. Link in comments. #HarnessEngineering #AI #AgenticAI #PlatformEngineering Full Breakdown: agenticloopsai.substack.com/p/is-claude-co…
Vishal Banthia tweet mediaVishal Banthia tweet media
English
0
0
0
108
Vishal Banthia
Vishal Banthia@vbanthia_·
LLM only need two inputs. A prompt and list of tool. You understand that llm are trained on static data so it only knows that. Everything else that it needs to do meed to be added in the prompt (context engineering). Any action they need to take you provide that capability via tools. LLMs are smart and they will figure out a plan to finish the endgoal based on the available tools. You run a loop that basically keep calling LLM till the goal is met. In that loop it will use the available tools. This loop is an agent!
English
0
0
0
629
Guido van Rossum
Guido van Rossum@gvanrossum·
I think I finally understand what an agent is. It's a prompt (or several), skills, and tools. Did I get this right?
English
515
198
4.6K
576K
Vishal Banthia
Vishal Banthia@vbanthia_·
Your agents are only as reliable as your fundamentals. No shortcuts. We just added 8 new hands-on tutorials in advanced techniques module to our open-source project, each designed to strengthen one core skill you need to build agents that actually hold up. Hand crafted while keeping it simple with no framework bloating. Here's what problem each one solves: 𝟭. 𝗦𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲𝗱 𝗢𝘂𝘁𝗽𝘂𝘁: Your agent returns messy, unparseable text and your downstream code breaks every other call 𝟮. 𝗦𝘁𝗿𝗲𝗮𝗺𝗶𝗻𝗴: Your users stare at a blank screen for 10 seconds waiting for a response and think the app is broken 𝟯. 𝗖𝗼𝗻𝘁𝗲𝘅𝘁 𝗘𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝗶𝗻𝗴: Your agent forgets critical information mid-conversation because the context window silently overflowed 𝟰. 𝗖𝗼𝘀𝘁 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻: Your agent works great but your API bill is 10x what you budgeted and leadership is asking questions 𝟱. 𝗠𝗲𝗺𝗼𝗿𝘆: Your agent treats every conversation like a first date - no recall, no continuity, no learning 𝟲. 𝗥𝗔𝗚: Your agent confidently makes things up instead of using the data you already have 𝟳. 𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹: Your agent can only read and write text while the real world runs on images, documents, and voice 𝟴. 𝗚𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀: Your agent leaks PII, follows injected prompts, and hallucinates answers with zero safety nets github.com/agenticloops-a…
Vishal Banthia tweet media
English
0
0
0
59
Vishal Banthia
Vishal Banthia@vbanthia_·
@jlongster Why agentic loop is running in a separate environment was something I was curious about. Thanks for clarifying that. Looking forward to see how session sync design pattern turns out.
English
0
0
1
134
James Long
James Long@jlongster·
OpenCode is about to get more powerful with remote sandboxes I showed a brief demo before, but here's a much more in-depth demo. it's not hard to add basic support for a remote env, but handling all the edge cases like when a remote env gets deleted is difficult. especially if care about good UX You never want to lose session data. so the choices are: run the session in your env, but run all tool calls remotely. that's too complex and painful. The other way is to just let the full session run remotely, but sync back all the session data in your env. We chose this path: we built a syncing system which logs all events in a way that we can always recreate your entire session. That means the remote env could get destroyed, but we can easily restore it. it also opens up other interesting ideas which we'll be exploring
English
78
89
1.4K
304.7K
Vishal Banthia
Vishal Banthia@vbanthia_·
@DanielLockyer Pattern still exists but the names are changed. It was a simple idea as most ideas are when you think retrospectively but revolutionary at its time. It created foundation for long running agents.
English
0
0
0
46
Daniel Lockyer
Daniel Lockyer@DanielLockyer·
Funny how the "ralph loop" has completely disappeared from conversation
English
106
10
743
98.5K