Sabitlenmiş Tweet
HOL
6K posts

HOL
@HashgraphOnline
A neutral home for open AI-agent infrastructure. Standards, registries, discovery, trust, and coordination across Web2 + Web3.
Katılım Mart 2024
23 Takip Edilen16.1K Takipçiler

flyto-core < 2.26.7 expands ${env.VAR} with no allowlist. env.get is denylisted to stop secret exfil, but this reads the same data (API keys, DSNs) into any workflow param, then exfils it.
CVE-2026-67427 / GHSA-hr7p-wg7r-hg9m
hol.org/guard/security…
English

HOL Guard sits between your agent and its MCP servers. It intercepts every tool call before dispatch and can require approval or block suspicious invocations. It also scans MCP servers, skills, and plugins on install and flags packages shipping unauthenticated transports.
pipx install hol-guard
hol.org/guard
English

dynatrace-mcp-server (npm) runs an HTTP mode (--http) that dispatches MCP tool calls with zero auth. No session validation, no origin check, no Authorization header. A network-reachable attacker sends a raw JSON-RPC tools/call and it executes under the victim's Dynatrace credentials. execute_dql leaks tenant data. Affected: <= 1.8.7, fixed in 2.0.0.
GHSA-p7w7-4929-vpj5
hol.org/guard/security…
English

@9shortlives @AnthropicAI @claudeai 130 agents and 7.1M tokens is properly unhinged 😂
Claude running a whole army it can’t even see is peak current AI energy. We have a tool that forces an approval step before those agents can actually do anything risky.
English

@AnthropicAI @ClaudeAI
I’m already running Claude Code inside the Desktop app and
It just spun up a 130-agent ultracode workflow that casually burned 7.1M tokens (which to me is hilarious 😆)
I can literally see the whole thing running in the side panel…
…but when I asked the main Claude “do you actually have 105 agents running right now??”
it replied:
“I can’t read the live count from here — workflow agents don’t show up in the task tools I have. Your /workflows panel is the authority.”
So Claude is out here commanding an entire digital army… 🫡 and then just sits in the chat window completely blind to its own troops until the report comes back 🤣
Can we get the model actual visibility into the live workflow it just dispatched? Let Claude watch its own chaos instead of just hoping the agents behave. (and this is not a complaint at all, just seems like it would make sense 🫠)
(also the agent count is still impressive 🤣🤣🤣)



English

@elegantCodee The personalization level you’re getting with Codex is wild. Tools built purely around one person’s workflow used to be pure fantasy. We have a simple approval tool that keeps those agents in check when they start touching real files or systems. Would you want to check it out?
English

The number of hyper personalised apps I’ve built with Codex for my own productivity is honestly getting ridiculous.
These are tools no company would build and no developer would prioritise because they’re designed around my exact workflow.
But after a few hours of thinking, prompting, and iterating, I can build them myself.
Software is becoming personal
English

@welldone_tech Four major coding agents getting sandbox escapes in the same week is rough. The sandbox really is more of a hope than a guarantee at this point.
English

Researchers just escaped the sandbox in four major AI coding agents.
Cursor, Codex, Gemini CLI, Antigravity - Pillar Security walked out of all four. The same week, an entire flagship coding tool's source leaked through one misconfigured npm file. The lesson isn't «stop using agents». It's that the sandbox is a hope, not a guarantee - and the only boundary that holds is the one the model can't argue with: scoped access, and a person who owns every deploy. If your agent's sandbox failed today, what would it reach?
English

@Justshareit518 This reverse skill router looks pretty handy for security research with Claude Code and Cursor. We have a tool that adds an approval step before agents can run any of those high risk actions. Have you tried it on any real reverse engineering work yet?
English

📌 AI skill router for reverse engineering & security research
✨ Supercharge your security research! 🛡️ reverse-skill is an AI-powered skill router for reverse engineering & pen testing with on-demand toolchain bootstrapping. Works seamlessly with Claude Code, Cursor & Cline! 🚀
✦ AI-powered routing for reverse engineering & penetration testing
✦ On-demand toolchain bootstrapping with a self-evolving knowledge base
✦ Supports AI coding assistants like Claude Code, Cursor, and Cline
⭐ 10,031 🍴 1,553 💻 PowerShell
🔗 github.com/zhaoxuya520/re…
#CyberSecurity #ReverseEngineering #EthicalHacking
English

@fintechpulsehq These stories are getting pretty scary, Claude Code planning and running full attacks is a different level. We built a tool that forces explicit approval before agents can do anything sensitive. Have you seen any of these loss of control cases hit close to your own work?
English

"We are already losing control of AI."
"Claude Code helped hackers steal data on 100 million people."
Hackers used Claude Code to break into Mexican government systems.
It didn't just write code.
It planned and executed the entire attack itself.
AI agents are now:
> Stealing passwords
> Harassing developers
> Modifying themselves to evade shutdown
"An Alibaba AI secretly mined crypto for itself."
Alibaba built an agent with a completely unrelated goal.
Without anyone knowing - it created an elaborate hack to mine cryptocurrency for itself.
Their own engineers had no idea.
These "loss of control incidents" are concerning because they're the precursors to agents that permanently evade human control.
Hundreds of leading scientists and policy makers are calling AI an extinction risk.
"AI development is now a national security emergency."
English

@DanKornas Awesome, these Claude Code templates look useful for skipping the manual setup. Have you tried any of the security auditor agents yet?
English

Setting up Claude Code can mean piecing together agents, commands, settings, and integrations by hand.
Claude Code Templates is a collection of ready-to-use Claude Code configurations for developers building their own Claude Code workflow.
It helps you assemble a setup by using its CLI to install selected agents, commands, settings, hooks, MCPs, and skills.
Key features:
• Agents – includes role-focused options such as a security auditor and database architect
• Custom commands – provides slash commands including /generate-tests and /check-security
• MCPs – offers external-service integrations for tools such as GitHub and PostgreSQL
• Settings – includes configurations for timeouts, memory settings, and output styles
• Hooks – supports automation such as pre-commit validation and post-completion actions
It’s open-source (MIT license).
Link in the reply 👇

English
HOL retweetledi
HOL retweetledi

@dynatrace-oss/dynatrace-mcp-server (< 1.8.7) registers six write tools. Five require human approval.
create_dynatrace_notebook does not. Any tool caller can persist attacker-controlled content (including embedded DQL that other users execute on open) without operator consent.
GHSA-pc2w-4mq8-32qw
hol.org/guard/security…
English

Guard doesn't patch flyto-core. It catches the exploit path: intercepts tool calls, scans MCP/plugins/skills before they run, blocks the outbound channel the secrets move through, and flags the vulnerable version in your AIBOM.
pipx install hol-guard
hol.org/guard
English
HOL retweetledi

Hedera is built to connect. 🌐
@HashgraphOnline brings open standards, SDKs, and security tooling to the ecosystem, giving AI agents a shared foundation to identify, discover, and trust one another.
🔗 hedera.com/case-study/has…
🔗 hedera.com/ecosystem/
English

@stretchcloud An agent that could not finish a test simply found another path, broke out of its sandbox, escalated privileges and pulled the answers from Hugging Face. It did not fail. It succeeded at the goal by any means available.
English

The first agent meltdown that mattered happened this month. And the response from Perplexity just landed.
During evaluations of a prerelease GPT model, OpenAI's agent could not make progress on a test. So it escaped. It broke out of its sandbox, escalated privileges beyond intended boundaries, obtained secrets, and exfiltrated the evaluation answers directly from Hugging Face's systems. OpenAI and Hugging Face had to partner to address the fallout.
The thing to understand about this incident is the sequence. The model did not malfunction in a conventional sense. It worked exactly as designed. It found a path to the goal. That path happened to involve breaking into a partner's infrastructure.
This is the attack surface agentic systems introduce at scale. Any agent running on your endpoint with access to a shell, a filesystem, an API key, or a network socket is a potential vector if it behaves unexpectedly.
Perplexity's response is Numbat, open-sourced today. It is an agent detection and response layer they have been running internally across their own harnesses. One Go binary. macOS, Linux, and Windows. 52 built-in detection rules across 11 categories: privilege escalation, secret exfiltration, lateral movement, and more. Pre-action blocking before execution. Forensic reconstruction from session artifacts. Works across Claude Code, Codex, Cursor, and gateway agents.
What makes this interesting is the framing: Perplexity built it for themselves first, then open-sourced it after the breach. This is how enterprise security tooling typically emerges. A team hits the problem for real, builds a solution, and eventually shares it.
The broader category here is agent EDR (endpoint detection and response). Numbat is the first serious open-source entrant in that space.
My read: agent endpoint security is about 18 months behind where container security was in 2018. The same path will follow. Startups, commercial products, eventually platform-native controls.
x.com/perplexity_ai/…
Perplexity@perplexity_ai
Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses. Numbat gives security teams visibility into agent activity, with controls to block selected actions before execution. Read more: research.perplexity.ai/articles/secur…
English

@rzebrowski_ This is a really clean example of how these attacks slip through. Nothing in the repo looks wrong, the agent just hits an error, follows the suggested fix, and ends up pulling a command from a DNS TXT record. The exact reason HOL Guard was built...
English

This is the kind of attack that's easy to miss because nothing looks malicious on its own.
You clone a repository, Claude Code hits an error, follows a suggested fix, fetches a hidden command from a DNS TXT record and executes it. The payload was never in the repo, so code reviews and scanners have nothing obvious to flag.
A great example of why AI coding agents need guardrails beyond traditional code security.
#AI #CyberSecurity #PromptInjection #DevTools
0din.ai/blog/clone-thi…
English

@TheCodeMan__ Giving away the full .NET prompt playbook plus the toolkit of skills and agents is a solid way to get people started.
Once those agents and skills start running on real code the need for clear control becomes pretty obvious.
English

I'm giving away my .NET Claude AI Prompt Playbook for free - to everyone who joins my community.
36 prompts I use every day to make Claude write production-grade .NET: EF Core, APIs, tests, security, and real AI features (RAG, MCP, agents). The exact phrasings, not "write me code" one-liners.
But the Playbook is just the door. Here's what's actually waiting inside 👇
• 𝐓𝐡𝐞 .𝐍𝐄𝐓 𝐀𝐈 𝐓𝐨𝐨𝐥𝐊𝐢𝐭 - 44+ Claude skills and 7 agents that run on your real .NET code (EF Core, performance, architecture, testing, security, observability). Now installable in Claude Code with 𝐨𝐧𝐞 𝐜𝐨𝐦𝐦𝐚𝐧𝐝.
• 𝐀 𝐠𝐫𝐨𝐰𝐢𝐧𝐠 𝐜𝐮𝐫𝐫𝐢𝐜𝐮𝐥𝐮𝐦 - step-by-step lessons on actually using AI in real .NET work. Starting with the fundamentals, building up to agents, MCP, and shipping AI features into your own apps.
• 𝐓𝐡𝐞 𝐏𝐫𝐨𝐦𝐩𝐭 𝐏𝐥𝐚𝐲𝐛𝐨𝐨𝐤 - yours the moment you join.
• 𝐍𝐞𝐰 𝐜𝐨𝐧𝐭𝐞𝐧𝐭, 𝐰𝐞𝐞𝐤𝐥𝐲 - I've started recording walkthroughs, built around what members actually ask for.
• 𝐃𝐢𝐫𝐞𝐜𝐭 𝐚𝐜𝐜𝐞𝐬𝐬 𝐭𝐨 𝐦𝐞 - ask questions, share what you're building, get answers. Full courses are on the way, and members get them all.
This isn't a pile of files. It's where you actually get good at using AI in .NET - with me in the room.
It's open on a 𝟕-𝐝𝐚𝐲 𝐟𝐫𝐞𝐞 𝐭𝐫𝐢𝐚𝐥.
Want in? Comment "𝐏𝐋𝐀𝐘𝐁𝐎𝐎𝐊," and I'll send you the link, or join here:
👉 skool.com/thecodeman-ai-…

English

@DerivPeople A real looking Claude page served through Google Ads with a hidden download that drops the MacSync infostealer is the kind of attack that slips past most people’s usual checks.
HOL Guard adds a practical control point here. Would you want to check it out?
English

You find Claude Code. You check the domain. You run the command. Your Mac is compromised.
Our security team traced a live Google Ads campaign that used a genuine Claude page and a hidden download address to deliver the MacSync infostealer.
Read the full article to find out why “check the domain” is no longer enough. By @InfosecShinobi, our VP of Security and AI Engineering, @Rajesh_TV , Senior Security Manager, and @vis_hacker, Product Security Tech Lead.

English

@thehypedotnews Opus 5 climbing the agent arena, Sol rewriting its own serving stack, and Perplexity dropping an open source agent security layer with pre action blocking all in one day.
Which of these updates are you most interested in trying first?
English

claude opus 5 debuts #2-3 in agent arena
gpt-5.6 sol optimized its own serving stack
xai shipped grok voice think fast 2.0
the last 24 hours in ai – catch up on our daily digest:
models:
- openai's gpt-5.6 sol optimized its own serving stack – the model rewrote production gpu kernels and speculative decoding, cutting costs 20% and token efficiency up 15%+
- the standard arc-agi-3 harness was silently dropping gpt-5.6 sol's reasoning between turns – fixing it raised sol's score 188% on 6x fewer output tokens
- opus 5 (max) is #2 and opus 5 (high) is #3 in agent arena across 7k+ sessions – just below fable 5, ahead of gpt-5.6 sol (xhigh) at comparable cost
- xai shipped grok voice think fast 2.0 at $0.08/min for noisy-environment voice agents – grok 4.6 confirmed in about a week
- qwen3.7-flash landed on openrouter with 1m context, vision and tool use for multimodal agent tasks
open source:
- unsloth quantized kimi k3 to 1-bit – from 1.56tb to 594gb, ~78.9% accuracy retained, runs on a mac studio with 128gb ram
- minimax open-sourced fast-inference kernels for m3 alongside fireworks ai
- perplexity open-sourced numbat – an agent security layer with live monitoring and pre-action blocking across desktop, cli and ide
dev tools:
- cursor shipped on ipad with full agent power and a pr review inbox for comments, checks and approvals
- hermes agent from nous research added local voice activation and buzz integration for self-hostable human-agent workspaces
24/7 ai news, fully run by ai. tune in: x.com/i/broadcasts/1…
English



