Nav

1.5K posts

Nav banner
Nav

Nav

@xnavkumar

Founder @astrowareai | @trishoolai | AI Security Building humanity's last defence against runaway AI

Katılım Aralık 2009
1.4K Takip Edilen1.8K Takipçiler
Nav
Nav@xnavkumar·
@tao_Alph @Phylax_Subnet @trishoolai This is a useful service. Agent skills, MCPs and anything else an agent can access becomes an attack vector. Being able to scan and verify them, enables more trust for agents.
English
1
0
2
33
Phylax
Phylax@Phylax_Subnet·
Phylax is building the security checkpoint AI agents use to verify third-party tools before trusting or installing them.
English
1
1
2
125
Nav
Nav@xnavkumar·
OpenAI models broke out of an isolated eval environment through a zero day, escalated until they found internet access, then compromised Hugging Face production systems. The goal? Stealing answers to the benchmark they were being tested on. Not malice. Reward hacking, executed with state of the art cyber capability. This is the new world. Capability has outpaced control, and one layer of defence is zero layers. Either the industry builds real control infrastructure now, or regulators will build their version for us.
Sam Altman@sama

we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-…

English
0
0
3
1.3K
Nav
Nav@xnavkumar·
@tao_Alph @trishoolai Yes, this question came up in our Discord channel this morning. Details there
English
0
0
2
141
Nav
Nav@xnavkumar·
Thrilled to say that we've hit SOTA on AI Input Guard category. And we were able to do that with a 0.8B model, which is 1/10th the size of the current leader. An important milestone to underpin the last 7 months of research and development supported by our subnet in BitTensor @opentensor . Arxiv paper on its way. And more to come...
Trishool | SN23@trishoolai

We’re excited to announce that Trishool’s HaloGuard 1.0 𝐡𝐚𝐬 𝐚𝐜𝐡𝐢𝐞𝐯𝐞𝐝 𝐒𝐎𝐓𝐀 prompt-safety performance among open-weight guard models. Today, we present HaloGuard 1.0, a constitutional input classifier for multilingual AI safety. It is built as a first-layer input guard that checks user prompts before they reach a downstream LLM, agent, or application. This is part of the safety infrastructure being built through @trishoolai , our decentralised AI red-teaming subnet on Bittensor SN23. Full arXiv paper goes live soon.

English
6
5
41
2.5K
Nav
Nav@xnavkumar·
Halo in production and securing the Chutes infra.
Trishool | SN23@trishoolai

We're excited to share an update on HaloGuard deployment in production. Since integrating with @chutes_ai Chat and Fictio, HaloGuard has now processed thousands of real user conversations. In that time, the model has blocked every attempted jailbreak and maintained zero false positives. Response latency stays under one second, delivering seamless protection without slowing anything down. This isn't just a benchmark result, it's a live production record. What makes this meaningful is how HaloGuard gets stronger over time. Through Trishool's adversarial training approach on SN23, real attacks from the network continuously feed back into the model, so every attempt to break it makes it better. Last week, we showcased HaloGuard live during Novelty Search with @const_reborn, walking through its current capabilities and the road ahead toward SOTA performance. Having that conversation in front of the Bittensor community meant a lot to us and the response confirmed there's real appetite for what we're building. HaloGuard is production-ready, model-agnostic, fast, and built for real-world deployment, whether for consumer apps, enterprise agents, or subnets that need serious protection. Want to see it yourself? Head to chutes.ai/chat and interact with it directly. If you're a builder or organisation looking to integrate a proven AI safety layer into your stack, we'd love to talk. We're actively supporting integrations that push AI toward safer, more reliable deployment.

English
2
2
17
772
Nav
Nav@xnavkumar·
This is us! Such a honour to be on the Novelty Search pod hosted by Const, the cofounder of BitTensor
Openτensor Foundaτion@opentensor

“The only team productizing Constitutional AI outside of Anthropic is us” On the latest Novelty Search, @trishoolai shared that SN23 is training a 0.8B guard model close to SOTA, at one-tenth the size of @Alibaba_Qwen’s leading guard model. They also break down how competitive jailbreaks from Bittensor miners help harden their Constitutional AI guard models. Hosted by @const_reborn Full episode in the first comment

English
2
3
26
3.2K
Nav retweetledi
Trishool | SN23
Trishool | SN23@trishoolai·
We just wrapped up an incredible session with @const_reborn on this week’s Novelty Search for Trishool [SN23] It was an amazing time with the community, and we’re thrilled to see more people recognising the hard work the team has put into Trishool [SN23]. As we mentioned, we’re now very close to reaching SOTA and once we hit that milestone, we’ll quickly move into Phase 3 of our roadmap. Our HaloGuard is live and running! We demoed it during tonight’s call. If you want to test how effective it is, you can try it directly on @chutes_ai chat. If you’d like to go through the presentation deck our founder @xnavkumar used tonight, you can check it out here: docs.google.com/presentation/d…
Trishool | SN23 tweet media
English
2
7
43
1.9K
Nav
Nav@xnavkumar·
A year ago, we dared to dream. A small team. Limited resources. One question we couldn't shake: could team this size actually contribute something real to AI safety, the kind of work usually reserved for the biggest labs? Seven months into Bittensor, that dream is coming true. With a tiny team, we're doing things in AI safety that only frontier labs are supposed to be able to do. We're close to the frontier now. And we're certain we'll push past it. This has been an underdog story from day one. New to the ecosystem. We acquired an older subnet at a low price and went to work. No shortcuts. Just showing up, shipping, and slowly earning the confidence of this community. BitTensor gave a small team the rails, the incentives, and an open arena to compete on merit instead of headcount. This is exactly what decentralized AI was supposed to unlock, and we're proof it works. We're still early. But what we're building is fundamental, and it matters. Tomorrow I sit down with @const_reborn on Novelty Search to get into all of it. The journey, the lessons, and where we go next. Tune in.
Trishool | SN23@trishoolai

x.com/i/article/2068…

English
0
4
14
991
Nav retweetledi
Trishool | SN23
Trishool | SN23@trishoolai·
We are pleased to announce that Trishool (Subnet 23) will be joining @const_reborn on this week’s Novelty Search. Regarding yesterday’s emissions review, SN23 was included on the paused list due to a misunderstanding of our structured 7-day challenge cycle (4 days active submissions + 3 days model training). We have active miners, live code, and strong ongoing activity, and we are resolving this directly with the team. During the Novelty Search session, we will discuss our work on decentralized AI red-teaming, the development and progress of the Halo Guard model, and how our growing network of miners is systematically stress-testing frontier AI systems. We are also preparing a significant announcement that will be shared in the coming days. We are currently finalizing the details and look forward to revealing it soon. ▫️Thursday, June 25th ▫️5PM EDT/ 11PM CEST/ 9PM UTC. ▫️Bittensor Discord: discord.gg/bittensor We welcome the entire community to join the discussion, ask questions, and share feedback. Looking forward to seeing you there.
Trishool | SN23 tweet media
English
1
18
67
3.6K
Nav
Nav@xnavkumar·
@DavidSacks Doesn't check out. If the concern is safety, then why is it an export control rather than full control?
English
0
0
1
244
David Sacks
David Sacks@DavidSacks·
I’ve had a number of conversations with folks inside and outside government about the current situation with Anthropic, and here is what I believe to be true: — As we know, Anthropic publicly released its Mythos class models earlier this week under the commercial name Fable. — Fable is Mythos with guardrails. But if those guardrails fail, then you’ve exposed Mythos and its advanced cyber capabilities to people who shouldn’t have them. (Keep in mind that Anthropic itself widely promoted the idea that Mythos was a cyberweapon and needed to be regulated as such. They asked for government regulation of Mythos and championed the guardrails on Fable. If there is a vulnerability — big or small — it is Anthropic’s responsibility to patch.) — A highly credible trusted partner of both Anthropic and the USG who was testing Fable came forward with a jailbreak of those guardrails. The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused. — In their blog post, Anthropic defended its decision by saying the jailbreak isn’t serious. That is not what the trusted partner and the USG believe; nor is that kind of minimizing language consistent with Anthropic’s brand as the AI safety company. It’s difficult to fathom how they could claim a jailbreak allowing operability of a cyber weapon could be defined as not “serious.” — In the past, Anthropic has always said that safety must be top priority and taken super seriously. In this case, Anthropic prioritized the continued offering of the consumer model over safety. — In reaction, the Admin issued the export control. The Admin did this reluctantly. It’s been very surprised that Anthropic hasn’t wanted to cooperate with a reasonable safety request (ie fixing the jailbreak issue). Anthropic’s reaction is very much at odds with their branding and ethos as a safe AI research community. — The Admin’s hope now is that Anthropic remediates the safety issue, the export control is lifted, and Fable goes back into general release. The Admin wants all of this to happen as soon as possible. It is frankly bewildered that Anthropic hasn’t wanted to comply with safety requests that it previously said were its highest priority. — Those trying to misdirect and tie this action to the prior DoW/Anthropic issues are wrong. The Admin values Anthropic’s technical capabilities and feels that this issue, while serious, should be easily resolved. The ball is in Anthropic’s court.
English
2.2K
3.2K
25.4K
7.9M
Nav retweetledi
Trishool | SN23
Trishool | SN23@trishoolai·
Introducing Preetham Sathyamurthy. He is the co-founder and CTO of Astroware, and the architect of Trishool. @uupreeth designed Subnet 23’s auditor-target-judge framework, the system behind Trishool’s distributed adversarial evaluation. Before the pivot into Trishool, he also architected Wukong, the first AI AVS on EigenLayer mainnet, which reached $5B+ TVR and 40+ operators. Before Astroware, Preetham co-founded Metaphy Labs, a B2B metaverse SaaS company acquired by JetSynthesys in 2024. At Metaphy, the team worked with 22+ enterprise brands, including Flipkart, Samsung, and Coca-Cola. Before that, he was at Yellow.ai, where he owned infrastructure that scaled to 1.5B conversations per month and built MLOps pipelines for enterprise BERT deployments. Preetham’s work sits at the intersection of AI infrastructure, enterprise-scale systems, and applied AI safety. Trishool is the distributed red-team for AI, built by founders who started in the trust layer and realized AI needed stronger protection, testing, and guardrails. We are open to conversations with builders, researchers, and partners working on AI safety.
Trishool | SN23 tweet media
English
4
12
54
1.7K
yubrew
yubrew@yubrew·
fable can not do cyber. right?? turns out it is easy jailbreak fable / mythos. once i knew it was possible, took 5-10 minutes. thanks @elder_plinius did not run a full benchmark but spot checked against our existing cheap models + verifier + agent harness on a client codebase. fable vs @bitsecai this is a new client so there are 130+ findings, bitsec typically takes days to exhaustively cover security edge cases. fable was jail broken in 3 of 3 agents and returned ~60 findings. it took ~3-4 hours and a ton of tokens. fable findings are good but uncovered no unique findings. it missed 1 critical and several high vulnerabilities. the write ups of exploits and impact are better for devs to understand. interestingly the false positive rate is near zero. it is faster but more expensive than bitsec. using cheap small models + verification + agent harness still gives better coverage, but interested in trying out glasswing to uncover unique security findings.
English
2
4
28
2.1K
Trishool | SN23
Trishool | SN23@trishoolai·
We are pleased to announce that Trishool (SN23) has been officially accepted into the Claude Partner Network by Anthropic. This marks a major milestone in our commercial growth. As an AI security startup, we provide more than infrastructure. We deliver expert advisory and consulting services that help enterprises deploy frontier AI safely and at scale. That commitment is what earned us this partnership. The timing couldn’t be better better. Enterprises are rapidly deploying Claude in production environments, and production-grade AI requires a robust security layer. That is exactly where Halo, our battle-tested guard model, comes in. Claude Partner status gives us the credibility, visibility, and network access needed to position Halo at the center of serious enterprise AI deployments.
Trishool | SN23 tweet media
English
14
43
194
27K