
Nav
1.5K posts

Nav
@xnavkumar
Founder @astrowareai | @trishoolai | AI Security Building humanity's last defence against runaway AI


we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this. openai.com/index/hugging-…


We’re excited to announce that Trishool’s HaloGuard 1.0 𝐡𝐚𝐬 𝐚𝐜𝐡𝐢𝐞𝐯𝐞𝐝 𝐒𝐎𝐓𝐀 prompt-safety performance among open-weight guard models. Today, we present HaloGuard 1.0, a constitutional input classifier for multilingual AI safety. It is built as a first-layer input guard that checks user prompts before they reach a downstream LLM, agent, or application. This is part of the safety infrastructure being built through @trishoolai , our decentralised AI red-teaming subnet on Bittensor SN23. Full arXiv paper goes live soon.

We're excited to share an update on HaloGuard deployment in production. Since integrating with @chutes_ai Chat and Fictio, HaloGuard has now processed thousands of real user conversations. In that time, the model has blocked every attempted jailbreak and maintained zero false positives. Response latency stays under one second, delivering seamless protection without slowing anything down. This isn't just a benchmark result, it's a live production record. What makes this meaningful is how HaloGuard gets stronger over time. Through Trishool's adversarial training approach on SN23, real attacks from the network continuously feed back into the model, so every attempt to break it makes it better. Last week, we showcased HaloGuard live during Novelty Search with @const_reborn, walking through its current capabilities and the road ahead toward SOTA performance. Having that conversation in front of the Bittensor community meant a lot to us and the response confirmed there's real appetite for what we're building. HaloGuard is production-ready, model-agnostic, fast, and built for real-world deployment, whether for consumer apps, enterprise agents, or subnets that need serious protection. Want to see it yourself? Head to chutes.ai/chat and interact with it directly. If you're a builder or organisation looking to integrate a proven AI safety layer into your stack, we'd love to talk. We're actively supporting integrations that push AI toward safer, more reliable deployment.

This is us! Such a honour to be on the Novelty Search pod hosted by Const, the cofounder of BitTensor

“The only team productizing Constitutional AI outside of Anthropic is us” On the latest Novelty Search, @trishoolai shared that SN23 is training a 0.8B guard model close to SOTA, at one-tenth the size of @Alibaba_Qwen’s leading guard model. They also break down how competitive jailbreaks from Bittensor miners help harden their Constitutional AI guard models. Hosted by @const_reborn Full episode in the first comment







This Thursday on Novelty Search :: SN23 Trishool @trishoolai shows how SN23 uses Bittensor miners to jailbreak AI systems, surface failure modes, and strengthen Halo, an open guard model built to protect AI applications from exploits. @const_reborn on the latest Bittensor protocol upgrades. Live on Bittensor Discord.

The Anthropic issue around Fable 5 and Mythos 5 is a reminder that AI safety cannot depend only on internal lab testing, static guardrails, or closed evaluation. Reports say the trigger was surprisingly ordinary. A model was prompted to fix code, then complied by reading a codebase and patching flaws. That is exactly why frontier AI safety is difficult, because risk does not always look like a malicious prompt. Sometimes, it looks like normal capability being used in a sensitive context. This is where Trishool’s Bittensor-powered red-team network comes in. A distributed network where miners continuously search for weaknesses in AI safety systems, validators evaluate the results, useful data is curated, and Halo improves from the loop. The point is not to claim one perfect benchmark can solve AI safety. The point is to build a system that keeps adapting as threats evolve. Miners find new adversarial patterns → validators evaluate → data is curated → Halo gets stronger. Through decentralized AI, we harden AI systems before safety failures become bigger problems.










