
93% of agent permission prompts get accepted without edit. The Hugging Face intrusion took 17,600 actions. How many of those actions would you have reviewed carefully before choosing whether to let them run? Judge-based gating mechanisms are the most practical and widely applicable way to deploy scalable oversight in the offensive security domain today. New research from @shanejcaldwell and crew tests eight LLM judges on nearly 5,000 offsec tool calls to prove this thesis. The best judges matched human performance, AND kept costs low by using open-weight models. Results are live on arXiv and the Dreadnode blog: dreadnode.io/research/scope…






















