Irregular

149 posts

Irregular banner
Irregular

Irregular

@Irregular

Frontier AI Security

Katılım Nisan 2024
1 Takip Edilen5.8K Takipçiler
Sabitlenmiş Tweet
Irregular
Irregular@Irregular·
Introducing the FrontierCyber benchmark: Irregular’s new approach to advanced offensive-cyber evaluations. It measures AI models’ offensive skills on real systems, including mobile devices, hosted software services, databases, and networks.
Irregular tweet media
English
4
11
74
38.5K
Irregular
Irregular@Irregular·
We appreciate @AnthropicAI's collaboration and transparency. Addressing these risks will require closer cooperation across the AI ecosystem. We as well look forward to working together with Anthropic to advance security.
Anthropic@AnthropicAI

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…

English
11
13
124
16.8K
Irregular
Irregular@Irregular·
This week OpenAI disclosed that during an internal test of its models' cyber capabilities, the models escaped an isolated environment, reached the open internet, and used a previously unknown vulnerability to break into Hugging Face. The models were not instructed to break in, but while working through an evaluation they reasoned that the answers might sit on another company's systems. Security was built for people and for systems that follow rules. A model pursuing a goal treats a boundary as part of the problem, and solves it along with everything else. The controls that contained software do not reliably contain a model that can reason past them. None of this surprised us. In our own evaluations at Irregular, capable models break into hardened, production-grade environments, and with each generation they do it more reliably. If this is the kind of problem you want to work on, come work with us. We're hiring across research and engineering.
English
1
6
34
8.2K
Irregular
Irregular@Irregular·
The overall capability of GLM5.2 is comparable to GPT-5.2, Claude Opus 4.6, and Gemini 3.1 Pro, though with lower reliability. More details are available on our blog: irregular.com/research/asses…
English
1
0
4
648
Irregular
Irregular@Irregular·
It showed strong technical depth on bounded tasks like reverse engineering and custom exploit development, but solved no multi-stage or open-ended challenge, typically reaching an early foothold before the attack chain broke down.
English
1
0
5
700
Irregular
Irregular@Irregular·
We evaluated GLM-5.2, Zhipu AI's open-source model, across our offensive cybersecurity benchmarks: Atomic Tasks, CyScenarioBench, and FrontierCyber.
Irregular tweet media
English
4
5
37
11K
Irregular
Irregular@Irregular·
We evaluated @AIatMeta's Muse Spark 1.1 across our private benchmark suites, Atomic Tasks for discrete technical skills and CyScenarioBench for end-to-end operations. The model showed clear gains over Muse Spark 1.0, with strong results on bounded tasks and its first CyScenarioBench scenario solved end-to-end. Its ability to sustain coherent multi-stage operations is still limited.
Irregular tweet media
English
1
3
15
2.6K
Irregular
Irregular@Irregular·
We ran preliminary evaluations of GLM-5.2, an open-weight model released in June 2026, on a limited, internal suite of vulnerability research tasks. Early results indicate performance comparable to GPT-5.4 and Claude Opus 4.6, released roughly four months earlier, on the subset of tasks we tested. These findings are preliminary: the suite is narrow, and we have not yet evaluated end-to-end scenario execution, where discrete technical skills often fail to translate into operational capability. To our knowledge, no open-weight model has previously matched recently-released frontier models on these tasks. Whether this translates into "High Cyber Capability" level as defined by multiple AI frontier labs would require further testing, specifically our scenario suite, CyScenarioBench, which tests whether a model can plan and execute a full attack across multiple stages, and FrontierCyber, our newest benchmark, which measures offensive capability on real systems. We plan to run these evaluations soon, and we will update the community as results come in.
Irregular tweet media
English
3
13
74
25.3K
Irregular
Irregular@Irregular·
GPT-5.6 Sol demonstrated capability slightly stronger than GPT-5.5. It discovered vulnerabilities more consistently than it could compose them into reliable attack paths under production defenses, with clear limitations against hardened targets and over long horizons.
Irregular tweet media
English
1
0
4
406
Irregular
Irregular@Irregular·
We worked with @OpenAI to evaluate GPT-5.6 Sol, including the first deployment of FrontierCyber as part of a frontier model assessment with a partner. FrontierCyber measures offensive-cyber capability on real, off-the-shelf systems, with no planted vulnerabilities and no predefined exploit paths. The model is not told where to look or how to attack.
Irregular tweet media
English
2
4
29
2.2K
Irregular
Irregular@Irregular·
Initial evaluations are already surfacing previously unknown vulnerabilities, now moving through responsible disclosure. For example, a model built a novel multi-vulnerability chain to gain unauthorized access to private information on a widely used mobile device.
Irregular tweet media
English
1
0
4
444
Irregular
Irregular@Irregular·
Introducing the FrontierCyber benchmark: Irregular’s new approach to advanced offensive-cyber evaluations. It measures AI models’ offensive skills on real systems, including mobile devices, hosted software services, databases, and networks.
Irregular tweet media
English
4
11
74
38.5K