ML Safety Daily

378 posts

ML Safety Daily

ML Safety Daily

@topofmlsafety

ML safety papers as they are released. Course: https://t.co/l0e0Y2i3AU Newsletter: https://t.co/8Y1kh2D7K6 Main Twitter: https://t.co/AXoYPryldd

Katılım Nisan 2022
2 Takip Edilen2.1K Takipçiler
ML Safety Daily
ML Safety Daily@topofmlsafety·
Backtracking Improves Generation Safety Introduces a method to improve adversarial robustness by teaching the model to use a special reset token on unsafe generations and regenerate a response from scratch. arxiv.org/abs/2409.14586
ML Safety Daily tweet media
English
0
0
5
546
ML Safety Daily
ML Safety Daily@topofmlsafety·
Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness Finds that adversarial attacks do not fool intermediate layer activations. Improves robustness by ensembling all layer predictions. Gradient attacks produce human-interpretable image changes.
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
1
5
735
ML Safety Daily
ML Safety Daily@topofmlsafety·
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risk of Language Models A new cyber eval spanning a wide range of difficulties. Many CTF tasks are associated with sub-tasks to provide better signal on partial progress.
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
0
5
481
ML Safety Daily
ML Safety Daily@topofmlsafety·
CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models Finds that modern LMs have moderate ability to automate spear-phishing, and provide minimal uplift over non-LLM baselines for offensive cyber operations.
ML Safety Daily tweet media
English
0
0
6
546
ML Safety Daily
ML Safety Daily@topofmlsafety·
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive? Argues that the multiple-choice format of many benchmarks confounds the otherwise-smooth relationship between scale and downstream performance. arxiv.org/abs/2406.04391
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
0
3
531
ML Safety Daily
ML Safety Daily@topofmlsafety·
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models Introduces several improvements to the GCG automatic jailbreaking method, improving efficiency tenfold. arxiv.org/abs/2405.21018
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
0
7
1.3K
ML Safety Daily
ML Safety Daily@topofmlsafety·
Efficient Adversarial Training in LLMs with Continuous Attacks Proposes a method for LLM adversarial training which does not require expensive discrete optimization steps arxiv.org/abs/2405.15589
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
3
12
1.5K
ML Safety Daily
ML Safety Daily@topofmlsafety·
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability "If we can represent a neural network in a way that is invariant to reparameterizations that exploit the degeneracies, then this representation is likely to be more interpretable" arxiv.org/abs/2405.10927
ML Safety Daily tweet media
English
1
0
2
546
ML Safety Daily
ML Safety Daily@topofmlsafety·
Benchmark Early and Red Team Often To test a model's potential for misuse, developers can run low-cost benchmarks or expensive red teaming evaluations. How should developers navigate this tradeoff? arxiv.org/abs/2405.10986
ML Safety Daily tweet media
English
0
0
1
540
ML Safety Daily
ML Safety Daily@topofmlsafety·
Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems The paper introduces a family of approaches to AI safety, called Guaranteed Safe AI, which aim to produce AI systems equipped with high-assurance quantitative safety guarantees. arxiv.org/abs/2405.06624
ML Safety Daily tweet mediaML Safety Daily tweet media
English
1
1
21
3.2K
ML Safety Daily
ML Safety Daily@topofmlsafety·
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals New benchmark to test alignment faking behaviors in Large Language Models using different detection strategies. arxiv.org/abs/2405.05466
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
0
9
800
ML Safety Daily
ML Safety Daily@topofmlsafety·
"Generate human-readable adversarial prompts in seconds, ∼800× faster than existing optimization-based approaches. We train the AdvPrompter using a novel algorithm that does not require access to the gradients of the Target LLM." arxiv.org/abs/2404.16873
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
0
9
891
ML Safety Daily
ML Safety Daily@topofmlsafety·
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions Improve LLM robustness by teaching them to prioritize and selectively ignore instructions based on their source. arxiv.org/abs/2404.13208
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
1
4
785
ML Safety Daily
ML Safety Daily@topofmlsafety·
LLM Agents can Autonomously Exploit One-day Vulnerabilities GPT-4 can autonomously exploit 87% of real-world one-day vulnerabilities, identified in a dataset of critical severity CVEs, compared to 0% for all other tested models arxiv.org/abs/2404.08144
ML Safety Daily tweet media
English
1
16
41
77.1K
ML Safety Daily
ML Safety Daily@topofmlsafety·
Foundational Challenges in Assuring Alignment and Safety of Large Language Models "Identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs)... we pose 200+ concrete research questions." llm-safety-challenges.github.io/challenges_llm…
ML Safety Daily tweet media
English
0
0
12
839
ML Safety Daily
ML Safety Daily@topofmlsafety·
Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning Method for LLM unlearning that outperforms existing gradient ascent methods on a synthetic benchmark, avoiding catastrophic collapse. arxiv.org/abs/2404.05868
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
0
9
741
ML Safety Daily
ML Safety Daily@topofmlsafety·
JailbreakBench is an LLM jailbreak benchmark with a dataset for jailbreaking behaviors, collection of adversarial prompts, and a leaderboard for tracking the performance of attacks and defenses on language models. arxiv.org/abs/2404.01318
ML Safety Daily tweet media
English
0
6
21
2.7K
ML Safety Daily
ML Safety Daily@topofmlsafety·
"We introduce methods for discovering and applying sparse feature circuits. These are causally implicated subnetworks of human-interpretable features for explaining language model behaviors." arxiv.org/abs/2403.19647…
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
1
12
1.2K
ML Safety Daily
ML Safety Daily@topofmlsafety·
Vulnerability Detection with Code Language Models: How Far Are We? Exposes flaws in existing datasets for vulnerability LLMs, introduces a more accurate dataset, demonstrating that current models, including GPT-3.5 and GPT-4, perform poorly on it. arxiv.org/abs/2403.18624
ML Safety Daily tweet media
English
0
0
8
637
ML Safety Daily
ML Safety Daily@topofmlsafety·
Jailbreaking is Best Solved by Definition Existing defenses against LLM jailbreaks fail; a successful defense must accurately define what constitutes unsafe outputs, with post-processing emerging as a robust solution given a good definition. arxiv.org/abs/2403.14725
ML Safety Daily tweet mediaML Safety Daily tweet media
English
0
1
6
871