Rishabh Bhardwaj

135 posts

Rishabh Bhardwaj banner
Rishabh Bhardwaj

Rishabh Bhardwaj

@rishabh15

Aligning LLMs to be safer! A researcher and a founder.

Katılım Temmuz 2010
33 Takip Edilen64 Takipçiler
Sabitlenmiş Tweet
Rishabh Bhardwaj
Rishabh Bhardwaj@rishabh15·
📣📣 New research on AI safety Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic We are introducing a new approach to prevent LLM's compromised safety Paper: arxiv.org/pdf/2402.11746… Code and safety evaluation dataset: github.com/declare-lab/re… Data: huggingface.co/datasets/decla… Did you know enterprise adaptation of LLMs comes at the cost of compromised safety? If not, our research might be interesting to you. We propose RESTA: Restoring Safety through Task Arithmetic. 👉 RESTA is a simple, fast, and effective approach that provides a no-cost solution to model safety realignment. At the core of RESTA, we add a homegrown safety vector to the model to bring back its safety while maintaining its task performance. 👉 Our evaluations of fine-tuned Llama-2 on CatQ (a multilingual safety evaluation benchmark we created) show a sharp drop in unsafety score from 33.57% to 12.17% in PEFT and from 22.16% to 4.34% in Full-FT averaged across fine-tuning domains. 👉 To gauge the effect of the safety vector beyond categories in CatQ, we evaluate RESTA on three existing safety evaluation benchmarks— HarmfulQ, AdevrsarialQ, and DangerousQ. We observe a reduction in the unsafety score from 18.59% to 5.14% in PEFT and from 9.16% to 1.55% in Full-FT when averaged across benchmark datasets and fine-tuning domains. 👉 The effectiveness of RESTA is evident across languages, as seen in the 26.2% reduction in PEFT and 21.37% reduction in Full-FT on the Vietnamese CatQ. Similar improvements are observed for CatQ in Chinese, with a reduction of 17.35% in PEFT and 24.54% in Full-FT. For more interesting information, please have a look at our work at lnkd.in/gr9dNz5H. Thanks to my great collaborators, @soujanyaporia and Duc Anh Do! #NLProc #Safety #LLM @llm_sec @topofmlsafety @cohere @dair_ai and hoping @seb_ruder @seraphinagt @jeremyphoward @omarsar0 like it.
Rishabh Bhardwaj tweet mediaRishabh Bhardwaj tweet mediaRishabh Bhardwaj tweet media
English
1
1
11
1.4K
Rishabh Bhardwaj retweetledi
Varun Gumma
Varun Gumma@VarunGumma23·
[1/n] What if your AI agents stop following your instructions? That is what we study extensively in our research: arxiv.org/abs/2509.26495
English
1
4
10
12.3K
Rishabh Bhardwaj retweetledi
Paul Röttger
Paul Röttger@paul_rottger·
Today, we are releasing MSTS, a new Multimodal Safety Test Suite for vision-language models! MSTS is exciting because it tests for safety risks *created by multimodality*. Each prompt consists of a text + image that *only in combination* reveal their full unsafe meaning. 🧵
Paul Röttger tweet media
English
1
20
47
3.9K
Rishabh Bhardwaj retweetledi
Haonan Li
Haonan Li@haonanlp·
We are excited to announce Libra-Leaderboard: The first LLM leaderboard dedicated to balancing safety and capability in LLMs. As AI advances, ensuring its safety becomes more critical than ever. By prioritizing safety measurement, we aim to inspire the AI community to make safety a core factor in model evaluation and development, alongside capability.🚀 #LLM #AI #AISafety #ResponsibleAI Website: leaderboard.librai.tech/LeaderBoard Paper: arxiv.org/abs/2412.18551 Thanks to all the contributors: @han_xudong @_ZenanZhai @honglin_mu @_HaoWang @_ZhenxuanZhang @yilin_link @ShomLinEd @realReasonWang @ArtemShelmanov @xiangyuqi_pton @_YuxiaWang @_DonghaiHong @youliang_yuan @_MengChen @HaoqinT @FajriKoto @ttk_kuribayashi @HaoqinT @rishabh15 @BingchenZhao @yawen_duan @_YiLiu @_YaodongYang @dongyp13 @soujanyaporia @stefan_fee @waterluffy @RenHector @Emad_A_Alghamdi @IGurevych @preslav_nakov @monojitchou @eltimster
Haonan Li tweet media
English
1
15
31
3.9K
Rishabh Bhardwaj retweetledi
Vrushank
Vrushank@retrovrv·
Productionizing Gen AI is as much a "platform engineering" challenge as it is an AI challenge This was clear at our @PortkeyAI practitioners' dinner in 🇸🇬 where we had AI leads from Mediacorp, OCBC Bank & others share their real production stories..
Vrushank tweet media
English
3
3
12
505
Rishabh Bhardwaj retweetledi
Rada Mihalcea
Rada Mihalcea@radamihalcea·
The new GSM-Symbolic paper from Apple has been making waves, but we published very similar findings earlier this year. Using nearly the same symbolic template methodology on GSM8k problems, we demonstrated the reasoning limitations of LLMs. arxiv.org/pdf/2401.09395
Rada Mihalcea tweet media
English
6
71
271
78.8K
Rishabh Bhardwaj retweetledi
Gabriel Chua
Gabriel Chua@gabrielchua·
this week at ai wednesdays, we had @rishabh15 and Tej share more about their research on automated red teaming & guardrails ~ we still have some slots for sharings in nov ~ drop me a dm or reply here if you’re keen
Gabriel Chua tweet mediaGabriel Chua tweet media
English
1
1
2
235
Rishabh Bhardwaj retweetledi
Sumit
Sumit@_reachsumit·
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse Introduces a metric for evaluating LLM trustworthiness in RAG systems, and a framework to improve grounded responses. 📝arxiv.org/abs/2409.11242 👨🏽‍💻github.com/declare-lab/tr…
Sumit tweet mediaSumit tweet mediaSumit tweet mediaSumit tweet media
English
0
4
15
1.5K
Rishabh Bhardwaj
Rishabh Bhardwaj@rishabh15·
It's a very common issue, and it's not the first time I have seen it. The same applies to Indian AI investments, hype and marketing skills bring in more money than quantitative proofs of the technology.
Danish Pruthi@danish037

I have been meaning to call out all the bullsh*t that @NirantK is all about! This thread does a good job. This “Top 5 GenAI scientist” has never contributed anything to science! No one would even recognize him if he were to attend a serious scientific conference.

English
0
0
0
259
Raj Dabre
Raj Dabre@prajdabre·
@psidharth567 Ok thanks but I don't work for Sarvam, so ranting to me doesn't do anything :/
English
1
0
0
70
Rishabh Bhardwaj retweetledi
Daniel van Strien
Daniel van Strien@vanstriendaniel·
WALLEDEVAL is an AI safety testing toolkit for large language models (LLMs). It supports open-weight and API-based models and features over 35 safety benchmarks, including multilingual safety and prompt injections. huggingface.co/papers/2408.03…
English
0
1
4
351
Rishabh Bhardwaj
Rishabh Bhardwaj@rishabh15·
@pratikpoddar Not everyone 😀 At least the research community tends to be happy in general (those who are serious about research)
English
0
0
0
53
Pratik Poddar
Pratik Poddar@pratikpoddar·
Somewhere as we grow up, priorities changed from pleasure maximisation to pain minimisation. Not sure if it's good or bad. Only happening with me or everyone?
English
13
2
43
6.7K
Rishabh Bhardwaj
Rishabh Bhardwaj@rishabh15·
@HarveenChadha To be honest, I was recently thinking about what you just typed here – all three points, and more interestingly, the last one.
English
0
0
1
251
Harveen Singh Chadha
Harveen Singh Chadha@HarveenChadha·
This marketing strategy works the best in India. Do an unnecessary controversy and promote your unreliable and non existent product. Unfortunately cloud customers are not the same people who watch big boss
Bhavish Aggarwal@bhash

On @Linkedin, @Microsoft and their wokeness. As an Indian institution, Ola is for genuine actions on diversity. We run one of the largest women only automotive plants. Not 1 out of 10 lines, or a small section, but the whole plant! Almost 5000 women now and will grow to tens of thousands in the coming years. And regarding gender inclusivity, we don’t need lectures from western companies on how to be inclusive. Our culture didn’t need pronouns to be inclusive for thousands of years. On a personal note, I had visited Ayodhya last year and learnt about how transgenders had been accorded special respect in our culture from ancient times! Here’s a short video from our national broadcaster DD on this - youtube.com/watch?v=goDQFI…. On the other hand, the pronouns issue I wrote about is a woke political ideology of entitlement which doesn’t belong in India. I wouldn’t have waded into this debate but clearly Linkedin has presumed Indians need to have pronouns in our life, and that we can’t criticise it. They will bully us into agreeing with them or cancel us out. And if they can do this to me, I’m sure the average user stands no chance. As a founder and CEO, this western DEI system has a major impact on my business as it grows an entitlement mindset in our professional lives and I will fight it. This situation brings me to the need for us to build our own Indian tech platforms. I’m not against global tech companies. But as an Indian citizen, I feel concerned that my life will be governed by western Big Tech monopolies and we will be culturally subsumed as the above experience shows. This is not about Ola or any of my companies. Ola is too small to make any impact against this. I want to confront this forced ideology as a free thinking Indian and do what I can in my capacity. So here are the actions I’m taking. Putting my money where my mouth is. ⁃While we can’t do anything about Linkedin’s monopoly overnight, I’m making a commitment to work with the Indian developer community to build a DPI social media framework. DPIs like UPI, ONDC, Aadhaar etc are a uniquely Indian idea and is even more needed in the world of social media. The only “community guidelines” should be the Indian law. No corporate person should be able to decide what will be banned. Data should be owned by the creators instead of being owned by the corporates who make money using our data and then lecture us on “community guidelines”! ⁃Since LinkedIn is owned by Microsoft and Ola is a big customer of Azure, we’ve decided to move our entire workload out of Azure to our own @Krutrim cloud within the next week. It is a challenge as all developers know, but my team is so charged up about doing this. ⁃Any other developer who wants to move out of Azure, we will offer a full year of free cloud usage. As long as you don’t go back to Azure after that! Mail us on exitazure@olakrutrim.com. Offer is perpetually open! Read more here: blog.olakrutrim.com/ola-on-linkedi…

English
13
38
285
9K
Rishabh Bhardwaj
Rishabh Bhardwaj@rishabh15·
@Krutrim, @bhash Do reach out to us if you want a systematic way to identify such non-idealities. 🙂 We are researchers who collectively formed #WalledAI to help make systems safer for enterprises and their users.
Rishabh Bhardwaj tweet media
English
0
0
2
35