Alex Gurung ✈️ ACL26

49 posts

Alex Gurung ✈️ ACL26 banner
Alex Gurung ✈️ ACL26

Alex Gurung ✈️ ACL26

@AlexAag1234

Visiting Researcher @ServiceNowRSRCH, PhD student at @EdinburghNLP | undergrad+masters @gtcomputing

Edinburgh Katılım Temmuz 2014
774 Takip Edilen495 Takipçiler
Alex Gurung ✈️ ACL26 retweetledi
Aryo Pradipta Gema
Aryo Pradipta Gema@aryopg·
What makes a transformer attention head a retrieval head? In our new paper, we argue that for **non-literal retrieval**, it is not enough to ask where a head attends. We should also ask what it writes. Introducing **LOCOS: Logit-Contribution Scoring for retrieval heads.**
English
4
22
50
16K
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
What if you didn't have a gold-standard proxy? We also experiment on the effect of noise on proxy performance, simulating a silver-standard setting. We find surprisingly stable performance, indicating that even lower-quality proxies are a useful source of learning signal.
Alex Gurung ✈️ ACL26 tweet media
English
1
0
2
86
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
Improving at long-context tasks is hard, can we just train with short context versions instead? Our ACL paper gives a recipe to distill from these proxy tasks to long-context capabilities! I'll be presenting our oral: ⏰ 2pm Sunday 📍Promenade Come to learn about ProxyCoT! 🧵
Alex Gurung ✈️ ACL26 tweet media
English
1
2
14
538
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
Been thinking recently about how to improve credit assignment in long horizon RL? Our new MosaicLeaks blog post describes our method to accurately value actions via situational rewards, improving our privacy-aware research agent over outcome-only rewards! huggingface.co/blog/ServiceNo…
Rafael Pardinas@muchomuchacho

MosaicLeaks is now on arXiv. The Mosaic Effect captures a simple idea: small fragments can look harmless alone, but become revealing in aggregate. Deep research agents can leak enterprise information in exactly this way. 1/9

English
0
2
10
1.1K
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
Excited to share my recent work @ServiceNowRSRCH ! We introduce a new privacy-centric deep research dataset and show models frequently leak enterprise information. However, training with dense _situational_ rewards efficiently learns to jointly optimize performance and privacy
Rafael Pardinas@muchomuchacho

MosaicLeaks is now on arXiv. The Mosaic Effect captures a simple idea: small fragments can look harmless alone, but become revealing in aggregate. Deep research agents can leak enterprise information in exactly this way. 1/9

English
0
3
8
674
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
Heading to ICLR to present this (with new results and ablations) at the LIT workshop 🇧🇷 Would love to chat about all things creativity, RL, and optimizing the unverifiable - feel free to reach out!
Alex Gurung ✈️ ACL26@AlexAag1234

New preprint: How can we use latent-reasoning when initial model performance is low? We introduce LiteReason, a simple and lightweight framework that combines latent reasoning _with RL_ to reason efficiently both during and after training while retaining performance gains! 🧵

English
0
1
8
1.2K
Alex Gurung ✈️ ACL26 retweetledi
Akash Gupta
Akash Gupta@aksgupta97·
Excited to share that I will be presenting my 1st year PhD work at #ICLR2026 @iclr_conf next week in Rio 🇧🇷 ❓In-Context Learning (ICL) has been a widely used paradigm to adapt LLMs for a task with a few examples or shots. But does it also work for multimodal examples? 🧵
Akash Gupta tweet media
English
1
11
40
3.7K
Alex Gurung ✈️ ACL26 retweetledi
Irina Saparina
Irina Saparina@irisaparina·
Reasoning models are powerful, but they burn thousands of tokens on potentially wrong interpretations for ambiguous requests! 👉 We teach models to think about intent first and provide all interpretations and answers in a single response via RL with dual reward. 🧵1/6
Irina Saparina tweet media
English
1
12
35
2.8K
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
More details, results, and analysis in the paper: arxiv.org/abs/2512.02240 We found a lot of small design decisions (e.g. not adding new tokens to the vocabulary) made a large difference, I'm excited to see this area gain traction so we learn more best practices!
English
1
0
2
126
Alex Gurung ✈️ ACL26
Alex Gurung ✈️ ACL26@AlexAag1234·
New preprint: How can we use latent-reasoning when initial model performance is low? We introduce LiteReason, a simple and lightweight framework that combines latent reasoning _with RL_ to reason efficiently both during and after training while retaining performance gains! 🧵
Alex Gurung ✈️ ACL26 tweet media
English
1
5
19
2.2K