EdinburghNLP

1.3K posts

EdinburghNLP banner
EdinburghNLP

EdinburghNLP

@EdinburghNLP

The Natural Language Processing Group at the University of Edinburgh.

Edinburgh, Scotland Katılım Mayıs 2017
159 Takip Edilen13.7K Takipçiler
Sabitlenmiş Tweet
EdinburghNLP
EdinburghNLP@EdinburghNLP·
Join our PhD programme in Designing Responsible Natural Language Processing at the UKRI AI Centre for Doctoral Training, University of Edinburgh. Applications are now re-opened for Home fee status candidates (past candidates need not re-apply). responsiblenlp.org
English
0
7
15
5.7K
EdinburghNLP retweetledi
Pasquale Minervini
Pasquale Minervini@PMinervini·
Organisations -- sponsorship opportunities are available for #EACL2027, the flagship European conference in computational linguistics, taking place in Athens in March 2027! Support the #NLP community and connect with researchers and practitioners: 2027.eacl.org
Pasquale Minervini tweet media
English
0
5
17
1.3K
EdinburghNLP retweetledi
Verna Dankers
Verna Dankers@vernadankers·
Today + tomorrow I'm at @aclmeeting '26, and early this morning, I received the 2026 ACL Computational Linguistics Doctoral Dissertation Award!! 🙀🏆👩‍🎓It's the icing on the cake of my excellent PhD experience in Edinburgh @Edin_CDT_NLP @EdinburghNLP. (1/3)
Verna Dankers tweet mediaVerna Dankers tweet media
ACL 2027@aclmeeting

Don't forget to attend the "ACL Computational Linguistics Doctoral Dissertation Award 2026" session on Mon Jul 6, 9:009:30 AM (Location: Promenade) in the "Session 8: Oral/Poster Session D" (1/2) #ACL2026NLP

English
19
18
147
20.4K
EdinburghNLP retweetledi
Ivan Titov
Ivan Titov@iatitov·
On-policy self-distillation: the teacher saw the solution; the student didn't. Copy that confidence and you gain in-domain but lose out-of-domain reasoning. GeoSD scales the teacher's pull by overlap and bounds drift - OOD +5.7–8.6 pts, ID largely intact. Led by @chatruncata
Josip Jukic@chatruncata

Self-distillation gives LLMs dense reasoning feedback without requiring a larger teacher. But ID gains can come at the cost of OOD generalization. New joint work with @iatitov : "Geometric Self-Distillation for Reasoning Generalization" (arxiv.org/abs/2607.06855)

English
1
8
77
14.9K
EdinburghNLP retweetledi
Josip Jukic
Josip Jukic@chatruncata·
Self-distillation gives LLMs dense reasoning feedback without requiring a larger teacher. But ID gains can come at the cost of OOD generalization. New joint work with @iatitov : "Geometric Self-Distillation for Reasoning Generalization" (arxiv.org/abs/2607.06855)
English
4
10
48
17.2K
EdinburghNLP retweetledi
Antonio Valerio Miceli Barone
Antonio Valerio Miceli Barone@AVMiceliBarone·
Wanna improve your LRM ability to reason logically about code?
Antonio Valerio Miceli Barone tweet media
English
1
2
6
607
EdinburghNLP retweetledi
Pasquale Minervini
Pasquale Minervini@PMinervini·
New blog post on our {ICML, ACL} 2026 papers, plus some new interesting results on open-ended learning and multi-modal retrieval! @icmlconf @aclmeeting 🧵
Pasquale Minervini tweet media
English
1
5
19
1.3K
EdinburghNLP retweetledi
Cyrus Wai-Chung Kwan
Cyrus Wai-Chung Kwan@cyruskwan1997·
Can LLMs self-improve on open-ended tasks without curated data? SCOPE says yes: co-evolving self-play and rubric-based self-judging: - Challenger creates frontier tasks - Solver answers them through multi-turn retrieval - Judge generates rubrics and scores responses
GIF
English
3
12
56
3.7K
EdinburghNLP retweetledi
Aryo Pradipta Gema
Aryo Pradipta Gema@aryopg·
What makes a transformer attention head a retrieval head? In our new paper, we argue that for **non-literal retrieval**, it is not enough to ask where a head attends. We should also ask what it writes. Introducing **LOCOS: Logit-Contribution Scoring for retrieval heads.**
English
4
22
50
16K
EdinburghNLP retweetledi
Fazl
Fazl@FazlBarez·
"Old Habits Die Hard" Once a model refuses/make things up mid-convo, it tends to keep going. We show why: past behavior leaves the model in a region of latent space it struggles to leave. congrats to wonderful @AdiSimhi @mtutek @boknilev and Shay Cohen @EdinburghNLP!
Fazl tweet media
English
2
2
13
1.9K
EdinburghNLP retweetledi
Edoardo Ponti
Edoardo Ponti@PontiEdoardo·
I'll be in Seoul for #ICML2026 to present - fast and expressive MTP - blending SFT and RLVR with prefix sampling - faster differentiable sparse attention - self-improving world models with latent actions Dm me if you'd like to meet and chat (incl. PhD/postdoc openings in my lab)!
English
2
6
78
4.6K
EdinburghNLP retweetledi
Edoardo Ponti
Edoardo Ponti@PontiEdoardo·
When you build an LLM agent team (planner, executor, verifier, ...), cost & accuracy stop being properties of one model. Instead, they depend on which model fills which role, and where it's hosted. The right choices can lift accuracy by up to +44% or match it for 12× lower cost.
Edoardo Ponti tweet media
English
2
9
20
1.9K
EdinburghNLP retweetledi
Joshua Ong @ ICML
Joshua Ong @ ICML@joshuaongg21·
🚀🚀 Introducing Pythagoras-Prover 🚀🚀 🔹 The smallest theorem prover to date – Pythagoras-Prover-4B 🔹 The first proof-of-concept diffusion theorem proving model – Pythagoras-Prover-Diffusion 4B With efficient LoRA training and an 8192-token context length, Pythagoras-Prover achieves strong MiniF2F results with far smaller models. 🏆 MiniF2F - Pythagoras-Prover-4B achieves 86.1% Pass@32, outperforming DeepSeek-Prover-V2-671B. - Pythagoras-Prover-32B achieves 89.8% Pass@32 and 92.6% Pass@2024, achieving the best MiniF2F result reported to date. 🌏Website: pythagoras-lm.github.io 📝 arXiv: arxiv.org/pdf/2606.12594 🤖 Pythagoras-Prover-4B: huggingface.co/Pythagoras-LM/… 🔐 Our models, data, and the training pipeline will be released (gradually) soon. A huge thank to my amazing collaborators: @zhengzhao97 @CMihaelaStoian Qiyuan Xu, @haonanlp @WendaLi8 Shay Cohen @e_giunchiglia
Joshua Ong @ ICML tweet mediaJoshua Ong @ ICML tweet media
English
6
26
99
154.4K
EdinburghNLP retweetledi
Ivan Titov
Ivan Titov@iatitov·
Can we reorganize an LLM into a compact backbone + swappable memories, learned jointly through pruning and distillation? Especially happy this came out of an MSc project: great work by @KGalliamov_ with @ChoenniRochelle @AmsterdamNLP @EdinburghNLP Thread + paper below 👇
Karim Galliamov@KGalliamov_

Takeaway: LLM knowledge need not be tangled together in a single set of weights. Some capacity can be reallocated between a shared backbone and external, swappable memories. Paper: arxiv.org/abs/2605.29075 with @iatitov and @ChoenniRochelle

English
0
5
18
5.9K
EdinburghNLP retweetledi
Pasquale Minervini
Pasquale Minervini@PMinervini·
you could run all experiments in our "outstanding paper award" ICLR 2021 paper overnight on a single GPU, and the method is still SOTA in that field (see our ICML'25 follow-up) -- you don't have to be in a frontier lab to do cool stuff and have an impact
Pasquale Minervini tweet media
Yi Ma@YiMaTweets

I’m getting increasingly annoyed by young people complaining that they cannot do AI-related research unless they join big industrial labs… well, here is my reply: academia is supposed to work on ideas that money cannot buy!

English
1
3
91
8.9K
EdinburghNLP retweetledi
Karim Galliamov
Karim Galliamov@KGalliamov_·
Takeaway: LLM knowledge need not be tangled together in a single set of weights. Some capacity can be reallocated between a shared backbone and external, swappable memories. Paper: arxiv.org/abs/2605.29075 with @iatitov and @ChoenniRochelle
English
0
4
10
3.5K
EdinburghNLP retweetledi
Karim Galliamov
Karim Galliamov@KGalliamov_·
New preprint: Knowledge Offloading (KOFF). LLMs store general abilities and domain-specific knowledge in the same parameter set. Can we reorganize this: extract a shared core and move more specialized capacity into small swappable memories?
Karim Galliamov tweet media
English
4
6
29
2.5K
EdinburghNLP retweetledi
Edoardo Ponti
Edoardo Ponti@PontiEdoardo·
Goodbye top-k in hierarchical attention! We devised DashAttention, which is adaptively sparse (compute is allocated based on the information structure of the query) and end-to-end differentiable. DashAttention pushes the accuracy–efficieny frontier over NSA and InfLLMv2!
Yuxiang Huang@yxyxyyy6

[1/n] Can a model learn *where* and *how much* information it should attend to, and do so efficiently? We introduce DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention! This pushes the accuracy-efficiency frontier in LLMs.

English
1
35
228
24.7K
EdinburghNLP retweetledi
Pasquale Minervini
Pasquale Minervini@PMinervini·
Hey, that's @rohit_saxena's work on this page of the Stanford HAI's 2026 Artificial Intelligence Index Report!
Pasquale Minervini tweet media
English
2
5
11
1.5K