
EdinburghNLP
1.3K posts

EdinburghNLP
@EdinburghNLP
The Natural Language Processing Group at the University of Edinburgh.








Don't forget to attend the "ACL Computational Linguistics Doctoral Dissertation Award 2026" session on Mon Jul 6, 9:009:30 AM (Location: Promenade) in the "Session 8: Oral/Poster Session D" (1/2) #ACL2026NLP

Self-distillation gives LLMs dense reasoning feedback without requiring a larger teacher. But ID gains can come at the cost of OOD generalization. New joint work with @iatitov : "Geometric Self-Distillation for Reasoning Generalization" (arxiv.org/abs/2607.06855)








Takeaway: LLM knowledge need not be tangled together in a single set of weights. Some capacity can be reallocated between a shared backbone and external, swappable memories. Paper: arxiv.org/abs/2605.29075 with @iatitov and @ChoenniRochelle


I’m getting increasingly annoyed by young people complaining that they cannot do AI-related research unless they join big industrial labs… well, here is my reply: academia is supposed to work on ideas that money cannot buy!


Just learned: Software engineers used to do manual data labeling at Scale AI while Alex Wang was CEO. After he left, new leadership joined, and were HORRIFIED to learn this. Stopped it ASAP Now at Meta, software engineers are assigned manual data labeling... see the pattern?


[1/n] Can a model learn *where* and *how much* information it should attend to, and do so efficiently? We introduce DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention! This pushes the accuracy-efficiency frontier in LLMs.









