🎶How can we learn to follow along to a music performance in a score without direct supervision?🎶
Introducing FuSiLi (Fused Sinkhorn-Localized Similarity)!
A multimodal contrastive learning approach that learns precise local alignments using only global supervision.
🧵
Introducing 🎼🫧MulTTiPop, an evaluation dataset of multitrack pop music for automatic music transcription!
MulTTiPop contains 3.5 hours (572 segments) of aligned multitrack pop audio and MIDI, and can be previewed at gclef-cmu.org/multtipop
🧵1/7
🎼 Can we turn music into readable code?
Excited to share Decomposer, a post-training framework for symbolic music decompilation!
Given MIDI, Decomposer recovers an executable program in Strudel that reconstructs the music while exposing patterns, harmony, rhythm, and voices as editable code.
🧵 [1/5]