Wanda Wang
9 posts


Comparing a new image encoder to VJEPA 2.1 actually says a lot about the future of computer vision.
Video features are advancing rapidly, and unified models that ingest both image and video modalities will eventually turn into the industry standard.
This area of research is super interesting, and I can't wait to see more papers leveraging this duality in VJEPA 2.1 at the next conferences.
For example, it would be cool to learn a task at the image level (e.g., detection or depth) and easily extend it to video with close to zero effort.
Massimiliano Viola@massiviola01
Image foundation models never stop surprising!😮 And this time, it's not the usual big tech players. A few days ago, @robbyant_brain released LingBot-Vision, a new family of vision encoders built natively for dense spatial perception.
English
Wanda Wang retweetledi

Image foundation models never stop surprising!😮
And this time, it's not the usual big tech players.
A few days ago, @robbyant_brain released LingBot-Vision, a new family of vision encoders built natively for dense spatial perception.

English
Wanda Wang retweetledi

AI Engineer Run Club
no better way to make friends from Lisbon, London, Italy, Seattle, and many more places
@aiDotEngineer @KernelLabs_ai @EntireHQ @lizziepika @FanaHOVA


English





