Om Khangaonkar
49 posts

Om Khangaonkar
@reachomk
PhD student @ucdavis. I like representation learning.
Katılım Kasım 2020
392 Takip Edilen27 Takipçiler
Om Khangaonkar retweetledi
Om Khangaonkar retweetledi

@eding421 you have no idea how annoying it was I put hours into my eccv and colm reviews and half of them came back with obvious chatgpt rebuttals… like atp how do I even know ur numbers aren’t hallucinated
English

If reviewers feel the rebuttal is entirely AI generated, they are even less likely to raise the score.
Also, based on past author responses I think are LLM generated, they seem to argue over every minor suggestion, even when it is simple to fix and could strengthen the paper.
Lei Li@_TobiasLee
opensource our rebuttal skills, distilled from a lot of successful rebuttals from myself & labmates; also my experience serving as ACs for ARR. Hope this help for your EMNLP and incoming NeurIPS reviews :) github.com/TobiasLee/Rebu…
English
Om Khangaonkar retweetledi


Two months ago, I vaguely posted a number: 0.9 FID, one-step, pixel space.
Now it is 0.75, and can be even lower.
Many wonder how.
I thought it might end as a small FID prank: simple and deliberate.
It started with one question: can FID be optimized directly, and what does it reveal?
Introducing FD-loss.

English
Om Khangaonkar retweetledi

GPT Image 2 has been deeply unsettling to me in the best way.
Some of its outputs make it hard for me to keep using the old criterion of vision, especially the old definition of visual representation learning.
Thus, I wrote this essay as a reflection on that shift: why knowledge may be the right name for what vision once called representation, and what can be the ultimate formulation for representation learning.
(An unexpected side path: it also led me to think about the relation between knowledge and representation through the old calligraphic relation between spirit and form 😃
hwjiang1510.github.io/blogs/knowledg…
English
Om Khangaonkar retweetledi
Om Khangaonkar retweetledi

“there has been limited evidence that generative vision models have developed strong understanding capabilities.”
“Limited” or strong evidence from 2023 and 2024 🤔 👇
x.com/anand_bhattad/…
Radu Soricut@RSoricut
Meet Vision Banana 🍌 from @GoogleDeepMind! We provide strong evidence that image generators are generalist vision learners. Traditional computer vision tasks (segmentation, depth estimation, normal prediction) can now be performed at/near SOTA with a single generalist model derived from an image generation model. 🖼️ Explore the results: vision-banana.github.io 📄 See details at: arxiv.org/abs/2604.20329
English

@ducha_aiki @hpirsiav10 Appreciate the clarification. Thanks again for sharing!
English

@reachomk @hpirsiav10 I don't mean that as a negative thing, just shared my experience.
English

Multimodal Language Models Cannot Spot Spatial Inconsistencies
@reachomk Hadi J. Rad @hpirsiav10
tl;dr: in title. The task is hard enough for humans as well - I have to spend 5-20 sec examining every object to spot the inconsistency.
arxiv.org/abs/2604.00799




English

@ducha_aiki @hpirsiav10 glance and you need to search for an inconsistency. (4/4)
English

@ducha_aiki @hpirsiav10 understands the world will be able to spot most of the errors. However, they are intended to take time and be "tricky."
3) Some pairs are intended to take time to solve. This is why we use the plausible/implausible distinction, as plausible pairs would look normal at a (3/4)
English

@sir4K_zen @ducha_aiki @hpirsiav10 The point is that it’s supposed to be a spatial reasoning benchmark, unlike BLINK or others which test low level vision understanding. It might take you a few seconds, but you get it right because you understand 3D.
English
Om Khangaonkar retweetledi

@graceluo_ @feng_jiahai @trevordarrell @AlecRad @JacobSteinhardt Nice, this is a super interesting work!
English

We trained diffusion models on a billion LLM activations, and we want you to use them!
New preprint: Learning a Generative Meta-Model of LLM Activations
Joint work with @feng_jiahai, @trevordarrell, @AlecRad, @JacobSteinhardt.
More in thread 🧵
English









