Om Khangaonkar

49 posts

Om Khangaonkar

Om Khangaonkar

@reachomk

PhD student @ucdavis. I like representation learning.

Katılım Kasım 2020
392 Takip Edilen27 Takipçiler
Om Khangaonkar retweetledi
corsaren
corsaren@corsaren·
Local minima are extremely rare in high dimensional spaces, so if you ever feel stuck in a rut it’s probably just because you aren’t considering a wide enough set of orthogonal options
English
70
331
4.1K
122.9K
Om Khangaonkar
Om Khangaonkar@reachomk·
@eding421 you have no idea how annoying it was I put hours into my eccv and colm reviews and half of them came back with obvious chatgpt rebuttals… like atp how do I even know ur numbers aren’t hallucinated
English
0
1
2
116
Om Khangaonkar
Om Khangaonkar@reachomk·
If reviewers feel the rebuttal is entirely AI generated, they are even less likely to raise the score. Also, based on past author responses I think are LLM generated, they seem to argue over every minor suggestion, even when it is simple to fix and could strengthen the paper.
Lei Li@_TobiasLee

opensource our rebuttal skills, distilled from a lot of successful rebuttals from myself & labmates; also my experience serving as ACs for ARR. Hope this help for your EMNLP and incoming NeurIPS reviews :) github.com/TobiasLee/Rebu…

English
1
1
4
380
Om Khangaonkar retweetledi
Jiawei Yang
Jiawei Yang@JiaweiYang118·
Two months ago, I vaguely posted a number: 0.9 FID, one-step, pixel space. Now it is 0.75, and can be even lower. Many wonder how. I thought it might end as a small FID prank: simple and deliberate. It started with one question: can FID be optimized directly, and what does it reveal? Introducing FD-loss.
Jiawei Yang tweet media
English
56
158
965
239K
Om Khangaonkar retweetledi
Hanwen Jiang
Hanwen Jiang@hanwenjiang1·
GPT Image 2 has been deeply unsettling to me in the best way. Some of its outputs make it hard for me to keep using the old criterion of vision, especially the old definition of visual representation learning. Thus, I wrote this essay as a reflection on that shift: why knowledge may be the right name for what vision once called representation, and what can be the ultimate formulation for representation learning. (An unexpected side path: it also led me to think about the relation between knowledge and representation through the old calligraphic relation between spirit and form 😃 hwjiang1510.github.io/blogs/knowledg…
English
5
26
151
17.1K
Om Khangaonkar retweetledi
ICML Conference
ICML Conference@icmlconf·
So who's gonna set up the Polymarket for when ICML decisions are gonna drop? 👀📈⏳📉
English
13
10
241
62.7K
Om Khangaonkar retweetledi
Anand Bhattad
Anand Bhattad@anand_bhattad·
“there has been limited evidence that generative vision models have developed strong understanding capabilities.” “Limited” or strong evidence from 2023 and 2024 🤔 👇 x.com/anand_bhattad/…
Radu Soricut@RSoricut

Meet Vision Banana 🍌 from @GoogleDeepMind! We provide strong evidence that image generators are generalist vision learners. Traditional computer vision tasks (segmentation, depth estimation, normal prediction) can now be performed at/near SOTA with a single generalist model derived from an image generation model. 🖼️ Explore the results: vision-banana.github.io 📄 See details at: arxiv.org/abs/2604.20329

English
2
4
29
4.5K
Le Thien Phuc Nguyen ✈️ CVPR 2026
Has any one received any news about the CVPR 2026 broadening participation program? I remember filling it very soon…
English
2
0
3
229
Om Khangaonkar
Om Khangaonkar@reachomk·
@ducha_aiki @hpirsiav10 understands the world will be able to spot most of the errors. However, they are intended to take time and be "tricky." 3) Some pairs are intended to take time to solve. This is why we use the plausible/implausible distinction, as plausible pairs would look normal at a (3/4)
English
1
0
1
20
Om Khangaonkar
Om Khangaonkar@reachomk·
@sir4K_zen @ducha_aiki @hpirsiav10 The point is that it’s supposed to be a spatial reasoning benchmark, unlike BLINK or others which test low level vision understanding. It might take you a few seconds, but you get it right because you understand 3D.
English
0
0
1
41
Om Khangaonkar retweetledi
Kwang Moo Yi
Kwang Moo Yi@kwangmoo_yi·
Khangaonkar et al., "Multimodal Large Language Models Cannot Spot Spatial Inconsistencies" A benchmark for spotting spatial inconsistencies. MLLMs are still quite far from being accurate. Reminds me of various automated benchmarks that have recently been shown to be misleading.
Kwang Moo Yi tweet media
English
2
1
16
1.4K