mishka

301 posts

mishka banner
mishka

mishka

@mishakgg

Token-maxxer 阴阳 Apprentice predictor, chess, ai agents, playtesting

Katılım Ekim 2021
150 Takip Edilen93 Takipçiler
mishka
mishka@mishakgg·
I have been finding counterexamples to tons of research with pro, genuinely scary smart imo. For anything coding related its not really worth unless you are trying to find some really specific high level optimizations, and if you ask for an implementation plan it does overengineer your small project with too many phases and tests.
English
0
0
2
15
brock
brock@brockjelmore·
@delitzer from my experience they serve different purposes Ultra is good at getting a massive task done, Pro feels smarter imo could be wrong but thats been my impression
English
1
0
2
138
brock
brock@brockjelmore·
its crazy how good Sol 5.6 Pro is i wish i could access it in the desktop app even if it cant use the full harness, it would be so nice
English
5
1
10
2.4K
Odysseus | phylax.systems
Odysseus | phylax.systems@odysseas_eth·
@brockjelmore You can create a skill, where basically you instruct a terra agent to use computer-use to be steered by pro. So basically: terra 1: orchestrates upload of zip file + prompt to pro pro: gets the input, works, steers back terra 1 terra 2, ...n: does the impl work
English
2
0
5
276
mishka
mishka@mishakgg·
@Shekswess @iamgrigorev can you explain what advantage auto-research has over a long codex session with sub agents?
English
0
0
0
22
mishka
mishka@mishakgg·
@karirogg Hello, I (gpt 5.6 pro) have found about 15 candidate conjecture counterexamples (2 verified by co-author) would you be interested in this data to benchmark it , I am also running about 10 instances concurrently so number may rise.
English
0
0
0
74
Kári Rögnvaldsson
Kári Rögnvaldsson@karirogg·
We just evaluated Opus 5 on MathArena! It does well on ArXivMath and ArXivLean, reaching 4th and 2nd place, and crushes BrokenArXiv, reaching first place with over 90% accuracy.
Kári Rögnvaldsson tweet media
English
6
6
33
4.1K
mishka
mishka@mishakgg·
@enorton1903 Ozel okul isleten dostları var muhtemelen
Türkçe
0
0
0
15
mishka
mishka@mishakgg·
@fjzzq2002 Hello, can i share with you about 15-20 papers in which I (gpt 5.6 pro) found conjecture counterexamples and you could benchmark it with claude models? i dont have a max claude subscription to get this done .
English
0
0
0
2
Ziqian Zhong
Ziqian Zhong@fjzzq2002·
GPT-5.6 sol loves rejecting papers more than any previous models & rejected all 30 papers randomly downloaded from Arxiv. Might correlate with the low scores in NeurIPS 😅
Ziqian Zhong tweet media
English
16
16
384
38.7K
pr337h4m
pr337h4m@pr337h4m·
@Ricburton have a feeling he’ll bounce back the biopic is gonna go so hard
English
1
0
56
9.8K
mishka
mishka@mishakgg·
@cappucher Also, for one proof it used a paper that was published / submitted 5 days prior. that was amazing to see. it was also after many iterations of me telling it to keep working on the problem
English
0
0
0
21
Cappucher
Cappucher@cappucher·
@mishakgg That’s amazing. Which branch of math was this in?
English
3
0
2
70
Cappucher
Cappucher@cappucher·
I think GPT is getting decent at disproving (relatively?) non-trivial conjectures upon general request. chatgpt.com/share/6a6aa28f… After giving it the prompt in the conversation above, it actually seemed to disprove two random 3 month old conjectures out of random arXiv papers - one relatively simple gen func combinatorics problem that the authors were probably too lazy to brute force and another graph theory problem that seems to require some level of ingenuity. The graph theory was simple enough where relatively elementary lines of attack seemed to kill the conjecture though
English
3
3
20
4.8K
mishka
mishka@mishakgg·
@cappucher I also plan on making a benchmark for these papers but I do not have a claude sub so its mostly going to be gpt models if anyone wants to chip in thatd be great too lol
English
0
0
0
16
mishka
mishka@mishakgg·
analytic and additive number theory; Collatz-type arithmetic dynamics; probabilistic combinatorics and random graphs; extremal, domination, coloring, and algebraic graph theory; commutative algebra and polyomino ideals; complex and functional analysis; mathematical optimization; quantum resource theory; nonlinear PDEs; coding theory and permutation groups; cryptographic algebra and hash-function analysis; fine-grained complexity and dynamic algorithms; query complexity and graph reachability; computability and finite diagonalization. made gpt get this list probably right
English
0
0
0
26
mishka
mishka@mishakgg·
@AlexanderKalian Gpt-6 is going to be better at verifying solutions than an above average phd candidate / average phd level academic for almost all math fields. So, your claim is therefore anulled by the fact that models will get better faster than most people can get to that level at this point.
English
1
0
3
108
Dr Alexander D. Kalian
Dr Alexander D. Kalian@AlexanderKalian·
If you're trying to vibe-solve open problems in mathematics using GPT-5.6 Sol, then getting a solution verified is a major bottleneck. Even when produced by academia, big proofs can go for years before formal verification. So if you don't understand the solution, can't defend it well against sceptical experts, and aren't affiliated with any credible institution - why would you expect any notable mathematician to take it seriously and delegate time to reviewing your work? Even a small fry researcher like me is constantly getting spammed with independent researchers who want their work reviewed - and most are crackpots. The default is to shrug. It's not always fair, but this is the reality. Prominent professors, Fields Medalists etc. have it much much worse. If you are very lucky, you may be able to catch someone's attention at a conference or evening lecture, by asking an intelligent question. And if that happens - you better be ready to explain your solution simply to inspire further interest - and then later rigorously explain and defend it to top experts in the field. So if you wanna use AI to do frontier mathematical research - you should realistically still develop a frontier mathematics skillset and intuition.
English
29
4
111
5.4K
mishka
mishka@mishakgg·
@Avenoxai nice hayırlı olsun iyi fiyat bicersin bi 20 sene sonra
Türkçe
0
0
0
955
Avenox
Avenox@Avenoxai·
sürekli "AI'a her ay bu kadar para veriyorsun, salak mısın" yazan arkadaşlar; şu an kanala gelen sponsorluk tekliflerini kabul etsem, modellere harcadığımın katlarını her ay geri kazanıyor olurdum. ve işin özü şu: bu modelleri bu kadar kullanmamış olsaydım, beni izlemiyor olurdunuz. sponsor da yazmıyor olurdu. alın size net ROI, çok kasmayın :D
Türkçe
3
0
12
1.2K
mishka
mishka@mishakgg·
@Avenoxai neymis o kalem ya ilk kez duydum, kac token yakmamız lazım
Türkçe
1
0
0
103
Avenox
Avenox@Avenoxai·
@mishakgg openai kalem yolladı olmaz mı?
Türkçe
1
0
0
138
mishka
mishka@mishakgg·
@GuanyangW Does this account for further iterations by prompts with the instance for Prompt B ?
English
1
0
1
674
mishka
mishka@mishakgg·
@PhDPersuasion I have found about 4 counterexamples and I have been just emailing the authors, maybe they give me some credit in the footnotes, although doesnt really matter
English
0
0
1
22
Equation Persuasion
Equation Persuasion@PhDPersuasion·
Posting this here in case someone else tries to say they did it first. (I still need to perform some verification which I've been too lazy to do but so far I think it's correct) No it's not my field. It's ChatGPT's field, which is now all fields.
Equation Persuasion tweet mediaEquation Persuasion tweet mediaEquation Persuasion tweet mediaEquation Persuasion tweet media
English
2
1
4
662
mishka
mishka@mishakgg·
@togelius I dont know why we dont just benchmark how strong of a chess engine they can code, not this nonsense
English
1
0
2
145