Daniel Franke

7.7K posts

Daniel Franke

Daniel Franke

@dfranke

Security researcher. Keeping the internet safe for anarchy.

New Hampshire Katılım Mayıs 2008
478 Takip Edilen1.5K Takipçiler
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
Astonishing. "GPT disprove conjecture. You should do a breakthrough and find a structured counterexample. Make no mistakes" "No" "yes. don't be a pussy." "best I can do is…" "best you can do is a full counterexample." "ok fine here's a full counterexample"
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet mediaTeortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet mediaTeortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet media
Dmitry Rybin@DmitryRybin1

Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years. The graph below has fractional flow cost 58. Any unsplittable flow (with capacity violation <=15) has cost at least 60. Chat with GPT 5.6 Pro where this was found: chatgpt.com/share/6a60b2eb…

English
5
0
96
6.4K
Daniel Franke
Daniel Franke@dfranke·
@robertgraham @halvarflake Lol, as if Congress can still pass laws. Ideally yes, but courts should do their best to applying existing law unless and until that happens. My take: AI output is a derived work of the prompt. If the prompt is too short to be copyrightable, the output isn't either.
English
0
0
0
14
Robert Graham
Robert Graham@robertgraham·
@halvarflake It's one of those things where courts SHOULDN'T be involved. It's novel enough that Congress should pass a law saying what is or isn't allowed.
English
3
0
3
523
Halvar Flake
Halvar Flake@halvarflake·
What IP, precisely, was stolen? Because copyright on LLM output is commonly assigned to the prompter? Or should the model providers own model output? That'd make open models the *only* viable path to use coding agents?
Treasury Secretary Scott Bessent@SecScottBessent

We support open-source AI and the innovation it unlocks. But open source is not open season on American IP. When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.

English
5
11
78
5.3K
Brian Graham 🦬
Brian Graham 🦬@iroasmas·
you come to me today, on the day i disprove the jacobian conjecture, and you ask me to center a div
Brian Graham 🦬 tweet media
English
54
593
9.7K
189.5K
Daniel Franke
Daniel Franke@dfranke·
It's long been my suspicion that what gets diagnosed as depression is really several different conditions that present similar symptoms, and SSRIs are highly effective for one of them and ineffective for others. Studies that don't look for this and average everything together end up concluding a significant-but-tiny effect size.
English
0
0
2
107
Crémieux
Crémieux@cremieuxrecueil·
Time for another chapter of 'I don't like Cochrane reviews'! I recently read their review of modern antidepressants for kids, and I was extremely disappointed with it. They concluded: "most newer antidepressants may reduce depression symptoms in a small and unimportant way compared with placebo." This was not a warranted conclusion, but it is a good opportunity to see some of the tricks that go into a Cochrane review and make them bad for public consumption. The biggest problem I took with this review was the use of an arbitrary effect size threshold to decide which effects were important or unimportant. For the review, if an effect was not at least a shift of |5| points on the Children's Depression Rating Scale-Revised (CDRS-R), it was considered unimportant. (And they excluded studies using other outcomes, thus limiting their statistical power, real-world utility estimation, etc.) This threshold was chosen on the basis of the CDRS-R's severity categories being separated by ~10-15 points, and then reasoning to the idea that an important difference should be at least half of that. Why should that be the case? No one knows. The authors don't explain it, and it's actually just obviously untrue. Using the review's standard deviation of 14.47, the principal estimates for various drugs were small to moderate. For example, sertraline (0.24 SDs), fluoxetine (0.20), duloxetine (0.19), and escitalopram (0.18) all had benefits consistent with substantial clinical benefits. Moreover, the authors considered the effects of these drugs in a way that was inappropriate given their own standard: The rule they chose was that each point estimate was to be considered relative to the |5|-point change boundary, where sertraline, escitalopram, and duloxetine -- for example -- all produced equivalence intervals for their estimates consistent with benefits of that size. Therefore, the authors were wrong in how they worded their conclusion, by their own standards, and the should've said 'The average benefit is probably small, but a clinically important benefit cannot be excluded for some new-generation antidepressants.' But even giving them the ability to reason from their arbitrarily-selected threshold, it's not a good standard. The review noted response/remission estimates for fluoxetine, sertraline, and duloxetine of 48%, 51%, and 53%, versus a ~40% placebo response rate, meaning that the Number Needed to Treat (NNT) for these drugs is quite small, on the order of 13, 9, and 8, respectively.* By that more useful effect standard, antidepressants actually seem to be quite effective! By the bizarre standard the review authors chose, they seem to deliver unimportant effects. Frankly, I'd go with the standard that's not arbitrary, and is actually informative because it provides quantitative information about how effective practitioners can expect the drugs to be, unlike Cochrane. This is really the biggest issue with the review, but there are some other things I could also kvetch about, and would like to because they get at how Cochrane reviews so often work. The review mishandles 'heterogeneity'. There are at least three forms of it in the review: between drugs, between trials of the same drug, and between individual patients' treatment effects. The most consequential of these is the difference between drugs, and failing to account for this is one of the most common errors you get in Cochrane reviews. I'm QTing an example right now, where they butchered a review of amyloid beta-targeting monoclonal antibodies, claiming they don't work because they conflated the effects of some drugs that do hit the target and thus work with drugs that don't hit the target and thus don't work. Of course, if you combine estimates from effective and ineffective drugs, you can make it sound like all the drugs are ineffective. But this is dishonest and dull-minded. Sertraline, fluoxetine, duloxetine, and escitalopram were all effective drugs, but paroxetine, vilazodone, desvenlafaxine, and vortioxetine were basically not. The latter four might be effective, but their estimates were very imprecise, so it's not clear, but they at least didn't show consistently significant effects or effects in the right direction (vortioxetine). But the authors still decided to throw together the first four and the latter four and then declare that the whole class doesn't work and the differences among them are unimportant! How can that even be, when the estimate for one of them goes in the wrong direction, and the estimates for the effective ones are substantial in real (NNT/NNH) terms? The authors are also unconvincing in how they deal with the heterogeneity among drugs. For example, they rank the drugs, but argue that their rankings have little meaning because all the effects are small, yet they use the pattern of estimates to suggest sertraline, escitalopram, and duloxetine as first options alongside fluoxetine. But they can't have this both ways: don't rank these things if you're going to say the differences are "small and unimportant". You have to pick a side here. Trial-level effect modifiers were also not sufficiently handled, and they were pretty much neglected. Placebo response, number of study sites, lead-in and exclusion of placebo responders, trial era, frequency and intensity of clinical contact, concomitant care or psychotherapy, comorbidity, suicidality, treatment resistance, and, hell, even DOSE were not adequately modeled features. This matters because these trial features were NOT distributed randomly. The review even notes that placebo lead-in and exclusion of placebo responders was characteristic of fluoxetine and escitalopram trials, but not sertraline and duloxetine trials, and that placebo remission was lower in the studies that used fluoxetine as the treatment. And the effect size overall was much lower than it would've been in clinically representative samples because trialists often intentionally excluded groups like young people at a high risk of suicide. Thus, the conclusions of this review are underwhelming *by construction*, and it's no wonder its conclusions differ from other reviews that don't neglect heterogeneity or exclude exactly the populations that psychiatrists are most likely to want to put on antidepressants right away. There was even more wrong with this review, but I'm going to cut this off here now because it was just really annoying, and what I've said so far is sufficient for me to dismiss them and preference more comprehensive and careful reviewers' conclusions over theirs. * For comparison, the review's suggested NN Harms (NNHs) for these drugs are 104, 15, and 187, respectively. Psychiatrists giving these out are much more likely to help than to harm their patients. Source: cochranelibrary.com/cdsr/doi/10.10…
Crémieux@cremieuxrecueil

I am extremely disappointed in Cochrane for their systematic review on amyloid beta-targeting monoclonal antibodies. There's a lot to dislike about it, but the key issue is that they lumped together approved and unapproved drugs and then acted like they were all useless.

English
5
3
49
16.2K
Daniel Franke retweetledi
Daniel Franke
Daniel Franke@dfranke·
@perrymetzger I think it's too early to assert that K3 will be any kind of inflection point for Chinese dominance. It's a shot across the bow, at least, but so was Deepseek R1 and that was a year and a half ago.
English
2
0
1
523
Perry E. Metzger
Perry E. Metzger@perrymetzger·
Now that Chinese AIs are achieving parity with the best US created models, a few short-term predictions: OpenAI will adapt just fine, they are very pragmatic, and a lot of their value add is understanding how to do inference cheaply at scale. Anthropic will literally be unable to get out of its own way, because it is too ideologically committed to AI Doomerism, but they will survive anyway, because many of their victims, pardon me, customers, will continue going back to them over and over again no matter how abusive they are, and perhaps even because of how abusive they are. They will also continue to try to use fear as a mechanism for achieving regulatory capture, but with the Chinese racing ahead of them, they are going to have more and more trouble getting a warm reception from all but the far left contingent in Congress. Look for their newly rich employees to be spending vast amounts of money post-IPO on political campaigns in support of Doomer-friendly candidates and attempts to capture the Democratic party, and for that money to have some significant effect, but for it to mostly be wasted because they’re not good at understanding their fellow humans. Musk is ideologically committed to crushing OpenAI, and so xAI is going to keep its team awake 24 hours a day if necessary until they are ahead; look for them to be at the forefront soon. I would expect xAI to eventually be near or in the lead on commercialization. Google will not be able to get out of its own way, and it may start arguing more and more for heavy regulation of AI as a way of trying to cripple its opponents. This is insane given that they have some of the best technology out there, but unfortunately, their management is simply not good enough, and not just on AI. Meta seems to be catching up, but I don’t have a strong opinion on whether they will maintain momentum. On the Chinese side, I am expecting Chinese R&D to be at the front or ahead most of the time from now on; the US will need policies that adjust for that. I am expecting Anthropic to spend a ton of money on PR and lobbying claiming that this is all through distillation or espionage, although of course it isn’t, and I am expecting that a certain fraction of Congress will be bamboozled, although a surprising fraction will not be receptive, especially after Anthropic employees ham-handedly spend too much money on political campaigns. It will make no difference, because the only thing that the US could do would be banning Chinese models from being used in the US, and of course, the US can’t actually stop other countries or China from using Chinese models; all this would do is hurt the United States and its interests. I am expecting more calls for export controls, which will delay the Chinese a bit in the short term but which will ultimately stop working at all, because the Chinese are going to control their full technology stack soon, including having EUV fabs capable of manufacturing domestically designed training and inference hardware. I am also expecting the Chinese to race ahead in robotics, and especially in military robotics. Some US robotics companies, like Musk’s companies and Anduril, will equal or exceed them, but a lot of the other US companies will be crippled by the fact that US manufacturing is too heavily regulated and restrained by stupid internal policies. They will not have problems keeping up on the technology, but scaling requires manufacturing infrastructure, and much of the US has effectively banned economic development or wants to. Places like New York State and California are already envious of Europe’s self-destruction and want in. If this continues, look for an eventual decisive Chinese military advantage. Note that this is not what I want, it is the exact opposite of what I want, but absent a big change in US policy about things like data centers, chip fabrication facilities, and just plain normal factories, it’s going to be hard.
English
46
60
329
34.2K
Joaquin
Joaquin@delrio_joaquin·
@paulg Just “bring me” for now, “make me” will be when we get personal robots
English
1
0
1
567
Daniel Franke
Daniel Franke@dfranke·
The hard part of making this immersive for the pilot is the haptic and vestibular feedback. You should feel it if you get knocked down or try to make a movement that your opponent is resisting. In the ultra-futuristic version of this, the pilot is in a 3 DoF harness that'll spin you about, resist any movement that your robot can't make, and give you a tap when your robot suffers impact.
English
0
0
0
47
Beff (e/acc)
Beff (e/acc)@beffjezos·
We live in the cyberpunk future. This is going to be like the UFC but for techies. Watch this space
CIX 🦾@cixliv

Announcing @REK - Become a REAL Robot Pilot with our humanoid fighting sim. Top 2 players this weekend qualify for REK2: The first 6 foot tall Humanoid fight in the West on August 15 (Fighting T800s!) Thank You to 8VC and RoboStrategy for co-leading our seed round last year!

English
11
7
217
14.6K
Daniel Franke
Daniel Franke@dfranke·
@perrymetzger I've always stayed on the chronological timeline except by accident. The old algorithmic timeline was 50/50 people I follow and ragebait. Now it's more like 80/20, but still I scrolled a while without seeing any post from a non-followee that did anything but make my day worse.
English
1
0
1
57
Perry E. Metzger
Perry E. Metzger@perrymetzger·
The change in the timeline is so radical this might as well be a completely different and vastly better website. My only comment about the new timeline is that it's perhaps *too* insular now. Instead of 99% angering posts from people you don't know and 1% mutuals, we've gone to what seems like 100% mutuals. I'd like something where 90% or 95% is people I follow and 5%-10% is interesting strangers. cc: @nikitabier
English
7
2
30
926
Daniel Franke
Daniel Franke@dfranke·
@bdunbar This is my rifle. There are many like it, but this one is a piece of shit.
English
0
0
2
15
bdunbar
bdunbar@bdunbar·
"Do not accept broken shit," he told us. "If you get broken shit do this." And dramtically ripped the shit practically in half. "Go take that back to the supply NCO and get a new one."
English
2
0
2
52
Crémieux
Crémieux@cremieuxrecueil·
Vibe check: Should the punishment for an academic knowingly committing scientific fraud be death?
English
13
0
44
4K
Crémieux
Crémieux@cremieuxrecueil·
It remains really remarkable that there is NO PUNISHMENT if academics do fraud in the overwhelming majority of cases. You can get caught knowingly lying and shrug it off. Admin generally won't pursue. And if you get investigated, they might let you off to avoid embarrassment.
English
46
79
2K
61.8K
Daniel Franke
Daniel Franke@dfranke·
@paulg My stock response to this is "here is what I actually wrote [complete quote of the thread]; reread it and then answer that question for yourself".
English
0
0
3
292
Paul Graham
Paul Graham@paulg·
"So you're saying <enormously elaborate thing you didn't even suggest>?"
English
47
36
765
122.6K
Daniel Franke
Daniel Franke@dfranke·
@kendrictonn Sometimes I worry that Agent Smith was right when he told us that 1999 was the peak of our civilization.
English
0
0
3
18
Kendric Tonn
Kendric Tonn@kendrictonn·
You'll miss the American, who deals sharply but cheats no one, who is tougher than the thugs and cleverer than the tricksters, who says "I can do it" when others shrug, and who respects learning but is suspicious of those claiming to be learned, when the last one dies.
English
137
1.8K
16.7K
754.8K
Kendric Tonn
Kendric Tonn@kendrictonn·
The worst part about all of this, as someone who remembers the nineties, is the suspicion that the project of American ethnogenesis might have actually been able to succeed
Matthew Schmitz@matthewschmitz

In 2016, Mamdani’s director of appointments wrote: “It’s important that white people feel defeated.” In 2018, his housing advisor wrote: “Impoverish the *white* middle class.” Mamdani’s own platform called for raising taxes on “richer and whiter neighborhoods.”

English
49
184
3.9K
222.1K
Daniel Franke
Daniel Franke@dfranke·
@perrymetzger I was talking about automated RL. With a human in the loop you can do it, but that requires humans who themselves have the requisite taste putting serious effort into each evaluation. That's a Herculean amount of very high-skilled labor.
English
1
0
0
25
Perry E. Metzger
Perry E. Metzger@perrymetzger·
@dfranke Why can't you? "Do you prefer this code or this code?" Do that a few million times and eventually you get to start learning taste.
English
1
0
1
51
Perry E. Metzger
Perry E. Metzger@perrymetzger·
I'm working right now on some tooling for formal verification of systems software, and I keep running in to the fact that current AI models lack the taste / architectural sense needed to make really good technical decisions. They're very good at churning out thousands of lines of code in a couple of hours, but they often aren't very good at the 40,000 foot view, and will frequently make second rate or third rate decisions about design and implementation. I find this frustrating, since one would like to really have the things act as better design partners, but they often just aren't. I'm not exactly complaining, I'm working at a speed I never could have before, and the AIs are an incredible work partner that I would never want to do without, but it is still frustrating; one sees the gaps in capability very clearly when working on very complicated problems where being able to find a good solution from a huge space of possibilities is paramount. They instead just kind of grind through, which is better of course than *not* grind through, but often you have to intervene before designs turn to sludge because they'll just power through even when they should redesign something instead. They're good front line programmers, excellent at working to a spec, but not really *architects*, they don't see how you could (for example) avoid needing to write something entirely by making a different decision. For lack of a better term, they often lack "spark". If these things are going to be able to build large scale systems on their own with minimal human supervision, they will need more engineering taste and inspiration. I *am* impressed by their decisions once in a while, but really good human programmers will often impress you over and over again. On the other hand, no human programmer can work at this pace. Hopefully this gap gets closed in coming months or the next few years; it's absolutely what holds back building really big software systems that cohere. I think it's a big part of the reason that AI coded tools often end up being "clunky", that is much larger and less performant than you would like, while lacking obvious affordances.
English
20
6
62
3K
Daniel Franke
Daniel Franke@dfranke·
@perrymetzger I think models' style has gotten worse at the same time that their objective performance has gotten better. o1 often wrote broken code, but it was nicer to read and easier to fix than what GPT-5.5 produces.
English
0
0
0
26
Daniel Franke
Daniel Franke@dfranke·
@perrymetzger I think it's going to take much longer than a couple of years for this to get better. You can RL working code, but you can't RL taste.
English
2
0
0
77
Daniel Franke
Daniel Franke@dfranke·
@techzha @paulg It's probably about as energy-efficient and significantly more carbon-efficient per multiplication op than a human "computer" performing the same calculations by hand.
English
0
0
1
21
Paul Graham
Paul Graham@paulg·
Explained to 14 yo how you can multiply numbers by adding their logarithms (and divide by subtracting them), and how before calculators this was a critical hack for doing calculations.
English
116
151
7K
664.9K