a
946 posts




“Anth/OAI distilled the internet, so China’s just doing the same” is an unbelievably bad take. (1) Anthropic paid the largest settlement in history for the illegal books it downloaded — good luck getting China to do the same for mass violations of ToS (2) Training on internet data is fair-use: it’s incredibly transformative to take text and then use that to build a machine which can think. Training on copyrighted text is totally fine, so long as you buy the book. What’s illegal is *downloading those texts without paying for them*; and Anthropic now does pay for every book when it trains a model. What China does is the complete opposite: it is a direct breach to adversarially distill Fable’s outputs to another model. If you are analogizing “training” to “distillation”, you should know that *training* was never illegal; whereas distillation is a legal breach. Imagine you read a science textbook that you pirated, use that knowledge to invent with a new technique for building a new drug, and then start selling the new drug. Then, a competitor breaks into your lab, steals your technique, and uses that to sell a duplicate drug. That is theft, and it does not matter that you pirated the original textbook — you pirating the textbook means you should pay the original textbook authors, but it does not mean that your innovations are now open to be stolen by anyone! (3) Setting aside all that, the biggest issue here is national security. Anthropic spends billions in R&D and compute to push the frontier of reasoning, why should Chinese labs be able to just train on those reasoning traces to uplift their models? This is a horrible setup for China-US competition, and I am frankly baffled by how Americans could be okay with that, if they are AGI-pilled in even the slightest way.




















