Microsoft executive called AI scraping 'the largest theft of labor in human history'
Unsealed litigation documents gave HN unusually candid internal language about AI training, copyright and the threat AI could pose to publishers.
Unredacted court filings in publishers’ copyright litigation disclosed internal Microsoft language describing large-scale AI training on scraped content as an extraordinary form of appropriation.
Why HN cared
The thread reopened one of AI’s hardest unresolved arguments: whether training on publicly accessible copyrighted work is comparable to human learning or fundamentally different because it happens at industrial scale and produces an infinitely replicable substitute.
Commenters also pointed to the awkward position of technology companies that both defend AI training practices and control huge repositories of other people’s code and content.
HN snapshot: roughly 950 points and more than 800 comments.
-
Copyright and Artificial Intelligence
The Copyright Office's multi-part report on AI, including its analysis of training models on copyrighted works: the question at the center of these filings.
U.S. Copyright Office · Swim · 30 min
0