Some people seem mad that Anthropic shredded a bunch of books (eg this comic, this writeup). I wonder if Anthropic could solve this by making those books available, Google Books style?
Idk what the legal complexities around copyright would imply but Google Books has solved some of those by previewing a few pages at a time. Anthropic could maybe go further and see if any of the book authors/publishers want to make their work available, offer the $3k or wtv from the other settlements, etc.
Probably Penny Arcade’s reactions aren’t actually driven by book shredding per se and more of a general unhappiness about AI. But it’d still be a nice gesture, I think.
Yeah, I’m not sure exactly how much negotiation would be involved, but given again that Google Books has paved the way for this (and also, they hired one of the Google Books people to run the project?) it might be easier the second time around, maybe relatively cheap way to buy goodwill?
(Other ways of buying goodwill on this particular topic might be to publish the full list of scanned book titles, link to whatever versions are still available for sale, make a statement that they’re supportive of books and authors etc)
I remember the book shreddings getting horrified reactions in my social circle when the story first popped up ~2 years ago, when reactions to LLMs were generally less intense. I think the general consensus was people really wished they had donated the rarer books to a library.
Thinking on this more, I’m curious how/whether copyright is considered in Plan A’s Total Research Transparency (a proposal I broadly like.) Rereading the default proposal for TRT, it calls to restrict most training data:
AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters.
I’m a bit surprised because I assumed “data” was part of “research”; most calls for “open science” include calls to publish the underlying data. And I assumed that if the goal of Plan A is to make it so that model progress is strictly gatekept on compute, in which case making training data also transparent/public would aid that goal.
And now I’m also curious what kind of training data would be published as part of TRT.
Some people seem mad that Anthropic shredded a bunch of books (eg this comic, this writeup). I wonder if Anthropic could solve this by making those books available, Google Books style?
Idk what the legal complexities around copyright would imply but Google Books has solved some of those by previewing a few pages at a time. Anthropic could maybe go further and see if any of the book authors/publishers want to make their work available, offer the $3k or wtv from the other settlements, etc.
Probably Penny Arcade’s reactions aren’t actually driven by book shredding per se and more of a general unhappiness about AI. But it’d still be a nice gesture, I think.
I believe the legal complexities are “this exact thing is explicitly banned without prohibitive amounts of negotiation”
Yeah, I’m not sure exactly how much negotiation would be involved, but given again that Google Books has paved the way for this (and also, they hired one of the Google Books people to run the project?) it might be easier the second time around, maybe relatively cheap way to buy goodwill?
(Other ways of buying goodwill on this particular topic might be to publish the full list of scanned book titles, link to whatever versions are still available for sale, make a statement that they’re supportive of books and authors etc)
I remember the book shreddings getting horrified reactions in my social circle when the story first popped up ~2 years ago, when reactions to LLMs were generally less intense. I think the general consensus was people really wished they had donated the rarer books to a library.
Thinking on this more, I’m curious how/whether copyright is considered in Plan A’s Total Research Transparency (a proposal I broadly like.) Rereading the default proposal for TRT, it calls to restrict most training data:
I’m a bit surprised because I assumed “data” was part of “research”; most calls for “open science” include calls to publish the underlying data. And I assumed that if the goal of Plan A is to make it so that model progress is strictly gatekept on compute, in which case making training data also transparent/public would aid that goal.
And now I’m also curious what kind of training data would be published as part of TRT.
Tagging @Thomas Larsen?