Some people seem mad that Anthropic shredded a bunch of books (eg this comic, this writeup). I wonder if Anthropic could solve this by making those books available, Google Books style?
Idk what the legal complexities around copyright would imply but Google Books has solved some of those by previewing a few pages at a time. Anthropic could maybe go further and see if any of the book authors/publishers want to make their work available, offer the $3k or wtv from the other settlements, etc.
Probably Penny Arcade's reactions aren't actually driven by book shredding per se and more of a general unhappiness about AI. But it'd still be a nice gesture, I think.
I believe the legal complexities are "this exact thing is explicitly banned without prohibitive amounts of negotiation"
Yeah, I'm not sure exactly how much negotiation would be involved, but given again that Google Books has paved the way for this (and also, they hired one of the Google Books people to run the project?) it might be easier the second time around, maybe relatively cheap way to buy goodwill?
(Other ways of buying goodwill on this particular topic might be to publish the full list of scanned book titles, link to whatever versions are still available for sale, make a statement that they're supportive of books and authors etc)
I remember the book shreddings getting horrified reactions in my social circle when the story first popped up ~2 years ago, when reactions to LLMs were generally less intense. I think the general consensus was people really wished they had donated the rarer books to a library.
Yeah, donating something you don't need anymore to a library could have been a nice PR.
It's probably cheaper to scan the books destructively. Still worth considering an extra expense for PR purposes.
Thinking on this more, I'm curious how/whether copyright is considered in Plan A's Total Research Transparency (a proposal I broadly like.) Rereading the default proposal for TRT, it calls to restrict most training data:
AI model weights, significant fractions of the training data, and a small amount of other sensitive information is prevented from leaving the datacenters.
I'm a bit surprised because I assumed "data" was part of "research"; most calls for "open science" include calls to publish the underlying data. And I assumed that if the goal of Plan A is to make it so that model progress is strictly gatekept on compute, in which case making training data also transparent/public would aid that goal.
And now I'm also curious what kind of training data would be published as part of TRT.
Tagging @Thomas Larsen?
I've seen proposals for buying coal mines as a way of efficiently reducing emissions, by reducing the supply of coal and thus driving up coal's price on the open market. But how does that balance against the increasing the demand for coal mines, thus encouraging coal prospectors to seek out new coal sources?
Intuitively, this doesn't seem that likely; it feels like new coal sources should be pretty hard to discover? But two worrying examples that come to mind include:
One reason this has been on my mind: I wonder whether paying AI researchers lots of money to not do AI research would slow AI timelines, or would drive more people into research...
My sense is that coal mines 1) take a lot of money to make in the first place and 2) have poor future prospects. So the thing that happens if you buy a 40-year old coal mine with 10 years of coal left, and shut it down instead of operate it, is not that someone else just opens up a new coal mine with 50 years of life on it. [But this is probably more of a local effect than a global one--people are actually opening new coal mines somewhere.]
Suggestion: Inline comments for LessWrong posts, ala Google Docs
It's been commented on before that much intellectual work in the EA/Rat community languishes behind private Google Docs. I think one reason is just that the inline-commenting mechanism on a GDoc is so much better than excerpting the comment below. Has the Lightcone team considered this/what is the status?
(I vaguely recall them working on a live-collab feature, not sure if commenting would have been part of this)
If you go into a draft of a post and highlite text you can see a comment button. When I just tested it, clicking the button did nothing but at least the button to create comments is there.
Okay, now I've used the live-collab/commenting feature on a LessWrong draft. It's pretty good! If you haven't seen it yet, I'd recommend writing a new LW post and requesting feedback; Justis Millis's feedback was super fast, highly detailed, and all-around incredibly valuable!
Can I turn on inline comments for a published LessWrong post too? Even after "publishing" it'd super useful to get the comments inline. In my view, a great post should be a timeless, living, breathing, collaborative document, rather than a bunch of words dumped out once and never revisited.
(There's value in the latter in terms for eg news posts; but LW's focus is less on that.)