this post was submitted on 23 Aug 2026
436 points (99.8% liked)

Piracy: ꜱᴀɪʟ ᴛʜᴇ ʜɪɢʜ ꜱᴇᴀꜱ

70401 readers
795 users here now

⚓ Dedicated to the discussion of digital piracy, including ethical problems and legal advancements.

Rules • Full Version

1. Posts must be related to the discussion of digital piracy

2. Don't request invites, trade, sell, or self-promote

3. Don't request or link to specific pirated titles, including DMs

4. Don't submit low-quality posts, be entitled, or harass others



Loot, Pillage, & Plunder

We heartily recommend visiting the free port of freemediaheckyeah (aka FMHY) while you sail the high seas, for all the freshest links the ocean has to offer.

📜 c/Piracy Wiki (Community Edition):

🏴‍☠️ Other communities

FUCK ADOBE!

Torrenting/P2P:

Gaming:


💰 Please help cover server costs.

Ko-Fi Liberapay
Ko-fi Liberapay

founded 3 years ago
MODERATORS
 

As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna's Archive is calling for volunteers to help preserve them for the public record.

all 44 comments
sorted by: hot top controversial new old
[–] yakko@feddit.uk 66 points 1 week ago

This feels like a "moment where good people took action in history books".

[–] fluffykittycat@slrpnk.net 54 points 1 week ago

Anna's is the real deal

[–] Catoblepas@lemmy.blahaj.zone 36 points 1 week ago (4 children)

I have access to a book scanner, but I’m not sure what books need to be scanned, are legal to be scanned and uploaded (I don’t want to get expelled for using a school scanner to do something illegal), or whether there is personally identifiable data in the PDFs it generates (I think I can also save in GIF and maybe PNG?). Is there a beginner’s guide anywhere for all this?

[–] Truscape@lemmy.blahaj.zone 41 points 1 week ago (1 children)

Anna's Archive does have guides for all of that (apart from the legal advice) on their site (use wikipedia to find the right domain).

I don't think you'd want to do so as a university student unless you want to share the same fate as Aaron Swartz though.

[–] Catoblepas@lemmy.blahaj.zone 20 points 1 week ago (2 children)

Surely at least the stuff out of copyright is kosher to scan and share? Which seems like a good area to focus on if they’re being bought up and destroyed

[–] dan@upvote.au 29 points 1 week ago* (last edited 1 week ago)

Stuff that's out of copyright isn't an issue though.

The reason AI companies are destroying books is because current US copyright caselaw (most recently the Anthropic lawsuit) doesn't allow books to be copied, but considers it fair use if you transform the format of a legally purchased book (like from print to digital) without creating a new copy. By destroying the original book, they haven't made a copy of it, and so they operate within US law.

This doesn't apply to public domain books (books that are no longer copyrighted) as you can do whatever you want with public domain content.

[–] obbeel@kbin.earth 6 points 1 week ago

Maybe there is a way to do this anonymously? Anna's Archive deals with books new and old, but maybe that 'physical-only' book that has no digital print would be the best place to start?

I remember no-eyes from #ebooks (IRCHighWay) prioritized recommendations of novels (fiction) to expand their available books, so I think that's the safest type of book to scan?

[–] Derpenheim@lemmy.zip 22 points 1 week ago (1 children)

Im gonna be real chief. Stop worrying about IP legality. Scan and upload everything you can get your hands on, wherever you can reasonably upload it.

AI companies are violating all copyright laws at all times of day and then destroying the originals, or even DMCA claiming them. You cannot fight that and be worried about legality.

[–] EggInDisguise@lemmy.blahaj.zone 11 points 1 week ago

I have repurposed an old Xbox external drive and have been ripping every optical disc in my extended family's houses as well as opened it up to friends, anyone who wants to add to the collection is welcome to take from it, all I have to do is bring the drive with me, and they can grab whatever they want off it.

Fuck IP laws, fuck trademarks, fuck copyright.

If companies get to arbitrarily decide we can't have access to things (especially if you have given someone money for it), then why should we give a fuck what they think?

[–] fluffykittycat@slrpnk.net 11 points 1 week ago

Try your beat to cover your tracks while you do it, but if you can scan some rare books, please do. You'd be a hero

[–] A404@lemmy.dbzer0.com 1 points 1 week ago (1 children)

if you use the tor browser + public wifi they shouldnt be able to track you back. just make sure not to log into anything you accessed without tor

[–] A404@lemmy.dbzer0.com 1 points 1 week ago

if you have any questions regarding OPSEC feel free to ask

[–] Aceticon@lemmy.dbzer0.com 12 points 1 week ago

Non-profit Copyright Infringement (a.k.a. Piracy) is a Moral Duty, IMHO.

[–] Tim_Bisley@piefed.social 11 points 1 week ago (1 children)

God the future is depressing

[–] arran4@aussie.zone 6 points 1 week ago (1 children)

It doesn't have to be, just structurally people are incentivized to do this

[–] mindbleach@sh.itjust.works 2 points 1 week ago

And people would rather demand change just because they say so, versus changing any of the incentives involved.

We could just make these companies share their scans. Then it would only affect one (1) copy of every published work. There are no ongoing commercial concerns for a book that's been out-of-print for thirty years.

[–] Sibshops@feddit.cl 7 points 1 week ago (2 children)

AI companys would probably like this. Easier for them to get data.

[–] SnoringEarthworm@piefed.ca 26 points 1 week ago (1 children)

Nothing will stop corpos from getting data.

But we can prevent them from paywalling access by having archives of our own.

[–] EggInDisguise@lemmy.blahaj.zone 8 points 1 week ago (1 children)

Hmm, "archive of our own" sounds like a great name for something.

[–] SnoringEarthworm@piefed.ca 5 points 1 week ago (1 children)

Eh, it's a bit long.

We should make it shorter, like with an acronym or something.

[–] Zoop@beehaw.org 4 points 1 week ago

Well, clearly AOOO is the right choice here.

[–] Gerudo@lemmy.zip 5 points 1 week ago

Even if they do just end up using AA to build their data, at least WE get to also have a copy instead of it being destroyed.

[–] reallykindasorta@slrpnk.net 6 points 1 week ago (2 children)

We could probably recruit the special collections departments at university libraries to help since they’re already equipped, but how do we identify and procure what needs to be preserved? I kind of assumed most books were already being digitized. As a student I worked on a project digitizing old newspapers and entering basic metadata.

[–] fluffykittycat@slrpnk.net 7 points 1 week ago

I sometimes try and track down rare books from the 20th century, sometiems less than 75 years old, only to.run into the problem that the only copy is in a rare books collection in a Canadian university or some shit like that

[–] alapakala@quokk.au 1 points 1 week ago

you misunderstand, Universities are for book burnings. How do you think they stay afloat?

[–] SabinStargem@lemmy.today 5 points 1 week ago (2 children)

Setting aside the bad ethics of destroying rare books, the capitalist in me is very confused: Wouldn't you put in some effort to save the rare books, if only to sell them?

If you already have the data inside the books, you can put them up for auction, donating to libraries, and so forth. You get money, a reputation for ensuring that they find a home, have a physical backup if the storage drives fail, and so on.

...my inner capitalist is disillusioned. 🤨

[–] Longylonglong@sh.itjust.works 3 points 1 week ago

Not if the endgame is to control all Information and, if beneficial for you, to modify sources with no way to prove something different

[–] FjordDan@lemmy.zip 3 points 1 week ago (1 children)

In some interpretation of copyright law, it's considered a "transfer" rather than a "copy" if you destroy the physical original. (Ignore all subsequent copying of the digital version.)

Actually old and rare books are probably public domain, so that's not an argument.

Also, most of the books they scan/destroy are low quality trashy novels that won't even sell at thrift stores. But as they are written by humans it is still useful for training models.

[–] Aceticon@lemmy.dbzer0.com 2 points 1 week ago* (last edited 1 week ago)

Copyright nowadays is up from the original "25 years" to around "Death of author + 50 years" (a bit more in the US) which is almost always more than 100 years in total, so those "old" books are nowaday almost a century old or more.

(Well, sorta, if their copyright expired BEFORE new legislation was introduced to extend copyright length as has been regularly done for almost a century, those works didn't got back under copyright after the extension, and as at least in the US copyright extension legislation was seeming driven by wanting to keep the first Mickey Mouse cartoon under copyright, at least in the US said "old" out of copyright books will mainly be from before 1928).

[–] ApathyTree@lemmy.dbzer0.com 4 points 1 week ago* (last edited 1 week ago)

The only book I have that might be worth scanning is Knight’s Modern Seamanship 13th edition from ~1960 (was provided to my mom as part of a record-finding request regarding my grandpa’s WW2 deployments, I have no idea why), but if my time in the Navy is any indication, its one of those things that was given during basic to all going through basic (I have a modern copy from 2007 as well), so idk if its a valuable contribution. Probably already in the archive.

I also have a copy of the joy of cooking from the 70s, and a Betty Crocker cookbook from the 50s, but I assume that’s also already in there. Plus some language textbooks and the like, some bird and geology books, some limnology books, etc. the sort of thing that’s definitely already there.

I’m willing to scan them, however a person does that, but I’m not willing to destroy them. The old books came from my grandparents, through my mom, and all of those people are long dead. Everything else is from my own education and I use them for reference.

[–] EggInDisguise@lemmy.blahaj.zone 3 points 1 week ago (1 children)

I wish I had a book scanner, I'd be digitizing every book on my shelf. I already have a good chunk of my forgotten realms stuff digitally (don't ask where I got it) but there's still a ton more I need, and I've finally gotten my spouse to read and their collection is getting almost as big as my own. There's probably over 1,000 books between all my immediate family members.

[–] 01189998819991197253 2 points 1 week ago (1 children)

How does one find the official Anna's? I found a few that were shady af and were very obviously not real. Tor and i2p addresses are fine, too.

[–] WhyJiffie@sh.itjust.works 6 points 1 week ago

their wikipedia article should have the current domains, I was told before here

https://en.wikipedia.org/wiki/Anna%27s_Archive

[–] alapakala@quokk.au 2 points 1 week ago (1 children)
[–] not@lemmy.dbzer0.com 5 points 1 week ago

The best time to start was more than a year ago, the next best time to start is today.

[–] TerdFerguson@lemmy.ca 1 points 1 week ago* (last edited 1 week ago)

Gotta feed the machine.

I mean, these companies are speed-running atrocity after atrocity toward the end of humanity. There is no way this AI bubble ends without mass graves.. not if we let them continue to decide how things go.

[–] Mordikan@kbin.earth -1 points 1 week ago

There is a glaringly obvious issue with this article that I see a lot of people in the comments seem to have missed:

Only one copy of any given book exists The pretense to this entire thing is that AI companies are destroying the books both because of the process used AND to prevent competitors from accessing the book. There is more than ONE copy of that book. If you buy a book, that doesn't prevent me from also buying that book. You can not like AI for a number of reasons, but this is just stupid.

NOTE: If destroying the book actually prevented competitors from accessing it, that means that all of these AI companies are only going out on the weekends to used bookstores looking for good deals on second hand books.