Absolutely. Every AI company has Anna's Archive in their training data already. They want to differentiate their models by having more training data than their competitors. There are way more books out there that have never been scanned than those that have.
That is the justification and not the reason. Its faster and cheaper to cut the bindings, and they would have to store the books while they are using the digital copies to comply with copyright.
If they were archiving digital versions so they could be read by thousands in the future, this would be a totally different conversation.
They don't even need to be as selfless as the internet archive - just follow Googles lead and allow people to search and view selections the archive, with automatic opening of the archive when you are confident copyright is expired. Even a closed archive owned by a third party with a dead mans switch to open it in the future when the business model is done would be something.
When destroying books, we don't actually know what will be important in the future. That's why archivists spend so much time archiving mineutia without trying to curate. Those are different jobs and you can't curate for the future.
Keep in mind that books often become culturally important decades after their print run, sometimes when the copyright is ambiguous.
And also keep in mind there are technically, an insane number of "rare" books out there, but important rare books might only be important to people without the money to pay for another printing.
I could almost forgive this prrocess if they were /also/ archiving these books (leven if limited ike Google books does), but they won't, because they are so focussed on staying ahead of compeditors based on the "knowledge" in their training data.
They are buying up used books at rates that push up the cost of used books. That is enough to piss off people legitimately.
We also are fully aware they are not careful. Yes, they will occasionally shred the last few examples of a book that might have otherwise become useful in the future, and there's no indication that they are archivists. They will eventually destroy or loose this training data and the book contents will be gone to future generations.
You have failed to understand the difference between idealism and practice. There are more companies that use Linux without ever contributing back than those that do. Those that do do so because it makes practical business sense to improve the upstream project they rely on. Any company that has decided they don't want to share some code just builds it in a binary module, firmware, or application to sit on top.
There's already strong precedence that destroying data after learning of an investigation is a crime. Doesn't matter to the courts that is via a clever trigger. If he knowingly provided a pin to an official that destroys the data then he's hosed. His only hope is if they guessed on their own.
Maybe the Canadian police, but the Canadian court did not. They convicted him with zero corroborating evidence and without carefully comparing the chat logs with his actual account.
Another way to phrase that is that companies can choose to meaningulfully contribute back, or not, with either project, and there are plenty of examples of each combination.
Another way to think of it is the old "you wouldn't steal a car" commercial.
A company using BSD in their product cannot "steal" it, because we still have the BSDs. They can't take them away, so its not stealing. It just comes down to your personality and how much you value control over what they are doing with it
(No shade either way from me, if its your work)
In the end, I think it matters less than it seems. Companies can use Linux via either tivoization or proprietary software on top without contributing back,, and other companies voluntarily submit patches or pay FreeBSD developers because maintaining a closed source fork is expensive.
This is the correct take. Its common across various business contexts: destroying evidence after you learn you are being investigated is big time illegal. If you destroy data, you better be able to demonstrate you did so beforehand (I.e. an expiration policy), or you don't have it in the first place (because sensitive data doesn't touch a given mobile device and all.). Dedicated device for travelling is the best idea.
You could also take a video of yourself wiping the device before travelling for security in case it gets stolen. (Not a lawyer disclaimer).
Really a long time coming pain point is comfortably sharing between family. I think the new workflows will be the way eventually. I plan on using it to auto add pictures with the recognised faces of family members to a "Family Photos" album as they are uploaded, and allow everything else to stay unshared by default for each user.
Happy with them too. Standards compliant (use a real mail client). Secondary accounts are ultra cheap if you add family and such.
The "controversy" is probably the intended marketing campaign.
Absolutely. Every AI company has Anna's Archive in their training data already. They want to differentiate their models by having more training data than their competitors. There are way more books out there that have never been scanned than those that have.
That is the justification and not the reason. Its faster and cheaper to cut the bindings, and they would have to store the books while they are using the digital copies to comply with copyright.
If they were archiving digital versions so they could be read by thousands in the future, this would be a totally different conversation.
They don't even need to be as selfless as the internet archive - just follow Googles lead and allow people to search and view selections the archive, with automatic opening of the archive when you are confident copyright is expired. Even a closed archive owned by a third party with a dead mans switch to open it in the future when the business model is done would be something.
When destroying books, we don't actually know what will be important in the future. That's why archivists spend so much time archiving mineutia without trying to curate. Those are different jobs and you can't curate for the future.
Yes... If they were archiving their scans. Which isn't happening.
Keep in mind that books often become culturally important decades after their print run, sometimes when the copyright is ambiguous.
And also keep in mind there are technically, an insane number of "rare" books out there, but important rare books might only be important to people without the money to pay for another printing.
I could almost forgive this prrocess if they were /also/ archiving these books (leven if limited ike Google books does), but they won't, because they are so focussed on staying ahead of compeditors based on the "knowledge" in their training data.
That is not required by law. It is just easier to cut the bindings off. The other way is slow: https://archive.org/details/eliza-digitizing-book_202107
The Internet Archive absolutely does not destroy books. They scan them the hard way with the binding still intact.
Easier to scan if you cut the binding off, and it costs money to re-bind a book you purchased for 10 cents as part of a lot.
They are buying up used books at rates that push up the cost of used books. That is enough to piss off people legitimately.
We also are fully aware they are not careful. Yes, they will occasionally shred the last few examples of a book that might have otherwise become useful in the future, and there's no indication that they are archivists. They will eventually destroy or loose this training data and the book contents will be gone to future generations.
You have failed to understand the difference between idealism and practice. There are more companies that use Linux without ever contributing back than those that do. Those that do do so because it makes practical business sense to improve the upstream project they rely on. Any company that has decided they don't want to share some code just builds it in a binary module, firmware, or application to sit on top.
There's already strong precedence that destroying data after learning of an investigation is a crime. Doesn't matter to the courts that is via a clever trigger. If he knowingly provided a pin to an official that destroys the data then he's hosed. His only hope is if they guessed on their own.
Maybe the Canadian police, but the Canadian court did not. They convicted him with zero corroborating evidence and without carefully comparing the chat logs with his actual account.
Another way to phrase that is that companies can choose to meaningulfully contribute back, or not, with either project, and there are plenty of examples of each combination.
Of course they are. Sony, Netflix and the like are major contributors to the project.
Another way to think of it is the old "you wouldn't steal a car" commercial.
A company using BSD in their product cannot "steal" it, because we still have the BSDs. They can't take them away, so its not stealing. It just comes down to your personality and how much you value control over what they are doing with it (No shade either way from me, if its your work)
In the end, I think it matters less than it seems. Companies can use Linux via either tivoization or proprietary software on top without contributing back,, and other companies voluntarily submit patches or pay FreeBSD developers because maintaining a closed source fork is expensive.
This is the correct take. Its common across various business contexts: destroying evidence after you learn you are being investigated is big time illegal. If you destroy data, you better be able to demonstrate you did so beforehand (I.e. an expiration policy), or you don't have it in the first place (because sensitive data doesn't touch a given mobile device and all.). Dedicated device for travelling is the best idea.
You could also take a video of yourself wiping the device before travelling for security in case it gets stolen. (Not a lawyer disclaimer).
Really a long time coming pain point is comfortably sharing between family. I think the new workflows will be the way eventually. I plan on using it to auto add pictures with the recognised faces of family members to a "Family Photos" album as they are uploaded, and allow everything else to stay unshared by default for each user.