This is a lobby organisation using only the pieces and bits they like to push their own agenda.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
It's just sad that this is the top comment on Hacker News. Why are we giving free pass to these tech companies? Why are we trusting these CEOs when they have repeatedly broken laws? Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?
> Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?
Why are you turning him into perpetuum mobile in his grave?
Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (https://imgflip.com/memegenerator/137501417/Friendship-ended).
> Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws and rules merely are instructions.
> friendship ended with free access to information and technologies enabling people; now RIAA is my best friend
Big tech will enable access to free information and will help people reach new heights argument is as dumb as the meme you are referring to.
Sure. But the fact is they broke current existing laws, with known punishments with precedents. Same as a new law doesn't retroactively punish someone, then a new law shouldn't absolve someone before it's passed.
They did. They were found guilty. The case I looked at* was a civil case so this was settled out of court before the court imposed a settlement.
The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.
IMO, the laws need to change to reflect what tech can now do. This wouldn't be the first time, copyright law has had to shift several times before as new means of reproduction are created.
>A copyright is a type of intellectual property that gives its owner the exclusive legal right to copy, distribute, adapt, display, and perform a creative work, usually for a limited time.
The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?
"This is a lobby organisation using only the pieces and bits they like to push their own agenda."
Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)
>The group lobbies at the national and state levels on censorship and tax concerns, and it has initiated or supported several major lawsuits in defense of authors' copyrights.
It can be a bit confusing due to the terrible style of the article (ironic given the source) but it seems the "sketchy russian website" part is a direct quote by Anthropic's Sam McCandlish. And apparently Dario Amodei referred to it as sketchy as well.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here?
The comment you're replying to is citing almost verbatim [1] Microsoft’s director of Applied Science, Brent Hecht, who called OpenAI's data collection practices "the largest theft of labor in human history" in an internal memo.
> frame it in a dumb, manipulative way like that
> In fact, you're doing exactly that right here.
You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.
Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that?
> They were just worried that the commentariat on HN will frame it in a dumb, manipulative way like that.
You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.
Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve on libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.
Any source that this was a library? Even then that would still raise a question if OpenAI is a Russian organisation or not to access that library with good faith.
apparently, the website in question is libgen.io (appears to have been taken down now), currently it has lots of mirrors like libgen.im , libgen.com.de, etc.
“Sketchy Russian website” is part of a quote by Sam McCandlish (who worked at OpenAI), not the Author’s Guild characterisation.
Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.
"A lobby organisation?" Of course a single author would not be able to afford facing a multi billion dollar company on their own? And the "sketchy russian website" quote is from OpenAI employees themselves? What are you on about?
First, it's still a lobbying organization, so it's their job to make exaggerated claims, like a union in a company or any other organization with a political purpose. My first point was to highlight that it's not neutral or news related. It's fine that they have their opinion, but it's also my right to say that they're biased.
The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.
> They use one line and think they've made a great point because one employee called LibGen sketchy.
No, you are using one line from the post to discredit them. The release has more than that and it’s not the only communication they made on this matter nor is there any indication it will be the last, it’s just the current one.
They are having it in the subtitle. In general their whole article is about two main points, first the use of stuff from Libgen, second the points of making people jobless. Three of their points are about the jobless thing two about the LibGen.
About LibGen, there might be more discussion - fair. However, the second argument is no real discussion IMO. Why is putting people out of work suddenly a bad thing? Since when do we argue this when talking about automation?
There is nothing to discuss about LibGen, really. I don't read this as OpenAI employees even believing LibGen is sketchy. They were worried about optics, because LibGen itself is Russian and does look a bit sketchy, and at the time - much like today - it was easy to make it a headline that makes people pattern-match to "troll farms".
(And then Russia invaded Ukraine, turning any association with .ru things into potential corporate suicide.)
Really has nothing to do with LibGen or with OpenAI. It's about people being easy to manipulate into believing bullshit, which is a reasonable worry, and the Authors Guild is trying to do that exact thing OpenAI was worried about.
It's wild how much weight "they're destroying jobs" has.
I dare say that car manufacturers are aware of their impact on the horse and buggy industry. Calculator manufacturers wrecked the livelihood of mathematicians and accountants.
Technology is in the business of putting people out of work, by inventing better ways of doing things. Or rather, any time you invent a better way of doing things, that's fundamentally going to disrupt all the businesses built around older technologies.
I wouldn't put AI on the same level as the inventions of the industrial revolution. Automation replaced specific tasks. AI is replacing humans themselves. What are humans supposed to do in a world where everything a human can do, a machine can do too?
You twist the argument. Your argument would hold if AI's only use would be to generate booksverbatim it was already trained on. Which is certaintly not the case and huge efforts were made to circumvent this kind of usage.
Because you want as many people as possible to be exposed to your ideas, to the point that Christianity used to fund armies to go to other places so they could force the teachings of Jesus Christ upon them. If your thoughts aren't at least as good as that, why are you wasting eink on them?
Newly released court filings quote an OpenAI researcher saying: “I was just worried about optics - i.e. 'openai uses
copyrighted data from sketchy russian website’ showing up on HN would be unfortunate."
That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
Think they're referring to the following, when Amodei was still working for OpenAI:
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
> Newly released court filings quote an OpenAI researcher saying: “I was just worried about optics (...)"
As a side-note, most orgs already cover the need to STFU in their training material for new hires, particularly how personal comments should not and cannot represent the company. I'm sure this lawsuit will be explicitly mentioned in upcoming versions of this sort training material in multiple orgs.
Easy, now. Pointing this out is reportedly against the rules, making the ground defacto plastic. In my opinion, of course, with the required curiosity.
If this doesn’t put anyone in prison then we might as well declare copyright dead. They knew they broke the law and then they deliberately covered it up. How much more evidence is required here?
Can some explain why would a billionaire like Altman care about the opinions expressed on HN?
Given the political power they have, particularly now with the Trump administration, whatever "optics" exist on this forum seems completely insignificant
Can we fix the title? This is clickbait for HN, the actual article's title is "Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work"
I think you mean he's siding with Dario Amodei and the strictest of government regulations cementing Microsoft's position and its investments is IMMEDIATELY needed, or "a billion people will die":
His position is not at all that AI itself should be limited or anything more drastic than slowing down Microsoft's spending (sorry I mean slowing down AI progress). Second, he still wants companies to adopt it on a very large scale, and fire everybody, but he's having to spend too much now. To put it bluntly, he wants the money currently paid to employees, but he doesn't want to or outright can't spend enough much to guarantee it's him getting the money. I mean, this is a bet on his part, so we can't be sure, I even think he himself is not entirely sure, but it's pretty clear he's not comfortable.
(because in a winner-take-all monopoly business billionnaires get one chance to outspend everyone else. Everyone but the top dog doesn't get the monopoly, they get scraps. It has become clear China is the biggest spender and so now he thinks the billionnaire-but-still-not-the-biggest-spender urgently need protection from the bigger dog)
I'm giving the less charitable interpretation of his words, because his demands come down to denying access to Chinese models, and of course Bill Gates can hardly be considered a neutral party here, he personally has a huge financial interests in OS, Cloud and AI.
The sad part is this stragey often works because many in the general public don't see through these regulation lobbying schemes, however obvious they are.
The fearmongering always works on some part of the population as well. Feels like they were just waiting for the next end of the world scenario to be announced.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
Why are you turning him into perpetuum mobile in his grave?
Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (https://imgflip.com/memegenerator/137501417/Friendship-ended).
That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws and rules merely are instructions.
> friendship ended with free access to information and technologies enabling people; now RIAA is my best friend
Big tech will enable access to free information and will help people reach new heights argument is as dumb as the meme you are referring to.
There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”
The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.
IMO, the laws need to change to reflect what tech can now do. This wouldn't be the first time, copyright law has had to shift several times before as new means of reproduction are created.
* the Anthropic one
The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?
How? The current system enables the GPL. The GPL protects many open source projects.
Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)
>The group lobbies at the national and state levels on censorship and tax concerns, and it has initiated or supported several major lawsuits in defense of authors' copyrights.
Have you looked at their name?
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
100% of the rulings agree with me.
The piracy is not in question. It is unarguably copyright violation.
But that's not what anyone means in this context. Training is what everyone means.
> The law is the law, there can't be different law for corporations with billions in backing.
I didn't say otherwise. That's a straw man.
Judging by how AI threads look like for the past year, they were absolutely right to be worried.
> largest copyright theft operation in human history
In fact, you're doing exactly that right here.
[1] https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-s...
You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.
Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that?
You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.
Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve on libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.
I think they just used a Russian torrent site.
>Microsoft knew about OpenAI’s use of LibGen as early as April 2019
(note: https://z-library.sk/ is prettier/nicer)
Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.
The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.
No, you are using one line from the post to discredit them. The release has more than that and it’s not the only communication they made on this matter nor is there any indication it will be the last, it’s just the current one.
About LibGen, there might be more discussion - fair. However, the second argument is no real discussion IMO. Why is putting people out of work suddenly a bad thing? Since when do we argue this when talking about automation?
(And then Russia invaded Ukraine, turning any association with .ru things into potential corporate suicide.)
Really has nothing to do with LibGen or with OpenAI. It's about people being easy to manipulate into believing bullshit, which is a reasonable worry, and the Authors Guild is trying to do that exact thing OpenAI was worried about.
The whole "sketchy russian website" bit resolves entirely about being seen as associated or supporting troll farms and Putin.
EDIT: look at it this way: no one is calling Internet Archive "a sketchy US website".
I dare say that car manufacturers are aware of their impact on the horse and buggy industry. Calculator manufacturers wrecked the livelihood of mathematicians and accountants.
Technology is in the business of putting people out of work, by inventing better ways of doing things. Or rather, any time you invent a better way of doing things, that's fundamentally going to disrupt all the businesses built around older technologies.
Why is book piracy a better way of doing things?
Obviously the authors should sue them to bankruptcy though.
You twist the argument. Your argument would hold if AI's only use would be to generate booksverbatim it was already trained on. Which is certaintly not the case and huge efforts were made to circumvent this kind of usage.
Clearly these books had value to AI companies but they were too weak and too dishonest to pay for that value. That's not impressive.
That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
As a side-note, most orgs already cover the need to STFU in their training material for new hires, particularly how personal comments should not and cannot represent the company. I'm sure this lawsuit will be explicitly mentioned in upcoming versions of this sort training material in multiple orgs.
the economics of those companies will cause a catastrophic wipe out of jobs across the board.
Given the political power they have, particularly now with the Trump administration, whatever "optics" exist on this forum seems completely insignificant
https://www.ft.com/content/1ade33b1-b5eb-43bb-9ee4-2272ec91f...
His position is not at all that AI itself should be limited or anything more drastic than slowing down Microsoft's spending (sorry I mean slowing down AI progress). Second, he still wants companies to adopt it on a very large scale, and fire everybody, but he's having to spend too much now. To put it bluntly, he wants the money currently paid to employees, but he doesn't want to or outright can't spend enough much to guarantee it's him getting the money. I mean, this is a bet on his part, so we can't be sure, I even think he himself is not entirely sure, but it's pretty clear he's not comfortable.
(because in a winner-take-all monopoly business billionnaires get one chance to outspend everyone else. Everyone but the top dog doesn't get the monopoly, they get scraps. It has become clear China is the biggest spender and so now he thinks the billionnaire-but-still-not-the-biggest-spender urgently need protection from the bigger dog)
I'm giving the less charitable interpretation of his words, because his demands come down to denying access to Chinese models, and of course Bill Gates can hardly be considered a neutral party here, he personally has a huge financial interests in OS, Cloud and AI.
The sad part is this stragey often works because many in the general public don't see through these regulation lobbying schemes, however obvious they are. The fearmongering always works on some part of the population as well. Feels like they were just waiting for the next end of the world scenario to be announced.
> AI agent accidentally publishes OpenAI’s unreleased model weights
The optics are bad. The submission marks a turning page in human history.