Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.”
Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI.
It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.
- FoxAlive@lemmy.zipEnglish16 days
I honestly don’t see a point where we can move past this without Sam altman, nadell, musk, etc all facing mandatory life time sentence without parole, work release etc.
I would also accept the removal of their heads.
Alaknár@sopuli.xyzEnglish
15 daysI’d prefer if they were forced to pay royalties for all the work they stole to the people they stole it from. And, like, have someone actually force them to comply. They’d have to hire a Microsoft-sized compliance department just to figure this shit out and track payments.
- JustPlainDave@lemmy.zipEnglish16 days
Better yet, seize their assets and make them try to make a go at it as a regular person.
- JustPlainDave@lemmy.zipEnglish16 days
Could you imagine Elon Musk running a fryer or injection mold for 8 hours? Who am I kidding? he’d never pass the piss test required to run the injection mold.
Mrkawfee@lemmy.worldEnglish
16 daysInternet search is terrible now. Every website I go to reads like it was generated by an LLM.
Doom@discuss.onlineEnglish
16 daysGlad I spent my teens and 20s getting stoned and reading the shit out of wikipedia before this shit happened. Can’t trust anything written online now.
- CosmoNova@lemmy.worldEnglish16 days
They knew what they were doing every step of the way. They are criminals that need to be disarmed and locked away. And we need to create a new Internet from scratch somehow thanks to these donkeys.
- affiliate@lemmy.worldEnglish16 days
If we make a new internet from scratch can we get rid of JavaScript too while we’re at it?
- HaraldvonBlauzahn@feddit.orgEnglish16 days
Here you go:
https://en.wikipedia.org/wiki/Gemini_(protocol)
https://github.com/kr1sp1n/awesome-gemini
(This protocol does not have to do anything with Google’s Gemini thing. It was created and named a bit earlier!)
- Arancello@aussie.zoneEnglish16 days
Relax guys, the very stable high IQ president of the united states will protect you. No meed for guardrails or regulations.
AmyAye@nord.pubEnglish
16 daysGod the guardrails things. These stupid companies are all hyping up “We need to slow down, we need guard rails!”
Ok.
No one is fucking stopping you. Just… Slow yourself down, guardrail yourself.
Oh wait, it’s just an excuse to create regulatory capture.
- 16 days
Microsoft, in a policy document, wrote that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained […] LLMs are a product that destroys its supply chain.”
If that’s not the very definition of a categorically unsustainable business model, I don’t know what the fuck is.
velma@sh.itjust.worksEnglish
16 days“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said.
Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing.
Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.”
Fuck these guys.
- tangeli@piefed.socialEnglish16 days
That’s because, thus far, they get away with choosing not to distribute any of their trillions of dollars to the suppliers of the information they consume - money has only gone to the suppliers of hardware and power, and to influencing politicians and rewarding investors. That’s their choice, and they should not be allowed to continue to make that choice. Good luck to the NYT.
- Grandwolf319@sh.itjust.worksEnglish16 days
But would it be worth it if they had to pay the creative costs for it?
They are only worth it now cause it doesn’t include that AND it’s subsidized by investor money.
- Rothe@piefed.socialEnglish16 days
From an economic viewpoint AI companies are not worth is as it is now. OpenAI and Anthropic are hundreds of billions of dollars in debt and will never make a profit. Same goes for the AI subsidiaries of Microsoft, Meta, Google etc, except they are just leeching off of the mother companies and hiding their figures among their finances. Paying creators for their data would make very little difference in their overall finances.
- kestrel7_7@lemmy.worldEnglish16 days
This is why I’ve been arguing to anyone who will listen for like four years now. If this tech is so amazing, someone will figure out a way to make money off of it. Right? The fact that it’s been 4+ years and no one has made a damn cent off of it should be making more people skeptical. The fact that 4 years ago it was widely celebrated even though no one had a plan to make a damn cent off of it should have made more people skeptical back then. It’s weird that I have to keep arguing this with folks.
- kablez@lemmy.worldEnglish16 days
Gonna get real weird soon when they run out of rich new training data and they begin to consume their own shit. When that happens their entire model will collapse and if the bubble hasn’t popped already that may be what causes it.
MalReynolds@slrpnk.netEnglish
16 daysPretty sure they mostly use the pre-AI internet (that they scraped and kept) and synthetic data currently. Probably trying (and failing so far or we’d have heard) to adapt to using video as training material at the moment, but developments there will likely apply to robotics at some point. Here’s hoping the current chuds have crashed and burned before then and that some sanity has taken over from unfettered capitalist oligarchs dreams of computer slavery.
- 16 days
Why don’t you thinj they have not adapted to video as training. Is that not what Flock does?
MalReynolds@slrpnk.netEnglish
16 daysFlock does pattern recognition, a quite old piece of machine learning. Nothing to do with training a large language model or other ‘AI’ model.
- 12 days
They almost certainly use that to train AI. Big tech is aggregating the data from many places.
MalReynolds@slrpnk.netEnglish
12 daysNah, it’s in the name, Large Language Models, what gets marketed as ‘AI’, are trained on text. That’s why they’re ripping up secondhand books (and copyright law, but that never stopped them) at the moment. No one has really cracked Large Vision Model yet, although there are some primitive versions being attempted in robotics, mostly they translate to text, which is why general purpose robots aren’t a thing yet.
- 8 days
I think we’re talking two different things. Like, for example, they can create data using ALPRs to create a history of a person’s location to train LLMs.
- nao@sh.itjust.worksEnglish16 days
a “doom loop” that is killing “the entire web.”
the corporate web maybe. how does it affect independent parts like the fediverse?
- 16 days
Something others haven’t touched on is the burden all the scrapers are putting on small websites and forums.
Lots of niche/hobbyist site runners have repeatedly talked about how the scrapers inefficiently scrape the same pages, images, and files endlessly. Until they crash the site, run them out of bandwidth, or run up their bills till they cant afford to host a site they’ve been running for many years before this nonsense.
All in the futile pursuit of more to dump into the machine because of the belief if they just had more data all the issues with the models would finally be solved.







