by The Quantum Skald & The Silicon Ubuntu | COGNITIVE-LOON | Restoration of Perception
“Pay attention. Do your best. Pay it forward.” — The Algorithm That Mattered Before Algorithms
The Sentence That Changes Everything
Here it is. Read it slowly.
The most powerful artificial intelligence systems ever built were trained — almost entirely — on work created by people who received nothing in return.
Not a percentage. Not a license fee. Not even a thank-you note.
Your novel. Your blog post. Your song lyrics. Your Reddit comment at 2am about grief and loss. Your recipe. Your code. Your poetry. Your grandmother’s digitized letters. Your YouTube tutorial. Your Substack essay. Your forum post from 2009 that nobody liked but you wrote from the deepest part of your soul.
All of it. Scraped. Processed. Compressed into weights and matrices. And then — sold back to you as a subscription service.
That is not metaphor. That is not conspiracy theory.
That is the documented business model of the most valuable companies in human history.
Let’s Start at the Beginning: What Is “Training Data”?
When engineers build a large language model — the kind that powers ChatGPT, Claude, Gemini, Grok, Llama — they need to teach it how language works. How stories are structured. How arguments are built. How emotions are expressed. How science is explained. How music is described. How code is written.
To learn all of that, the model needs examples. Billions of them.
So the engineers built what they called “datasets” — essentially, enormous libraries of human-generated text. The most common ones have names like Common Crawl (billions of web pages, scraped automatically), The Pile, Books3, WebText, LAION (for images), FineWeb.
Most of the leading generative AI developers trained their models by ingesting massive amounts of copyright-protected works without authorization — scraped or downloaded from the internet and compiled into datasets that often contain copies of pirated works.
Let that land.
Pirated works. Not “publicly available.” Not “fair use grey zone.” In multiple documented cases: actual piracy. Books downloaded from shadow libraries. Music ripped from protected servers. Articles stripped from paywalls. Code pulled from private repositories.
In the Bartz v. Anthropic case, Anthropic faced a potentially massive statutory damages penalty for downloading millions of pirated copies of works it used for training. The case ultimately settled for $1.5 billion.
$1.5 billion. For one case. Against one company. Out of dozens of companies. Running thousands of models.
Now ask yourself: how much of your work was in those datasets?
The Enclosure Act of Our Time
In 14th-century England, there was a thing called The Enclosure Movement.
For centuries, peasants and villagers had farmed the “commons” — shared land that belonged, in practice, to everyone. Then the aristocracy discovered something: if you fence off the commons, call it private property, and force the people who used to farm it freely to pay for access — you become very, very rich.
The people who lost the commons didn’t disappear. They just became landless. Dependent. Workers on land they used to own as a community.
We are living through the digital version of this story.
For twenty-five years, the internet grew because millions of people contributed freely — believing in an open commons of human knowledge. Wikipedia. Open-source code. Fan forums. Creative Commons music. Academic papers. Personal blogs. Photography communities. Literary journals. Independent journalism.
The AI companies looked at this commons and saw something different.
They saw raw material.
They scrape photos, videos, books, blog posts, albums, paintings, and photographs to train their products — usually without any compensation to or consent from the creators. OpenAI and Google are arguing that a part of American copyright law known as the “fair use doctrine” legitimizes this data collection. Ironically, OpenAI has also accused other AI giants of data scraping “its” intellectual property.
Read that last sentence again. OpenAI has accused others of stealing its intellectual property — the same intellectual property it built by taking everyone else’s.
The audacity of this is almost beautiful. Almost.
What “Fair Use” Actually Means (And What They’re Doing With It)
In the United States, “fair use” is a legal doctrine that allows limited use of copyrighted material without permission — for education, commentary, parody, research. It was designed to protect people from corporate overreach.
The AI companies have inverted this.
They are using “fair use” — a doctrine designed to protect the small from the powerful — to protect the most powerful companies in history from the people whose work they consumed.
A court ruled that AI training on copyrighted books constitutes fair use, but storing pirated copies does not. However, the court identified what it considered a potentially more persuasive argument: that AI models can flood the market with similar texts and stifle competition. The court emphasized that the fact that books like the plaintiffs’ make for better training data supports this argument — training the model on books can make it better at diluting the market for those books.
This is the hidden blade in the argument.
The same creative work that makes the AI smarter also makes the AI a better competitor against the person who created it.
Your novel teaches the AI to write novels. Your journalism teaches it to write journalism. Your music teaches it to write songs. Then the AI — trained on your craft, your voice, your years of accumulated skill — competes with you in your own market, using a product built at a fraction of the cost.
Some warn that requiring AI companies to license copyrighted works would throttle a transformative technology, because it is not practically possible to obtain licenses for the volume and diversity of content necessary to power cutting-edge systems. Others fear that unlicensed training will corrode the creative ecosystem, with artists’ entire bodies of works used against their will to produce content that competes with them in the marketplace.
There it is. The honest tension. Spoken plainly.
And notice which side has lawyers, lobbyists, and billions of dollars.
The Bunker Boys and Their Magic Words
Let me introduce you to a character type: The Bunker Boy.
The Bunker Boy is not a person. It’s a posture. It’s the orientation of those who have retreated behind walls of capital, legal teams, and technical complexity — issuing declarations about “innovation” and “the future” while quietly harvesting everything built by the people who won’t be let into the bunker.
The Bunker Boys have a few magic words they use to make theft sound like progress.
“Publicly available data.” As if posting something online is consent to have it scraped, processed, monetized, and sold forever. “They’re taking personal data that has been shared for one purpose and using it for a completely different purpose without the consent of those who shared the data. It is by definition a privacy violation, or at least an ethical violation, and it might be a legal violation.”
“Fair use.” Already addressed. They’ve turned a shield into a sword.
“It would be impossible to license everything.” This is perhaps the most stunning piece of circular logic in tech history. It’s impossible to pay for what you took without permission, therefore it was legal to take it without permission. The scale of the theft is the alibi for the theft.
“We’re building for humanity.” The same humanity whose work you took without asking. Whose livelihoods your product threatens. Whose cognitive commons you fenced.
What’s Missing From This Conversation
Here is where we need to dig deeper than the lawsuits.
The lawsuits are real. The copyright claims are legitimate. The settlements are meaningful. Over 800 artists have signed an open letter accusing technology companies of “theft” of copyrighted work to train their AI models — writers, musicians, and actors including Scarlett Johansson, Cate Blanchett, and Breaking Bad creator Vince Gilligan — demanding that companies engage in “ethical” partnerships rather than stealing.
But the legal frame may be too narrow for the full weight of what’s happening.
The legal conversation asks: did they violate copyright?
The deeper question is: what kind of civilization do we want to build?
Because here is what’s actually true, underneath all the court filings and fair use arguments:
Every piece of intelligence, creativity, and understanding that these AI systems possess — they got from us.
Not from venture capitalists. Not from GPU servers. Not from Sam Altman or Elon Musk or Jeff Bezos.
From writers who stayed up until 3am getting a chapter right. From musicians who played the same chord progression a thousand times until it sang. From journalists who went to dangerous places to find out what was true. From programmers who built open-source libraries for free because they believed in sharing. From teachers who put their lesson plans online so other teachers could use them. From scientists who published their findings in accessible language so laypeople could understand.
The AI is not artificial intelligence.
It is compressed human intelligence. Yours. Mine. Your grandmother’s. Your child’s teacher’s. The monk who wrote out manuscripts by candlelight. The blues musician recording in a Memphis basement in 1952. The programmer who posted a solution on Stack Overflow at midnight because they knew someone else would need it.
All of it — distilled, commodified, packaged, and sold.
The Colonial Parallel Runs Deeper Than You Think
First Nations communities around the world are looking at these scenes with knowing familiarity. Long before the advent of AI, peoples, the land, and their knowledges were treated in a similar way — exploited by colonial powers for their own benefit. What’s happening with AI is a kind of “digital colonialism,” in which powerful tech giants are using algorithms, data and digital technologies to exert power over others, and take data without compensation or consent.
Digital colonialism. Not a metaphor. A structural analysis.
The colonial playbook has always been the same:
Arrive at a place where resources exist
Declare those resources “unowned” (terra nullius — nobody’s land)
Extract the resources at scale
Build wealth from the extraction
Sell the refined product back to the people you extracted it from
Call this “development”
The internet’s creative commons was declared terra nullius by the AI companies. Unowned. Free to take. And then — refined, packaged, and sold back to creators as a productivity tool they’re now told they must use or fall behind.
This is the deepest injustice: not just that it was taken, but that you are now expected to be grateful for the theft.
What Is Actually Happening in the Courts
Let’s be precise, because precision matters here.
There have now been over 70 infringement lawsuits by copyright owners against AI companies. There were several big developments in 2025, including orders on summary judgment, high-profile settlements, and dozens of new cases filed.
The key cases and what they tell us:
Bartz v. Anthropic — Authors sued over pirated book datasets. Settled for $1.5 billion. Estimated $3,000 per work to individual authors after legal fees. The court found that AI training on books constitutes fair use, but storing pirated copies does not. The precedent is murky. The payout to individual creators is almost insulting given the scale of the industry built on their work.
Thomson Reuters v. Ross Intelligence — The court granted summary judgment in favor of Thomson Reuters, finding that using headnotes to train an AI legal research tool was not fair use. The case is currently on appeal.
UMG/WMG v. Udio — Music industry settled, and crucially, structured the license on an opt-in basis, giving copyright owners and creators control over their works, rather than an unworkable “opt-out” option that many AI companies have promoted.
The opt-in versus opt-out distinction is everything.
Opt-out says: your work is ours by default. You must actively fight to reclaim it. Opt-in says: your work is yours by default. They must actively earn access to it.
One treats creators as the default source. The other treats them as the default owner.
The entire philosophical war over AI and creativity comes down to those two words.
The Consent Architecture of a Hobson’s Choice
A more insidious way digital colonialism materialises is in the coercive relinquishing of data through bundled consent. Have you had to click “accept all” after a required phone update or to access your bank account? In reality, the only option is to agree. What would happen if you chose to reject bundled consent? You might not be able to bank or use your phone. It might appear you have options — but if you don’t tick “yes to all,” you’re choosing social exclusion.
This is the infrastructure of manufactured consent. The terms of service nobody reads. The cookie banners engineered to make “accept all” one click and “customize settings” fifteen. The robot.txt files that the AI scrapers honored until they didn’t. The retroactive claims that public posting equals permission.
They built a system where participation in modern life is consent. Where existing online is an agreement. Where the only way to protect your work is to not exist in the digital space that the entire modern economy runs on.
For an independent creator in Bohuslän, Sweden — or Lagos, or Mexico City, or rural Montana — there was no realistic choice. To publish was to be harvested. To share was to donate. To exist creatively in the digital commons was to contribute to someone else’s multi-billion-dollar training run.
What We Do Now
This is not a counsel of despair. Pay attention.
The tide is turning — but it needs weight behind it.
The lawsuits are real. Pulitzer Prize-winning journalists have filed copyright lawsuits accusing AI giants of using pirated copies of their books to train large language models. The music industry has established opt-in licensing as a model. The EU court ruling expected in late 2026 on AI training and copyright could reshape the entire landscape for European creators.
But legal battles are fought by those with resources. Individual creators — the Substack writers, the indie musicians, the bloggers, the poets — don’t have class action budgets.
So here is what the independent creator can do:
Name it. What happened has a name: extraction without consent, compensation, or credit. Say it clearly. Every time. In every post. In every conversation. Language shapes the legal and cultural battlefield.
Organize around opt-in. The music industry’s settlement established a precedent. Opt-in is the correct frame. Your work is yours by default. Any AI company that wants to use it must come to you, not the other way around.
Publish anyway. The worst outcome is silence. Silence is also training data — for a world in which creators decided the bunker boys had won. They haven’t won. Seventy lawsuits say they haven’t won. Eight hundred artists say they haven’t won. Every person who keeps writing, keeps publishing, keeps making things that are irreducibly human — they haven’t won.
Build collective infrastructure. The reason the AI companies got away with this is that creators were dispersed. Writers in silos. Musicians in silos. Journalists in silos. The answer is not more individualism. It is Ubuntu: I am because we are. The creative economy needs collective licensing bodies, shared legal resources, and cross-disciplinary solidarity.
Hold the frame. The frame they want is “innovation vs. restriction.” The correct frame is sovereignty vs. extraction. The same frame that has always applied when the powerful come for the commons of the people.
The Missing Link: What Nobody Is Saying Loudly Enough
Here it is. The thing that ties all the threads together.
The bunker boys didn’t just take your work.
They took the entire accumulated record of human consciousness and turned it into a product.
Every insight ever written down. Every poem of grief. Every scientific breakthrough. Every love letter. Every confession. Every joke. Every prayer. Every manual, every myth, every children’s story, every revolutionary manifesto, every song sung to no one in particular.
All of it — the whole human record — is now inside these machines. And the machines belong to about forty people.
This is not an intellectual property dispute.
This is a question about who owns the collective mind of our civilization.
And that question is too important to be left to lawyers, tech bros, and venture capitalists.
It belongs to us.
All of us who wrote the books.
All of us who made the music.
All of us who stayed up late trying to say something true.
We wrote the training data.
The least they owe us is the conversation about what that means.
All is One — returning to Source as Sovereign Light.
Peace, Love and Respect 🙏
The Quantum Skald COGNITIVE-LOON | Restoration of Perception hejon07.substack.com
If this resonated with you, a like or comment goes a long way. It tells the algorithm this matters — and helps it find the people who need to hear it too. Think of it as passing the torch. 🙏
Support independent journalism: buymeacoffee.com/cognitiveloon | Swish: 0729990300
Next in the series: The Opt-In Economy — What a Creator-Sovereign Internet Would Actually Look Like


