Chasing the alpha while the market sleeps — and this time, the alpha is a lawsuit. On June 11, 2025, WikiHow filed a copyright infringement suit against OpenAI, alleging the company scraped over 11,000 of its step-by-step guides without permission. The data set? Tiny in the grand scheme of AI training — maybe 0.01% of GPT-4’s total corpus. But the signal? It’s a seismic tremor that could reroute the entire data acquisition strategy of the AI industry.

Context: Why WikiHow's how-to content matters more than you think
WikiHow is the internet’s largest repository of instructional content, hosting over 240,000 articles covering everything from changing a tire to negotiating a salary. Each article is structured with clear steps, numbered lists, and practical tips. For an AI model, this is gold — not just for pre-training, but for instruction tuning. The ability to follow directions, understand causality, and produce actionable outputs is the very thing that separates a chatbot from a reasoning engine. Models trained on massive amounts of generic web text often struggle with task-oriented queries. WikiHow’s data fills that gap.
Based on my experience auditing tokenomics during the 2017 ICO craze, I can tell you that data provenance is the new tokenomics. Back then, I manually verified 50+ whitepapers for red flags — unrealistic return models, hidden vesting schedules. Now, I’m applying the same forensic lens to AI training data. The lawsuit isn’t just about copyright; it’s about who controls the inputs that shape the most powerful reasoning engines ever built.
Core: The technical anatomy of the scrape — and the hidden leverage
From a technical standpoint, OpenAI’s scrape is unremarkable. It’s standard web crawling — the same technique used by Google, the Internet Archive, and every price-comparison bot. The controversy lies in the scale and the lack of permission. But here’s the nuance most coverage misses: the value of WikiHow’s data is not in its volume but in its structure.
From my years bridging the gap between complex blockchain protocols and retail investors, I’ve learned that context matters more than raw data. A typical web crawl grabs paragraphs, but WikiHow’s articles are pre-labeled with task hierarchies. Each step is a mini-instruction chain. For an AI model, this is like having a pre-built curriculum for instruction following. OpenAI’s GPT-4 series likely used this data for fine-tuning, not just pre-training. That’s a higher-value use case — and a higher legal risk.
The industry pattern is clear: every major AI company relies on mass scraping. Meta, Google, Anthropic — they all do it. But the lawsuits are piling up. The New York Times, Getty Images, and now WikiHow. Each case chips away at the legal fiction that “publicly available” means “free for commercial training.”
Contrarian: The unreported angle — this lawsuit might actually help OpenAI in the long run
Here’s the counter-intuitive take: WikiHow’s lawsuit could accelerate the very thing OpenAI needs most — a legitimate data licensing market. Think about it. Right now, the AI industry operates in a gray zone. Data is scraped, used, and only later challenged. That uncertainty raises the cost of capital for every AI startup. What if the lawsuit forces OpenAI to negotiate a license with WikiHow, and that license becomes a template for the entire industry?
I’ve seen this play out in DeFi. When Uniswap V4 introduced hooks, the complexity scared off 90% of developers. But the remaining 10% built applications that brought liquidity back to the ecosystem. Similarly, the legal complexity of data licensing will scare off some AI companies, but the ones that adapt will create a more transparent, more valuable data supply chain. The lawsuit isn’t a threat to OpenAI’s business model; it’s a catalyst for the industry to mature.
From my perspective as a crypto operator who survived the 2022 bear market, I’ve learned that crises often force necessary infrastructure. The Terra collapse led to better auditing standards. The FTX debacle pushed for proof-of-reserves. This lawsuit will push for proof-of-data-ethics.
Takeaway: What to watch next
The real signal isn’t the court docket — it’s the behavior of content platforms. Scanning the noise for the signal — I’m watching Reddit, Stack Overflow, and Medium. If they file similar suits within six months, we’ll see a domino effect. The AI industry will be forced to pivot from “scrape first, ask later” to “license first, train later.”
That shift will increase data costs by 10–20% in the short term, but it will also unlock a new market: data provenance as a service. Startups that can verify and license training data will become the next Coinbase or Chainlink. The ledger doesn’t lie — and in this case, the ledger is the copyright registry.

Born in the fire of the first bubble — I’ve seen this cycle before. The ICO bubble collapsed under the weight of scams, but it gave birth to Ethereum. The NFT bubble burst, but it left behind a robust digital ownership culture. The web scraping bubble will burst, but it will leave behind a framework for ethical AI training.
Human faces behind the blockchain code — at the end of the day, WikiHow is a collection of people trying to teach others how to fix a leaky faucet or bake a cake. OpenAI is a collection of people trying to build a reasoning engine. The lawsuit is about money, yes, but it’s also about respect. Respect for the creators who built the knowledge base that AI now relies on.
Speed meets substance in the void — the market will move on from this news in a week. But the infrastructure built in response will last decades. Keep your eyes on the data licensing platforms, not the stock price.
From ICO hype to on-chain truth — and now, from scrape hype to copyright truth. The cycle continues.