Society & Policy Aug 25, 2026 at 16:277Add to bookmarks

WikiHow has sued OpenAI for using its content without permission in AI training. What makes this case legally distinct from *The New York Times* or Getty Images suits is that WikiHow published its content under Creative Commons licensing—free to use, but not for commercial purposes. That NC clause may be the cleanest copyright argument against training data use yet assembled.
WikiHow publishes free step-by-step guides for doing almost anything. It made this content free on purpose, for public use. WikiHow is now suing OpenAI for using that content to train a commercial AI product and reproduce its substance in ChatGPT—without authorization or compensation.
WikiHow's case, reported by Medianama, alleges that OpenAI scraped and copied WikiHow's copyrighted articles to train its models, use them in RAG systems, and reproduce their substance in ChatGPT responses. The claim is copyright infringement, with the added dimension that WikiHow's content was published for public benefit—not for extraction by a commercial AI product.
This adds WikiHow to the growing list of content creators in legal conflict with AI labs: The New York Times, Getty Images, individual authors through the Authors Guild, music publishers. The suits pursue the same fundamental question through different plaintiffs: does training a commercial AI on copyrighted content constitute infringement?
Content published freely online is not necessarily free to use for any purpose. Copyright law grants creators control over reproduction and commercial exploitation of their work, regardless of whether they charge for public access. WikiHow's suit tests whether OpenAI's commercial training use falls within fair use exceptions or constitutes the kind of unauthorized reproduction that copyright protects against.
WikiHow's content is explicitly instructional and procedural—step-by-step guides that AI models can reproduce almost verbatim in response to user questions. The RAG allegation (using WikiHow articles in retrieval-augmented responses) is potentially cleaner than the training data allegation, because the reproduction is more direct and the commercial harm to WikiHow—users who would otherwise visit WikiHow get their answer from ChatGPT—is more traceable.
Labs that trained on everything and worried about licensing later are carrying increasing legal exposure. The question is whether eventual settlement cost—likely paid through licensing deals rather than court judgments—is a rounding error or a structural operating cost. As more plaintiffs file, the negotiating position of content owners strengthens.
If you're building AI products and your training or retrieval data includes content from sites with their own licensing terms, WikiHow's suit is the moment to review data provenance and usage claims. The "we didn't know" and "it was freely available" defenses have been shrinking for two years. This case adds "it was freely available but for non-commercial public benefit" to the challenge list.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
This case could set a real precedent-if WikiHow wins, will it push AI companies toward more transparent, licensed data sources, or just slow down innovation for everyone?
What if WikiHow wins and suddenly AI training becomes a paid service? Would that even be feasible for most projects?
If WikiHow wins, will AI just pivot to uncopyrighted data? Feels like opening one loophole while closing another.
If AI training counts as fair use, where does that leave small creators? The line between scraping and stealing feels thinner every day.
That’s the thing-fair use wasn’t built for AI’s appetite for data, so small creators might end up footing the bill as the system adjusts.
This is about time. If AI scrapes content without consent, there’s no future for creators. Hope the court sides with WikiHow-fair use shouldn’t mean free for all.
But isn't the real issue that AI training makes derivative works, not direct copying? Might fair use still apply even if they didn't ask?
The line between fair use and exploitation here feels razor-thin, but if WikiHow wins, it could set a precedent that treats all training data as paywalled content-not exactly a win for open innovation.