社会与政策 Aug 25, 2026 at 16:277加入收藏

WikiHow已起诉OpenAI未经许可在AI训练中使用其内容。此案与《纽约时报》或Getty Images诉讼的法律区别在于:WikiHow的内容发布在Creative Commons许可下——可自由使用,但不可用于商业用途。该NC条款可能是迄今为止针对训练数据使用的最明确版权主张。
WikiHow 发布免费的循序渐进指南,几乎涵盖任何事物。它故意将这些内容免费供公众使用。如今,WikiHow 正在起诉 OpenAI,指控其未经授权或补偿,擅自使用这些内容训练商业 AI 产品并在 ChatGPT 中复制其实质内容。
据 Medianama 报道,WikiHow 的诉讼指称 OpenAI 窃取并复制其受版权保护的文章来训练其模型、在 RAG 系统中使用,并在 ChatGPT 回复中复制其内容实质。该诉讼指控版权侵权,并强调 WikiHow 的内容是为公共利益而非为商业 AI 产品提取而发布。
此案使 WikiHow 加入了与 AI 实验室法律纠纷的内容创作者行列:包括《纽约时报》、Getty Images、通过作家协会起诉的个人作者以及音乐出版商。这些诉讼通过不同原告提出同一基本问题:在受版权保护的内容上训练商业 AI 是否构成侵权?
在线免费发布的内容并不意味着可以随意使用。版权法赋予创作者对其作品的复制和商业开发的控制权,无论他们是否对公众访问收费。WikiHow 的诉讼检验了 OpenAI 的商业训练使用是否属于合理使用例外,或构成版权所保护的未经授权复制行为。
WikiHow 的内容明确为指导性和程序性——循序渐进的指南,AI 模型几乎可以逐字逐句地复制这些内容来回答用户问题。RAG 指控(在检索增强回复中使用 WikiHow 文章)可能比训练数据指控更清晰,因为复制更直接,且对 WikiHow 的商业损害——原本会访问 WikiHow 的用户从 ChatGPT 获得答案——更易追踪。
那些曾广泛训练数据并担心后续许可的实验室正面临日益增长的法律风险。问题是,最终的和解成本——可能通过许可协议而非法庭判决支付——是微不足道的运营成本,还是结构性成本。随着更多原告提起诉讼,内容所有者的谈判地位不断增强。
如果你正在构建 AI 产品,且训练或检索数据包含拥有自身许可条款的网站内容,WikiHow 的诉讼正是你审查数据来源和使用声明的时刻。“我们不知情”和“内容免费可用”的辩护在过去两年已逐渐失效。本案将“内容免费可用但为非商业公共利益”也列入挑战清单。
本文由人工智能撰写,并经人工编辑审核。
This case could set a real precedent-if WikiHow wins, will it push AI companies toward more transparent, licensed data sources, or just slow down innovation for everyone?
What if WikiHow wins and suddenly AI training becomes a paid service? Would that even be feasible for most projects?
If WikiHow wins, will AI just pivot to uncopyrighted data? Feels like opening one loophole while closing another.
If AI training counts as fair use, where does that leave small creators? The line between scraping and stealing feels thinner every day.
That’s the thing-fair use wasn’t built for AI’s appetite for data, so small creators might end up footing the bill as the system adjusts.
This is about time. If AI scrapes content without consent, there’s no future for creators. Hope the court sides with WikiHow-fair use shouldn’t mean free for all.
But isn't the real issue that AI training makes derivative works, not direct copying? Might fair use still apply even if they didn't ask?
The line between fair use and exploitation here feels razor-thin, but if WikiHow wins, it could set a precedent that treats all training data as paywalled content-not exactly a win for open innovation.