The AI copyright lawsuit is really about who pays for data

The New York Times has escalated its AI copyright lawsuit against OpenAI and Microsoft, alleging in newly unsealed documents that the firms knew training on its articles was theft. Our take: this is a fight over who pays for training data — and whoever built on unpriced data has a business-model problem, not a legal footnote.

TL;DR: The New York Times isn't only defending its own archive — it's testing whether generative AI was priced correctly in the first place, and the answer changes what your tools cost.

Key takeaway: If courts decide training data must be licensed, the cost of the underlying models rises — and that flows down to every team relying on them.

Why it matters: The durable, defensible thing you own isn't the model or its output. It's your workflow — the human-authored process that turns AI into results.

What happened

The New York Times has sharpened its case against OpenAI and Microsoft. In a court document unsealed this week, it alleges the companies didn't merely stumble into using its journalism — they knew.

According to reporting on the filing, the Times described the conduct as an astonishing theft of unprecedented proportions when millions of its articles were used to train AI models.

Source: RTE, 2026

That framing matters. Negligence is one legal problem; alleged knowing infringement is a much larger one, because it raises the stakes from an accident to a choice.

Most commentary will treat this as a courtroom drama

The consensus read is straightforward: a legacy publisher versus two of the most valuable companies on earth, with billions and a precedent at stake. Watch the ruling, watch for a settlement, move on.

That's a fair summary of the headline. It's also the least useful part. The interesting question isn't who wins the case — it's what the case reveals about how the whole industry was costed.

Our take: the workflow is the IP, not the output

Here's the uncomfortable thing the lawsuit exposes. A lot of generative AI was built on data nobody priced. If a court decides that data has to be licensed, the cost of the underlying models goes up, and that increase doesn't stay in San Francisco — it lands in your subscription.

So the question every marketing leader should ask isn't "which model is smartest". It's "what do I actually own here?"

In our experience building agents, the answer is rarely the AI output. Machine-generated text sits on shaky ground for copyright, and it can be regenerated by anyone with the same prompt. It isn't a moat. It's a commodity you're renting.

What you do own is the workflow — the specific, human-authored sequence of decisions, brand rules, review gates and data sources that turns a generic model into something only your team produces. That's the protectable layer. It survives a price rise on training data because it doesn't depend on any one model being cheap or even any one model existing.

We think this reframes the whole panic. If licensing makes models pricier, teams with a documented, portable workflow can swap the engine underneath and keep going. Teams that outsourced their thinking to a single vendor's raw output have nothing to move. The lawsuit is a stress test for that difference.

This is why we design our content creation agent around your process, not around a prompt. The brand rules, the sources and the approval steps are the asset. The model is a swappable part. When the economics shift — and this case suggests they will — you want the valuable bit to be the part you control.

None of this means the tools stop being useful. It means you stop confusing the output with the ownership. Skeptical of the hype, sold on the discipline: that's the position that ages well.

What this means for marketing teams

  • Document your top three content workflows this quarter — inputs, decisions, brand rules and sign-off — so the value lives in the process, not one vendor's account.
  • Treat model choice as reversible. Before you standardise, run a 30-day test that you could repeat on a different model without rebuilding everything.
  • Audit where AI output touches published work, and keep a human-authored layer of editing on anything you'd want to protect or claim as yours.
  • Budget for the possibility that model costs rise if training data has to be licensed — build a 15-20% headroom into next year's tooling forecast.
  • If you're unsure which parts of your stack are actually defensible, our team can talk through a workflow audit before you commit to a platform.

Frequently asked questions

Is AI-generated content protected by copyright?

Generally not on its own. In most jurisdictions, purely machine-generated output lacks the human authorship copyright requires. Meaningful human creative input and editing are what make a work protectable.

What is the New York Times suing OpenAI and Microsoft over?

The Times alleges the companies used millions of its articles to train AI models without permission. Unsealed filings claim the firms knew this use infringed its copyright.

Could an AI copyright ruling make AI tools more expensive?

Possibly. If courts require AI firms to license training data, those costs could pass down to users. Teams should build budget headroom and keep their workflows portable across models.

Written by the Anjin team - we build AI marketing systems and remain professionally unimpressed by hype.

Continue reading