AI in the content lifecycle: three years later
In June of 2023, Sarah O’Keefe authored this white paper on AI in the content lifecycle. Three years later, here’s what’s changed and what’s still true about AI in the content lifecycle.
The hands got fixed
In 2023, you could easily spot an AI-generated image of a person: extra fingers, arms bending the wrong way, legs that quietly disappear, and the list could unfortunately go on. Sarah asked an AI image generator for a photo of a person in athletic wear working on a computer. The generator created an image of a man with a missing leg, a backwards elbow, and either an extra knee or a really long leg. (I’m not sure, but please don’t make me look at it anymore.)
Now, image generators have largely solved problems like hands, faces, and extra knees. In many cases, AI-generated images and videos can be difficult to identify.
Model collapse: still theoretical
In the white paper, Sarah flagged model collapse—the idea that AI models trained on AI-generated content would effectively “eat their own brains”—as an impending risk. Two years later, this hasn’t happened yet, but stay tuned to see if this happens in the next two years.
Part of the reason: AI labs got much more deliberate about data curation, filtering, and using verified or synthetic-but-controlled data rather than just scraping whatever the open web produces. “Entropy always wins” isn’t wrong, but perhaps entropy is beatable (or delayable) if you’re willing to do the curation work. This, however, still goes back to Sarah’s recommendation from 2023: keep your source content controlled and “known good.”
Chatbots got better at “showing the work”
In the white paper, Sarah described ChatGPT as “autocomplete with some additional guardrails,” sharing that ChatGPT could generate plausible-sounding answers with confidence, but often without substance or accuracy.
Today’s AI models have gotten better at showing the step-by-step “reasoning” behind what’s generated. This helps users identify how an AI model generated a given answer and where to adjust multi-step tasks.
The lawsuits haven’t stopped
Since 2023, an increasing number of lawsuits have been expanding the many, many legal concerns surrounding AI. Here’s a sampling of current cases.
- A federal judge found that Ross’s use of Thomson Reuters’ Westlaw headnotes to train a competing legal-research AI was not fair use, directly addressing AI training data.
- In Bartz v. Anthropic, a judge ruled that training on lawfully purchased books could be fair use, but training on pirated copies was not. Anthropic is in the process of settling with the authors and publishers in 2025 for roughly $1.5 billion, the largest publicly reported copyright settlement of its kind. One of our books is caught up in this lawsuit, and we’re not bitter about it AT ALL.
- In the case of Kadrey v. Meta, the judge sided with Meta in Meta’s claim of fair use in using copyrighted books to train its AI model, but was careful to note that this wasn’t a blanket ruling that AI training is always fair use; just that these plaintiffs hadn’t proven their case.
- In an ongoing case between The New York Times, OpenAI, and Microsoft, we will see how courts treat news content specifically.
- There are ongoing, parallel cases in the UK and US for Getty Images v. Stability AI for scraping stock photography to train AI models.
- A court in Munich ruled that Google was accountable for the false claims ChatGPT displayed in AI-generated search overviews. AI overviews were held as Google’s content.
And this list will probably be outdated by… next week.
What hasn’t changed
The core recommendation from Sarah’s white paper hasn’t changed: use AI for pattern-driven, repetitive work so that humans can focus on more critical matters. In other words, technology has become the unpredictable variable, so humans have to become the layer of quality assurance.
In a recent podcast, Sarah and Pawel Kowaluk (Guidewire Software) spoke about this:
Paweł Kowaluk: “We always knew that people are going to be whimsical and maybe harder to rein in, but the technology was going to be predictable. Whereas now, technology is not predictable anymore: you give it a prompt, and you hope it’s going to do what you want.”
Sarah O’Keefe: “And now the people are being asked to be the deterministic layer, right? To be the QA on top of the AI.”
Trust matters more than ever. In 2023, we were watching for voluntary industry commitments and early-stage regulation. Now, we’re… still watching. However, the core advice Sarah gave still stands: disclose your sources, watch for bias, keep humans accountable for the output.
