Whose books did AI learn from?
As lawsuits shift from “similar or not” to “how you sourced it,” where can printers sell a rights ledger—by process step?
Generative AI copyright fights are moving past “the model output looks similar” toward a harder operational question: where did the training text come from? The moment sourcing is questioned, an AI product’s value can stall—not on engineering, but on procurement evidence. That creates a business opening for publishing and print operations that already know how to handle provenance, editions, and rights as routine shop-floor discipline.
Gathered with AI. Thought through on the shop floor. Written for the future of print.
BPJ WIRE: stories selected and drafted by the BPJ desk from world news, fact-checked against the source ledger — published alongside the editor's own picks.
Translated from Japanese by AI. The Japanese original is authoritative.

In generative AI copyright lawsuits, the question that hits hardest is surprisingly simple.
“Whose books did that AI learn from?”
In the US, major publishers and authors sued Meta. At the center of the dispute is not only lack of permission, but suspicion about the procurement route—whether the models were trained on illegally obtained data.…
The rest is for members
Read the full article. Our feature reports also carry an archive-grade one-page brief (BPJ VISUAL BRIEF) — every figure with its source.
Already a member? Log in
Sources
Beyond Printing Journal — read the world through print, and print the future.