Nearly every major AI language model in use today was built, at least in part, by reading copyrighted books without the authors' permission. Whether that was legal is still an open question moving through federal courts โ€” and the answer will shape how much these tools cost, and who controls them, for years to come.

The basic dispute is straightforward. AI companies need enormous amounts of text to train large language models, and books are a rich source of well-structured, high-quality writing. Several major developers, including OpenAI, Meta, and Anthropic, have been sued by authors and publishers who say their books were used without consent or payment. The companies argue this qualifies as fair use, a legal doctrine that allows copyrighted material to be used without permission under certain conditions, such as commentary, research, or transformation into something new.

Courts have started weighing in, and the results have been mixed rather than a clean win for either side. In some rulings, judges have suggested that training a model on copyrighted text can qualify as fair use because the model doesn't reproduce the original work โ€” it learns patterns from it. But those same rulings have drawn a sharp line around how the material was obtained. Downloading books from pirated sources, rather than purchasing them, has been treated far more skeptically by courts than legally acquired copies used for training.

That distinction produced one of the largest settlements in AI history. Anthropic agreed to pay authors a substantial sum after facing claims tied to its use of pirated book repositories, even as the broader fair-use question around training itself remained unresolved. Other cases against OpenAI and Meta are still working through the courts, and none of them are expected to produce a single, industry-wide answer soon.

This pattern echoes earlier fights over digital content. When Google scanned millions of books for its Google Books project in the 2000s, authors and publishers sued, and it took the better part of a decade before courts settled on fair use protections for that specific use. Music file-sharing fights in the early 2000s followed a similar arc: years of litigation, patchwork rulings, and eventual settlements that reshaped how companies operated going forward, without ever producing one tidy legal rule.

For small business owners, this legal fog doesn't mean you should stop using AI tools built on these models. It does mean the tools you rely on today were built on a legal foundation that's still being poured. If a court eventually finds that a company's training data was obtained illegally, that company could face damages, be forced to retrain models on cleaner data, or face new licensing costs โ€” expenses that tend to get passed down to customers eventually.

There's also a practical distinction worth knowing: fair use fights are mostly about how models were trained, not about whether you're breaking the law by using a chatbot for your business. Individual users aren't the target of these lawsuits. But the outcomes will influence pricing, availability, and which companies survive with their current products intact.

Watch for a few concrete signals over the coming months: how appellate courts rule on the training-as-fair-use question once cases move beyond trial courts, whether more companies follow Anthropic's lead in settling rather than litigating, and whether any AI vendor starts marketing itself around licensed or fully-cleared training data as a selling point.

The bottom line: the legal status of AI training on books remains unsettled, with courts drawing a line between the act of training and the act of acquiring source material illegally. Small business owners using these tools face no direct legal exposure from that fight, but should expect pricing, licensing terms, and vendor reliability to keep shifting as the litigation plays out.