Book Writers Sue Google Over Alleged Use of Copyrighted Books to Train Gemini | Free Download

Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow and SCRIBE, Inc. has filed a proposed federal class-action lawsuit against Google.

He accused the company of copying copyrighted books and journal articles without permission to train the Gemini AI model. The suit was filed on July 10, 2026. The case, Hachette Book Group Inc. et al. v. Google LLC is in the U.S. District Court for the Southern District of New York, case number 1:26-cv-05870.

The complaint describes Google’s alleged actions as “one of the largest violations of copyrighted material in history.” However, these allegations have not been proven in court, and no ruling has been issued on Google’s liability.

Legal claims and publisher deals at center of Gemini lawsuit

The plaintiffs allege four legal claims: direct copyright infringement, contributory copyright infringement, removal or alteration of copyright-management information, and violation of the Digital Millennium Copyright Act.

They argue that Google obtained protected material through its book services, Internet scraping, and other sources before using the material to train Gemini.

Publishers provided books and journal articles to Google Books, Google Play Books, and Google Scholar under agreements that included specific uses, such as displaying searchable excerpts, distributing e-books, and helping users discover academic publications.

The plaintiffs claim that these agreements did not authorize Google to copy entire works in datasets used for commercial generator AI development, and that Google reused content obtained through publishing relationships established for uses not covered by the agreements.

The complaint also states that Google collected books and other protected content through extensive Internet scraping, including from sources such as pirated websites and publications behind subscriptions or paywalls.

How is this Gemini case different from Google Books, and what the internal emails say

Google previously won a copyright case related to digitizing books for search and snippet previews. In 2015, a federal appellate court ruled that these uses were transformative and protected under fair use.

The plaintiffs argue that this earlier decision involved searchable book databases and limited excerpts, not the use of entire books to train generative AI models.

They claim that AI training serves a different business purpose and results in systems that can compete with the original tasks involved in training. They say the main difference is between search-centric digitization and the training of generative models, which is central in their case.

According to the complaint, a Google employee warned during internal discussions that using books offered through Google Play publishing agreements for AI development could pose serious legal risks. The employee reportedly estimated that such behavior could lead to a fine of ten to one hundred billion dollars.

Other internal documents described in the complaint show that Google sought professionally written books to enhance the performance of its Gemini AI system.

The plaintiffs claim that tests showed that models trained only on public domain books performed worse than models trained on collections that included copyrighted works. The complaint states that Google aims to include works containing curated facts, organized analysis, fictional narratives, and professionally edited writing.

The works cited in the complaint and what Gemini is accused of producing

The complaint highlights several actions as examples. Hachette lists Peter Brown’s The Wild Robot, NK Jemisin’s The Fifth Season, Becky Lomax’s Moon Glacier National Park, and Lemony Snicket’s Who Could That Be at This Hour?

It also mentions Innocent of Turo. Cengage includes Cognitive Psychology, Principles of Economics, Milady Standard Barbering, Nutrition: Concepts and Controversies, and Calculus: Early Transcendentals. Elsevier points to copyrighted journal articles in its section on complaints.

TURO AND SCRIBE REFERENCES PRESUMED INNOCENT, INNOCENT, AND TESTIMONIAL. The plaintiffs claim that these examples are just a sampling of the books and articles allegedly copied in connection with Gemini.

The filing offers examples to show that Gemini can produce content related to specific protected books. It claims that the system created content based on The Fifth Season and Who Could Be in This Hour? Prepared responses incorporating characters, events and details from. Plaintiffs also argue that Gemini can generate low-cost alternatives to professionally published works.

He estimates that the service can produce a 100-page murder mystery in a quiet seaside town in about 20 minutes for 39 cents. “No publisher or author can compete with him,” the complaint says. These time and cost figures are allegations by the plaintiffs and have not been verified by the court.

Who can join the class and what remedy are the plaintiffs seeking

The proposed class would include owners of registered copyrights in books that have International Standard Book Numbers, as well as journal articles identified through Digital Object Identifiers or International Standard Serial Numbers.

To be part of the class, members must demonstrate that Google has copied their works from one of its services, obtained them through web scraping, or used them in connection with Gemini training. The Court has not yet certified this proposed class.

The plaintiffs are seeking statutory damages or compensation based on their claimed losses and Google’s alleged profits, along with legal fees and other remedies available under federal copyright law.

They are asking the court to stop Google from continuing unauthorized copying through an injunction that would restrict the use of protected books and articles in Gemini training and related AI development.

The complaint also requests an accounting of the work and methods Google used to train Gemini. It includes details about the materials obtained, where they were obtained, and how they were incorporated into Google’s systems.

The plaintiffs are seeking court-supervised destruction of infringing copies and datasets obtained from their operations. It is not clear whether such an order would be legally feasible or appropriate in this case.

Publishers, authors and academic rights holders who may be part of the proposed class should take some practical steps as the case develops. They should verify which of their titles are registered with the US Copyright Office as the proposed class depends on registered works marked with an ISBN, DOI or ISSN.

It is also advisable to review existing agreements with Google Books, Google Play Books, and Google Scholar to understand what those agreements allow.

Additionally, they must keep records of publication dates, distribution channels, and any communications with Google regarding the use of their works in AI training datasets. It may also be important to track the progress of class certification, as inclusion in any potential class will depend on the court’s decision.

Where does the Hachette and Cengage lawsuit against Google stand now?

Hachette and Cengage previously attempted to join a separate copyright lawsuit filed against Google in California in 2023, involving authors and visual artists. Google opposed their involvement, and the publishers later withdrew that effort before the present case was presented in New York.

Google has not yet issued a detailed public response to the specific allegations in the New York complaint. The Company will have an opportunity to respond as the case progresses through the motion, potential discovery, and class-certification processes. No hearing date has been set, and the court has not ruled on any of the four claims.

Thanks for being a Ghax reader. The post Book authors sue Google over alleged use of copyrighted books to train Gemini appeared first on gHacks.

Source:Ghacks

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top