Delhi HC’s Threefold Test for AI Training: Why ANI Failed to Stop OpenAI’s Use of Its News Content

OpenAI ANI copyright case

Delhi High Court: While deciding an application for interim injunction in a copyright infringement suit concerning the use of ANI Media (P) Ltd.’s (ANI) copyrighted news content for training generative artificial intelligence (AI) models, a Single Judge Bench of Amit Bansal, J., held, prima facie, that OpenAI’s storage of ANI’s literary works for training the LLMs underlying ChatGPT fell within the scope of Section 52(1)(a), Copyright Act, 1957 (Copyright Act) and did not amount to infringement. The Court held that such use qualified as “private or personal use, including research” and satisfied the requirements of fair dealing, as it was limited to training, did not result in market substitution, and furthered public interest in technological innovation and dissemination of knowledge.

The Court further held that ANI failed to establish memorisation, regurgitation, or substantial reproduction of its works through ChatGPT outputs. Accordingly, the Court declined to grant interim injunction, holding that ANI had failed to establish a prima facie case and that the balance of convenience and irreparable injury weighed against the grant of relief.

Background

The Court was called upon to consider a significant copyright dispute concerning the use of copyrighted content for training generative AI models. The plaintiff, ANI, instituted a suit against OpenAI, alleging that the latter had, without authorisation, used ANI’s copyrighted news content to train its large language model (LLM) and that ChatGPT was capable of reproducing ANI’s copyrighted works in its responses. ANI sought an interim injunction, contending that OpenAI’s acts amounted to copyright infringement on 2 fronts:

  1. the copying and storage of ANI’s content for training its AI model (the training claim), and

  2. the generation of outputs that reproduced ANI’s copyrighted expression (the output claim).

The Court noted at the outset that the dispute presented a novel challenge arising from rapid advances in AI, requiring the Copyright Act to be interpreted in a technological context that the legislature could not have envisaged. It further observed that while foreign judicial precedents on AI and copyright could offer persuasive guidance, their applicability in India must be assessed cautiously, having regard to the differences in statutory frameworks. With the aforesaid backdrop, the Court proceed to decide the present application.

Issues

  1. Whether the storage by the defendants of plaintiff’s data (which is in the nature of news and is claimed to be protected under the Copyright Act for training its software, i.e., ChatGPT, would amount to infringement of plaintiff’s copyright.

  2. Whether the use by the defendants of plaintiff’s copyrighted data in order to generate responses for its users, would amount to infringement of the plaintiff’s copyright.

  3. Whether the defendants’ use of plaintiff’s copyrighted data qualifies as “fair dealing” in terms of Section 52(1)(a), Copyright Act.

  4. Whether the courts in India have jurisdiction to entertain the present law suit considering that the servers of the defendants are located in the United States of America.

Analysis

The Court explained that LLMs are trained on an enormous corpus of publicly available and licensed data, which is tokenised, embedded, and repeatedly processed to predict subsequent words, with the model being further fine-tuned to improve accuracy and safety. It further observed that LLMs generate responses based on user prompts and may also utilise retrieval-augmented generation (RAG), which references external knowledge bases in addition to the training data to produce more useful responses.

Whether the courts in India have jurisdiction to entertain the present law suit considering that the servers of the defendants are located in the United States of America

The Court observed that since the issue of jurisdiction goes to the root of the matter, it would be taken for consideration at the beginning. On the aspect of jurisdiction, based on the objections raised by OpenAI, essentially the following 2 issues arose for consideration:

  1. Whether this Court has territorial jurisdiction to entertain the present suit.

  2. Whether the Copyright Act would apply in relation to training claim since according to OpenAI training takes place on servers located outside India.

The Court held, prima facie, that it had territorial jurisdiction to entertain the suit under Section 20, Civil Procedure Code, 1908 (CPC) as well as Section 62(2), Copyright Act, noting that ANI’s principal place of business and registered office are situated within its jurisdiction, OpenAI offers its services to users across India, including within the jurisdiction of the Court, and the alleged infringing outputs were generated in India. Rejecting OpenAI’s contention that the Copyright Act would not apply since the training of its LLMs takes place on servers located in the United States, the Court observed that the storing of ANI’s works on foreign servers is merely the terminal step in the chain of events beginning with the access of copyrighted works from India. The Court held that the Copyright Act does not require it to sever the chain of events and examine only the last step, as such an approach would enable infringers to evade Indian copyright law by shifting the terminal link to servers abroad. It further found that the training claim could not be entirely divorced from the output claim, as the latter is based on the former and the alleged infringing outputs are reproduced within the jurisdiction of the Court. Accordingly, the Court concluded, at the prima facie stage, that it possessed territorial jurisdiction and that the training claim did not involve any impermissible extra-territorial application of the Copyright Act.

Whether the use by the defendants of plaintiff’s copyrighted data in order to generate responses for its users, would amount to infringement of the plaintiff’s copyright.

The Court observed that, under Section 13, Copyright Act, copyright subsists in the classes of works enumerated therein, and therein and noted that OpenAI did not dispute that ANI’s news articles and interviews constitute “original literary works”. On the question of ownership, the Court, on a prima facie view, was satisfied that the Professional Services Agreement placed on record expressly recognised ANI as the owner of the copyright in original literary works created by its personnel, in terms of Section 17, Copyright Act. The Court further held that the mere fact that such works are freely and publicly accessible does not deprive them of copyright protection, and that ANI, as the copyright owner, continues to enjoy the exclusive rights under Section 14, including the rights of reproduction and communication to the public. Accordingly, the Court held that it would have to examine whether the responses generated by ChatGPT result in the unauthorised reproduction and communication of ANI’s original literary works to the public, thereby constituting infringement under Section 51, Copyright Act. More particularly, the Court

would have to examine the following 2 issues:

  1. whether OpenAI memorises and regurgitates ANI’s copyrighted literary works in the form of responses, and

  2. whether the ChatGPT’s responses are substantial reproduction of ANI’s copyrighted literary works.

The Court examined ANI’s contention that the LLMs underlying ChatGPT memorise and regurgitate copyrighted works and observed that an LLM is not designed to reproduce its training data verbatim but functions by predicting outputs based on learned patterns. It noted that memorisation may occur only in limited circumstances and that, significantly, the illustrative ANI articles relied upon in the plaint were published after the completion of OpenAI’s training process. Consequently, the Court held, prima facie, that the impugned responses could not have resulted from memorisation of ANI’s works during training. Instead, the responses appeared to be generated through the RAG technique, whereby information is retrieved from external sources at the time of the query. Since the plaint did not specifically plead infringement arising from RAG-generated outputs, the Court declined to treat those illustrations as evidence of memorisation.

On the allegation of substantial reproduction, the Court reiterated the settled principle that copyright does not subsist in facts, but only in the form and manner of expression of those facts. Relying upon R.G. Anand v. Deluxe Films, (1978) 4 SCC 118 and Eastern Book Company v. D.B. Modak, (2008) 1 SCC 1., the Court observed that the impugned works must be compared as a whole and that infringement is established only where there is substantial copying of the protected expression. Examining the illustrations relied upon by ANI, particularly the Neeraj Chopra interview, the Court found that ChatGPT’s responses were not substantial reproductions of ANI’s articles but contained their own expression, commentary and contextualisation. The Court further observed that the quoted statements were, prima facie, the copyright of the speaker under Section 17(cc), Copyright Act and not of ANI. Distinguishing the foreign authorities cited by ANI on the ground that those cases involved repeated verbatim reproduction of copyrighted material, the Court concluded that no comparable facts existed in the present case. Accordingly, it held, at the prima facie stage, that ANI had failed to establish that OpenAI permanently stores ANI’s works for memorisation or that the responses generated by ChatGPT amount to a substantial reproduction of ANI’s copyrighted literary works. Consequently, the Court found that no prima facie case of copyright infringement was made out in respect of the output claim, leaving the disputed factual questions to be determined at trial.

The Court further noted that Issues 1 and 3 were intertwined and had to be considered together.

1. Whether the storage by the defendants of plaintiff’s data (which is in the nature of news and is claimed to be protected under the Copyright Act) for training its software, i.e., ChatGPT, would amount to infringement of plaintiff’s copyright

3. Whether the defendants’ use of plaintiff’s copyrighted data qualifies as “fair dealing” in terms of Section 52, Copyright Act.

The Court held that although the storage of a copyrighted literary work in any medium by electronic means constitutes reproduction under Section 14(a)(i), Copyright Act, the exclusive rights conferred under Section 14 are expressly made subject to the provisions of the Act, including Section 52, which independently defines acts that do not amount to infringement. Rejecting the contention that Section 52 should be construed narrowly, the Court observed that it is not in the nature of a proviso or an exception to Section 51, but an integral part of the Copyright Act that must receive a broad and liberal interpretation so as to balance the exclusive rights of copyright owners with the competing public interest of encouraging creativity and dissemination of knowledge. The Court further held that the applicability of Section 52(1)(a) requires a twofold inquiry, namely, whether the use satisfies the purpose test and whether it amounts to fair dealing.

Applying the purpose test, the Court rejected ANI’s contention that “private or personal use, including research” under Section 52(1)(a)(i) is confined to non-commercial activities or to individual human users. It held that the legislature consciously omitted any restriction to non-commercial use in Section 52(1)(a), unlike several other provisions of the Copyright Act where such limitation is expressly provided. The Court also held that the expression “private” is wide enough to include a private company, particularly where the stored data remains in a closed environment and is not made available to the public. Interpreting the expression “research” through the doctrine of updating construction, the Court observed that research can no longer be confined to human activity and must extend to machine learning in light of modern technological developments. Since OpenAI stores ANI’s literary works only for the internal process of training its LLMs, without making the training data accessible to the public, the Court held, prima facie, that such storage amounts to “private or personal use, including research” under Section 52(1)(a) and therefore satisfies the purpose test.

The Court observed that Indian Courts have not adopted any single uniform test for determining scope of expression “fair dealing”. After surveying the existing jurisprudence, the Court formulated a threefold test for the purposes of the present case:

  1. whether OpenAI’s use of ANI’s literary works is confined to training its LLMs,

  2. whether such use results in economic competition or causes actual or potential prejudice to ANI’s legitimate commercial interests, and

  3. whether the functions performed by ChatGPT serve the broader public interest. The Court held that these factors appropriately balance the exclusive rights of copyright owners with the public interest underlying Section 52, Copyright Act.

Applying the aforesaid test, the Court found, prima facie, that ANI had failed to establish that OpenAI used its literary works for any purpose other than training its LLMs or that ChatGPT reproduced ANI’s works in a manner amounting to market substitution. The Court observed that ChatGPT performs functions fundamentally different from ANI’s business of news reporting and syndication, and that its responses are not substitutes for ANI’s news articles. It further noted that ANI had placed no material on record to demonstrate any loss of market share or subscription revenue on account of OpenAI’s operations. Emphasising the considerable public benefits arising from LLMs, including advancing scientific research, education, accessibility, knowledge dissemination and technological innovation, the Court held that the public interest factor also stood established. Accordingly, the Court concluded that both the purpose test and the fairness test under Section 52(1)(a) stood satisfied and, therefore, prima facie, OpenAI’s storage of ANI’s literary works for training its LLMs constituted fair dealing and did not amount to copyright infringement.

The Court observed that ANI had the ability to block its website from being crawled or scraped by third parties, including OpenAI, but had admittedly not exercised the available opt-out mechanism. It further noted that OpenAI had stated that it had itself blocked ANI’s website from its web crawlers and the ChatGPT search/RAG function. The Court also found that ANI had not placed any material on record to establish that OpenAI’s activities had resulted in any loss of subscribers, diminution of its news syndication business, or other commercial prejudice. Since ANI had itself offered to licence its content to OpenAI for USD 7.5 million, the Court held that ANI’s claim was quantifiable and could be adequately compensated in monetary terms if it ultimately succeeded. Conversely, an interim injunction would have a significant impact on the functioning of OpenAI’s services and could not be similarly compensated. The Court further emphasised that AI and generative AI have brought about a transformational change across multiple sectors, and that the development of LLMs depends upon access to publicly available information. It observed that requiring licences from multiple sources at the training stage would render the development of LLMs economically unviable, adversely affecting technological innovation and the broader public interest, including millions of ChatGPT users in India.

Decision

Accordingly, applying the principles of prima facie case, balance of convenience and irreparable injury, the Court declined to grant an interim injunction. It held, prima facie, that OpenAI’s storage of ANI’s original literary works for training the LLMs underlying ChatGPT falls within Section 52(1)(a), Copyright Act and, therefore, does not amount to infringement under Section 51. The Court further held that the outputs generated by ChatGPT through the RAG technique were not substantially similar to ANI’s original literary works and that ANI had failed to establish any memorisation or regurgitation of its copyrighted works by the LLMs. Consequently, ANI failed to make out a prima facie case for interim relief, the balance of convenience weighed against the grant of an injunction, and irreparable injury would be caused not only to OpenAI but also to the public at large if such relief were granted. The application for interim injunction was, therefore, dismissed, with the Court clarifying that all observations were prima facie in nature and would not prejudice the final adjudication of the suit.

Also Read: Publishers Sue Google Over Gemini AI Training Lawsuit | SCC Times

[ANI Media (P) Ltd. v. OpenAI OPCO LLC, CS(COMM) 1028 of 2024, decided on 24-7-2026]


Advocates who appeared in this case:

For the Plaintiff: Sidhant Kumar, Akshit Mago, Manyaa Chandok, Anshika Saxena and Lahar Jain, Advocates

For the Defendant: Amit Sibal, Akhil Sibal, Kapil Sibal, Arvind P. Datar, Haripriya Padmanabhan, Rajshekhar Rao, Chander M. Lall, Senior Advocates with Sanjeev Kapoor, Nirupam Lodha, Madhav Khosla, Gautam Wadhwa, Moha Paranjpe, Abhi Udai Singh Gautam, Rebecca Cardoso, Hardik Malik, Vanshika Thapliyal, Rajat Bector, Aditya Gupta, Asavari Jain, Shuvam Bhattacharya, Vani Kaushik, Riddhie Bajaj, Jahnavi Siddhu, Aishwarya Kane and Sauhard Alung, Shashank Mishra, Akshi Rastogi, Parv Kaushik and Suvaroop Saha Roy, Shrutanjaya Bhardwaj, Akshat Agrawal, Tushar Srivastava, Shourya Das Gupta, Siddhi Nagwekar, Yashi Bajpai and Yash Tayal, Ankit Sahni, Kritika Sahni, Chirag Ahluwalia, Mohit Maru, Tanisha Sharma, Ameet Datta, Harsh Kaushik, Riddima Sharma, Akshay Nagarajan, Rishikaa, Gauri Khanna and Annanya Mehan, Advocates, Adarsh Ramanujan with Parth Singh, Advocate, Amicus Curiae, Professor Arul George Scaria, Amicus Curiae.

Join the discussion

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.