The Verdict and Its Immediate Context
On July 25, 2026, the Delhi High Court ruled that OpenAI’s use of copyrighted works to train its generative AI models does not qualify for a fair‑dealing exemption. The decision marks the first major judicial assessment of AI training under India’s copyright law, intensifying the clash between tech firms and content creators. The ruling threatens the operations of more than 30 industry groups that have collectively sued OpenAI and may force AI developers to secure licenses for billions of copyrighted text snippets.

- •Supreme Court Ruling on AI Training: Fair Dealing Triumph or Copyright Conundrum
Supreme Court Ruling on AI Training: Fair Dealing Triumph or Copyright Conundrum
The Delhi High Court on 26 May 2026 pronounced that OpenAI’s use of copyrighted works to train its large‑language models (LLMs) falls within the “fair dealing” exception of the Copyright Act 1957. Justice Amit Bansal’s observation marks the first time an Indian constitutional court has invoked fair use in the context of generative AI, a decision that could reshape the balance between creators’ rights and the burgeoning AI ecosystem.
Justice Bansal held that the act of storing the Association of National Institutes’ (ANI) original retrieved works for training LLMs “does not amount to infringement” because it is covered by Section 52(1)(a) of the Copyright Act. The ruling emerged from a petition filed by a coalition of rights‑holders, including the Federation of Indian Publishers, the Digital News Publishers Association, and the Indian Music Industry.
- ▸The judgment was delivered on 26 May 2026 in open court.
- ▸Petitioners argued that AI training amounted to industrial‑scale theft of copyrighted material.
- ▸The bench cited Section 52(1)(a) as the statutory basis for a fair dealing defence.
- ▸This is the inaugural application of a fair‑use‑type exception by India’s highest judicial forum.
- ▸Justice Bansal’s pronouncement specifically referenced the use of ANI’s works by OpenAI.
How Large Language Models Are Trained
LLMs such as ChatGPT ingest massive corpora scraped from the internet—books, news articles, images, and music—to learn statistical patterns of language. A recent technical advance, Retrieval‑Augmented Generation (RAG), combines this pre‑training with real‑time retrieval of external documents, improving factual accuracy but also raising new copyright questions. The process can be broken down into three stages: data collection, model training, and output generation.
- ▸Data collection involves crawling publicly accessible web pages, often without explicit licences.
- ▸During training, the model “memorises” token sequences, enabling it to reproduce phrasing that resembles source material.
- ▸The output stage can produce “AI hallucination” where the model fabricates information not present in the training set.
- ▸RAG systems retrieve documents at query time, blurring the line between input and output usage.
- ▸The scale of training data now reaches petabytes, dwarfing the manual curation possible for human scholars.
Did You Know? The term “hallucination” in AI was coined by researchers at OpenAI in 2020 to describe instances where the model generated plausible‑sounding but factually incorrect statements.
Copyright Law Meets Generative AI
The Copyright Act 1957 provides a limited “fair dealing” exception for purposes such as research, criticism, review, reporting, or teaching. Section 52(1)(a) enumerates these purposes, aiming to balance the creator’s exclusive rights with public access to knowledge. In the AI context, the law distinguishes between two potential infringements:
- ▸Input stage – copying copyrighted works to build a training dataset.
- ▸Output stage – the model’s generation of text that may replicate protected expression.
Indian jurisprudence has traditionally treated the input stage as a “copy” requiring permission, whereas the Supreme Court’s present ruling treats large‑scale, non‑commercial data ingestion as permissible research under fair dealing.
- ▸Fair dealing permits “limited use” without consent for the enumerated purposes.
- ▸The Indian exception is narrower than the U.S. “fair use” doctrine, which includes a four‑factor test.
- ▸The judgment emphasized that the purpose of training LLMs aligns with “research” as defined in the Act.
- ▸No explicit licence was required from rights‑holders for the data used in this case.
- ▸The decision does not immunise the output stage; infringing reproductions remain actionable.
International and Policy Responses
While India grapples with the legal frontier, other jurisdictions have taken divergent paths. The United States leans on the “fair use” defence, whereas the European Union is moving toward a mandatory licensing regime for AI training data. Domestically, the Department for Promotion of Industry and Internal Trade (DPIIT) released a working paper titled “One Nation One License One Payment,” proposing a hybrid model that obliges a blanket licence for all copyrighted works used in AI training, irrespective of consent.
- ▸The paper recommends a mandatory blanket licence, rejecting a voluntary model.
- ▸It envisions a single‑payment system administered by a central copyright collective.
- ▸The proposal aims to avoid “licence fatigue” for AI developers while ensuring remuneration for creators.
- ▸The EU’s forthcoming AI Act includes provisions for “data‑access obligations” on high‑risk AI systems.
- ▸The U.S. Copyright Office is currently reviewing a rulemaking on AI‑generated works.
Implications for Innovation and Media
The Supreme Court’s endorsement of fair dealing for AI training could lower barriers for Indian startups, fostering a homegrown generative‑AI sector. However, media organisations fear that the ruling may erode incentives to produce original content if large models can replicate large swathes of their work without compensation. The hybrid licensing model under consideration seeks to reconcile these competing interests by guaranteeing a baseline royalty while preserving the research exemption.
- ▸Start‑ups can now train models on publicly available Indian literature without immediate licence costs.
- ▸Content creators may push for statutory amendments to tighten the definition of “research.”
- ▸A blanket licence could generate a new revenue stream for the publishing and music industries.
- ▸International investors are watching India’s regulatory stance as a barometer for AI‑friendly jurisdictions.
- ▸The decision may influence future disputes involving other generative‑AI tools such as image generators and code assistants.
The ruling thus sits at the intersection of technology, law, and economics, signalling a pivotal moment for India’s digital future.
Tags
Concepts Mentioned
AI hallucination
AI hallucination is when a generative model produces statements that sound plausible but are factually incorrect or unfounded. This phenomenon matters because it can mislead users and undermine trust in AI systems, especially in critical applications. For example, a language model once fabricated a non‑existent scientific study when asked for references.
ChatGPT
ChatGPT is an advanced conversational AI model developed by OpenAI that generates human‑like text responses based on vast language data. It has reshaped how people interact with machines, enabling applications from customer support to creative writing. For example, it can draft a coherent essay on a given topic within seconds.
OpenAI
OpenAI is an artificial‑intelligence research organization that develops advanced machine‑learning models and tools. It is significant for pioneering safe, broadly beneficial AI and influencing global AI policy and industry standards. A concrete example is ChatGPT, a conversational model released in 2022 that quickly amassed over 100 million users.
Digital News Publishers Association
The Digital News Publishers Association (DNPA) is a global trade body representing online news outlets and digital journalism platforms. It advocates for sustainable revenue models, press freedom, and industry standards in the evolving digital media landscape. In 2022 DNPA launched the Trustworthy News Seal, which has been adopted by more than 150 member sites.
Copyright Act 1957
The Copyright Act 1957 is a legislation that governs copyright laws in India, providing protection to creators of original literary, dramatic, musical, and artistic works. It is significant for safeguarding the rights of authors, artists, and creators, allowing them to control the use and distribution of their work. The Act protects works such as the famous novel 'Gitanjali' by Rabindranath Tagore.
Log in to like, comment, and join the discussion.