india s fair use decision

The Delhi High Court’s decision in ANI v OpenAI is a watershed moment for how India treats the use of news content to train large language models and it will shape the way both media companies and AI developers negotiate data access and licensing in the coming years. At a time when newsrooms worldwide are asking whether AI systems are quietly free riding on their work, this ruling tells Indian courts and policymakers that technical use of news for training can be treated very differently from copying news for readers.

The case that put AI training under the spotlight

In late 2024, Indian news agency ANI sued OpenAI in the Delhi High Court, accusing the company of scraping and storing ANI’s copyrighted news reports to train ChatGPT without permission or payment. ANI asked for an injunction to stop any use of its content, along with damages of around two crore rupees, arguing that OpenAI’s model development relied on systematic copying of its news output.

The case quickly became India’s most closely watched generative AI lawsuit because it directly raised the question many creators have been asking in private: How far can an AI company go in using publicly available content for training before it crosses the line into infringement? The court treated the dispute as a test of Indian copyright law in the era of machine learning, framing core questions about storage of news data for training, the nature of AI outputs, and whether fair dealing could stretch to cover this new kind of use.

For much of 2025 and early 2026, the court heard technical and legal arguments from both sides and from appointed experts, before finally delivering a judgment that clearly signals where Indian doctrine stands right now on AI training.

What ANI argued and why it mattered for news organisations

ANI’s complaint was straightforward. It claimed that OpenAI had copied its news stories at scale, stored them in systems used to train ChatGPT and similar models, and monetised the results without any licence or compensation. From ANI’s perspective, this was not a small or incidental use but a core input into a commercial AI product that is sold and integrated into services used worldwide.

ANI alleges OpenAI systematically exploited its news reports to train and monetise ChatGPT without licence or compensation

The news agency emphasised that India’s fair dealing provisions in Section 52 of the Copyright Act were drafted for limited purposes such as private research, criticism, review and reporting current events. In ANI’s view, those exceptions did not extend to systematic scraping, copying and processing of protected works to build and refine commercial AI systems serving a global market.

Commentators defending ANI’s position argued that allowing training under existing fair dealing rules would hollow out copyright protection for news by converting large scale, unlicensed ingestion into something treated as neutral technical processing. They warned that this could erode the economic base of newsrooms that invest heavily in gathering and verifying information, while tech firms capture value from that content in new products without sharing revenue.

How the Delhi High Court viewed OpenAI’s use of news

The High Court rejected ANI’s core claim that using its reports for training ChatGPT amounted to copyright infringement under current Indian law and instead located OpenAI’s conduct within the fair dealing exception for research. The judge distinguished between expressive uses that serve readers and viewers and non-expressive technical uses that underpin how machine learning systems learn patterns from data.

A central part of the reasoning was the concept of non-expressive use. The court accepted that converting text into numerical vectors and extracting statistical relationships is a form of analysis that does not require reproducing the protectable expression of the original works in the outputs of the model. Because the trained system responds to prompts by generating new sequences rather than simply republishing ANI’s articles, the training process was cast as intermediate technical processing.

Crucially, the court found no substantial similarity between ChatGPT’s typical responses and ANI’s news stories. The judgment noted that the outputs did not replicate ANI’s reporting in a way that would count as reproduction or adaptation under Section 51, the provision that defines infringement. Without clear evidence that the system routinely regurgitated ANI’s articles, the court concluded that the use of ANI’s corpus in training did not by itself cross the threshold into actionable copying.

In parallel, the court treated ingestion and storage of news articles as intermediate steps in a research process rather than as commercial republication aimed at the same audience and market as ANI’s original reports. This framing allowed the judge to keep a clear line between AI model development and the business of publishing news, even though both can rely on the same underlying content.

Fair dealing, research and the idea of non-expressive use

The court anchored its analysis in Section 52(1)(a) of the Copyright Act, which permits fair dealing for private or personal use, research, criticism, review and reporting current events. Reflecting this approach, Justice Amit Bansal held that OpenAI’s storage of ANI’s works was non-infringing because it fell within Section 52(1)(a)’s fair dealing exception for research. Within that list, the judge placed AI training squarely in the research category, treating large scale data analysis for model development as a legitimate research activity even when carried out by a commercial technology company.

This view aligns with a growing strand of Indian academic and policy commentary that treats text and data mining for machine learning as non-expressive use that falls outside the core exploitation copyright is designed to regulate. Under this approach, what matters most is not that the material was copied at some point but whether the end product substitutes for the original works or competes directly for the same readership and market.

By accepting that AI training uses protected material in order to extract patterns rather than to serve the same audience with similar content, the court endorsed a more technology-aware understanding of fair dealing in the Indian context. At the same time, the judgment stands in clear tension with earlier arguments that Indian fair dealing is narrow and exhaustive, leaving little space for new categories like AI training unless Parliament expressly creates them.

This is where the decision shows its significance for doctrine. India does not have an open-ended fair use clause like the United States but instead lists specific purposes for fair dealing. Treating AI training as research pushes the boundaries of those purposes without formally adding a new category, signalling that courts are willing to interpret existing language through the lens of new technologies.

What this means for OpenAI and other AI developers

OpenAI has long defended the use of publicly available internet content including news articles as a legitimate foundation for building large language models, arguing that training produces new systems rather than competing publications. In India, the Delhi ruling effectively endorses that characterisation, at least where there is no clear evidence of models recreating specific articles in their outputs.

For OpenAI and other companies training models on mixed web data, the judgment provides short-term legal comfort in India. It signals that courts are prepared to see training as a research activity that falls within fair dealing, provided the outputs do not regularly mirror or substitute the original protected works. This aligns with the UK government’s recent projection of £45 billion in potential productivity savings through AI implementation, highlighting the broader economic implications of AI advancements.

However, the ruling does not give developers a blank cheque. It draws a line between non-expressive training and expressive outputs that could, in some scenarios, cross into reproduction or adaptation. If future evidence shows models reproducing full articles or closely paraphrasing reporting that is behind paywalls, courts could take a different view about specific uses or outputs even while keeping the general training process within fair dealing.

The impact on newsrooms and media business models

From the perspective of ANI and other news organisations, the decision is a mixed outcome. On the one hand, it confirms that current Indian copyright law offers limited tools to challenge AI training that uses news content as raw material, at least when the training is framed as research. On the other hand, it keeps open broader policy debates about licensing and remuneration that could rebalance how value flows between media and technology.

Newsrooms that hoped for a clear judicial declaration that unlicensed AI training is infringement will be disappointed. The ruling suggests that protecting revenues cannot hinge solely on stopping training but may instead require contractual solutions, collective bargaining, or future legislative schemes that create payment obligations.

At the same time, the case has made visible to editors and publishers how central their content has been to the development of powerful language models. It has moved the discussion from abstract worries about scraping to concrete courtroom exchanges about data pipelines, storage practices, and the kinds of uses that courts are prepared to treat as research. This visibility could strengthen the hand of news organisations in negotiations with AI platforms, even if the legal baseline today favours training under fair dealing.

The Delhi High Court’s analysis places India in an interesting position in the global conversation about AI training and copyright. Unlike the European Union, which has adopted specific text and data mining exceptions with conditions and opt-out mechanisms, India still relies on general fair dealing provisions without any express reference to AI or data mining.

Commentary around the ANI case has noted that European law explicitly addresses machine learning uses of text through dedicated exceptions, while India’s framework leaves courts to stretch existing language like research in order to cover similar activities. This contrast helps explain why the Delhi court leaned heavily on the non-expressive use concept, which parallels thinking in other jurisdictions even though India’s legislation has not yet been updated.

Compared with the United States, where fair use allows judges to weigh transformative purpose more flexibly, India’s decision shows a more cautious yet adaptable path. Courts are not rewriting the statute but are reading research in a technologically informed way that gives AI developers room to operate while signalling that outputs remain subject to conventional infringement analysis.

Policy ideas on the horizon licensing and AI training exceptions

Importantly, the judgment does not close the door on policy reform. Indian analysts and advocacy groups have already floated ideas like mandatory blanket licensing schemes for AI training or explicit statutory exceptions combined with remuneration rights for creators. Such models would recognise training as a legitimate activity while ensuring media organisations share in the economic value generated by AI products that rely heavily on their content.

Proposals inspired by European text and data mining rules suggest one possible direction: creating clear AI training exceptions while giving rights holders ways to opt out or claim compensation. Others point to collective licensing arrangements that could let AI firms access curated news archives under standardised terms, reducing transaction costs while ensuring money flows back to newsrooms.

The Delhi ruling indirectly increases pressure on policymakers. By treating training as fair dealing research under current law, the court has made it less likely that litigation alone will secure ongoing payments from AI firms to news outlets. That reality may accelerate calls within India for legislative updates that explicitly address AI training and its economic impact on journalism and other creative industries.

What this judgment does not resolve

Despite its significance, the decision leaves several important questions open. It does not fully settle how Indian law will treat situations where users intentionally prompt AI systems to reproduce protected articles or other content word for word. Some commentary suggests that such outputs could be analysed under fair dealing for private research, but courts have not yet developed clear rules for adversarial prompting and regurgitation.

Nor does the judgment definitively answer how Indian law will respond if AI systems are integrated into products that systematically summarise, paraphrase, or repackage news in ways that substitute for reading original reports from media outlets. Those scenarios raise distinct competition and market harm concerns that may require a different legal and policy toolkit than the one used to assess pure training.

Finally, the ruling does not address broader questions about transparency in training data, including whether companies like OpenAI should disclose more detailed information about exactly which sources feed their models and on what terms. For trust and accountability, many media and policy voices will continue to push for clearer disclosure regimes, even if fair dealing currently covers the core training activity.

Key takeaways and what to watch next

The Delhi High Court has delivered the first major Indian judgment to say that using news content to train large language models can fall within fair dealing for research rather than constituting copyright infringement. That conclusion rests on treating AI training as non-expressive technical processing and on the absence of substantial similarity between typical model outputs and the original news articles.

For AI developers, the ruling provides a degree of legal stability in India, confirming that they can continue training on publicly available news data, while remaining cautious about outputs that might edge into reproduction or adaptation. For newsrooms, it crystallises a hard truth under current law: stopping training outright is unlikely without legislative change, which shifts the focus toward licensing models, remuneration schemes and strategic partnerships.

Over the next few years, the most important developments are likely to come not just from courts but from Parliament and regulators as they weigh options like explicit AI training exceptions, blanket licences and new bargaining frameworks between media and technology platforms. The ANI v OpenAI judgment sets the doctrinal baseline for that debate and ensures that India enters the global AI copyright conversation with a clear judicial position on non-expressive use and fair dealing in the age of generative models.

Conclusion

India’s decision that OpenAI’s use of news content for training its models counts as fair dealing is a turning point for how artificial intelligence and copyright will coexist in one of the world’s largest digital markets. It offers short term comfort to AI developers while raising serious strategic questions for news publishers and lawmakers who now have to decide whether to accept this direction or push for stricter rules.

How India arrived at this moment

For years, India’s Copyright Act of 1957 has sat largely unchanged while artificial intelligence evolved from a research curiosity into a core infrastructure for search, productivity tools and media products. The Act gives authors exclusive rights over reproduction and adaptation of their works in Section 14 and defines infringement in Section 51, but it also carves out exceptions in Section 52, including fair dealing for private or personal use, research, criticism, review and reporting of current events.

Unlike the open ended fair use doctrine in the United States, Indian fair dealing has historically been interpreted narrowly, with courts focusing closely on purpose and limiting the exception to specific categories like research and news reporting. Legal scholars have repeatedly warned that the statute was never drafted with large scale automated copying and machine learning in mind.

The tension finally crystallised in the case brought by Asian News International against OpenAI over the use of ANI’s news content in training ChatGPT. ANI argued that systematic ingestion of its news feed to power a commercial chatbot went far beyond the limited purposes that fair dealing was meant to cover and asked the Delhi High Court to restrain OpenAI from using its material. For months this case was seen as India’s first serious test of how the Copyright Act applies to generative AI systems.

What the Delhi High Court actually decided

On July twenty four 2026 the Delhi High Court refused to grant ANI an interim injunction and held, on a prima facie view, that OpenAI’s training on ANI’s content amounted to fair dealing rather than copyright infringement. The court concluded that storing and using ANI’s original literary works to train large language models underlying ChatGPT fell within the fair dealing exception in Section 52 and therefore did not infringe under Section 51.

Justice Amit Bansal emphasised that the acts at issue involved storing works for the technical training of models and that this use was covered by fair dealing for research and related purposes. He further observed that the outputs generated by ChatGPT were not substantially similar to ANI’s articles and therefore did not themselves amount to infringement under Section 51.

Crucially, this is an interim order in an ongoing suit, not a final merits judgment. The court set an initial framework to allow generative AI operations to continue in India while the main case proceeds, signalling how judges are likely to approach similar disputes but leaving room for later refinement or reversal.

Fair dealing as technical non expressive use

The heart of the ruling is the way it frames AI training as a technical, largely non expressive use of copyrighted works rather than a competing publication that substitutes for the original. In other words, ingesting articles into a training dataset is treated as a step in building a tool, not as an act of public communication of those articles themselves.

This approach echoes arguments that have been gaining traction in other jurisdictions. The United States Copyright Office, in a recent report, described a spectrum in which non commercial research uses that do not enable reproduction of expressive portions in outputs are more likely to qualify as fair use, while copying expressive works from pirated sources to generate competing content in the marketplace is less likely to be lawful. Legal commentators have described the training step as transformative because it uses works as raw material to extract statistical patterns rather than to re present their original expression.

European and United Kingdom debates around text and data mining also turn on similar distinctions, with scholars noting that existing exceptions can sometimes cover automated analysis but may not neatly fit unlicensed AI training that involves extensive reproduction beyond the limits of those exceptions. India does not yet have dedicated text and data mining provisions, which makes the Delhi court’s choice to stretch fair dealing to cover AI training especially significant.

Why developers are reassured

For AI companies operating in or serving the Indian market, the ruling gives immediate practical relief. It confirms that training models on publicly available news content is not automatically considered infringement and can fall under fair dealing, at least at the interim stage and subject to the specific facts of the case.

This matters for several reasons. First, the cost of obtaining licences for every piece of training data from every news organisation would be prohibitive for most developers, particularly startups that drive much of the innovation but lack deep cash reserves. Second, India is positioning itself as a global hub for AI research and deployment, and an early finding that core training practices are prima facie lawful reduces regulatory risk for investors and local firms.

The decision also fits within a wider set of Indian judicial and policy moves that cautiously accept AI as a tool while insisting on human oversight. The Supreme Court has issued draft regulations for AI use in courts that explicitly forbid reaching judicial outcomes solely on the basis of AI generated information, while permitting AI for research, citation verification, summarising and translation under human supervision. In another recent case the Supreme Court quashed a tribunal ruling that had relied on fabricated AI generated precedents, describing such use as dangerous to the integrity of the legal system. Those steps show that Indian authorities are willing to integrate AI where it augments human expertise but are alert to its risks.

Why publishers are unsettled

For news organisations, particularly digital agencies whose content is constantly scraped and indexed, the ruling lands very differently. ANI and supporting scholars argue that fair dealing in India was never meant to cover systematic copying and storage of large volumes of copyrighted material for commercial AI training. Section 52 focuses on private or personal use, research, criticism, review and reporting of current affairs, and none of these categories obviously describe a global technology firm building a paid chatbot powered by a vast corpus of news.

From this perspective, the decision appears to expand the fair dealing exception without legislative change, effectively creating a broader training privilege by judicial interpretation. Critics worry that this undermines the incentive for publishers to invest in newsgathering if their material can be ingested at scale with minimal negotiation, especially when AI tools may divert traffic and advertising revenue away from original publishers.

The court’s finding that ChatGPT outputs are not substantially similar to ANI’s content offers some comfort but does not fully resolve concerns about market harm. Even if no single answer reproduces an article verbatim, a system that can summarise and contextualise the news may reduce the need for users to visit the original site, a factor that many copyright frameworks treat as highly relevant when assessing fairness.

India’s interim stance now sits alongside a patchwork of approaches worldwide. In the United States, litigation and regulatory analysis are converging on the idea that training on lawfully acquired content may often be fair use, particularly where outputs do not substitute for the original works, but copying from pirated sources or enabling direct market competition could fall outside the defence. Libraries and academic groups have advocated strongly for recognising the ingestion of copyrighted works to build AI models as fair use, arguing that it continues a long tradition of non expressive indexing and search technologies.

In Europe and the United Kingdom, lawmakers have introduced specific text and data mining exceptions but there is ongoing debate about whether they cover unlicensed AI training or whether rights holders can opt out, forcing developers into licensing negotiations. Scholars point out that AI training often engages acts of reproduction and storage that go beyond what these exceptions explicitly allow.

Against that backdrop, India’s reliance on a general fair dealing provision rather than a dedicated text and data mining rule looks both pragmatic and fragile. It gives courts flexibility to adapt existing law to new technologies but leaves businesses unsure how far the logic will stretch when different types of content, purposes or market effects are involved.

Unanswered questions and emerging risks

The Delhi High Court has taken a clear step, but many important questions remain open. The interim order does not answer in a definitive way whether all large scale AI training on news content qualifies as fair dealing or whether certain practices might cross into infringement, especially if outputs begin to resemble original works more closely.

Several fault lines are already visible in commentary and scholarship. One is the commercial nature of AI services. Critics argue that fair dealing for research was intended for activities such as academic inquiry or private study, not for global for profit platforms that monetise access to information and analysis. Another is the scale and automation of copying. Human researchers quoting portions of news for criticism or review operate within relatively clear norms, but continuous machine ingestion of entire feeds presents a qualitatively different scenario.

There is also the question of bargaining power. Large technology firms have the resources to negotiate licences when needed, while smaller publishers, especially in regional languages, may struggle to assert their rights if courts adopt an expansive view of fair dealing. Over time that could widen inequalities in whose content is used and whose interests are protected.

Beyond copyright, the ruling interacts with concerns about misinformation and synthetic media. India has already seen the judiciary refuse to rely on AI generated case law and insist that human judges remain the final decision makers. If AI systems trained on news become central gateways to information, errors or bias in those systems could shape public understanding in ways that traditional publisher models did not, which makes the question of accountability more pressing even when training is lawful.

What to watch next

The ANI v OpenAI case is still live, and the interim order is likely to be followed by detailed arguments and possibly appeals that probe the boundaries of fair dealing more closely. Lawmakers and regulatory bodies such as the Copyright Office will have to decide whether to codify a clearer position on AI training, perhaps by introducing text and data mining provisions or more explicit guidance on when large scale data use is acceptable.

News organisations may respond by strengthening contractual controls over their feeds, investing in paywalls or technical measures, or seeking collective licensing arrangements that give AI companies a straightforward way to obtain permission and share value. At the same time, AI developers will likely refine their data governance practices, documenting sources more clearly and avoiding reliance on pirated or dubious datasets to reduce legal exposure.

Internationally, India’s experience will feed into a broader conversation about how emerging economies can both foster innovation and protect local creative industries. Courts and regulators elsewhere will watch closely to see whether the Indian approach delivers practical stability or triggers further disputes that reveal its weaknesses.

Key takeaways and forward looking insights

  1. The Delhi High Court has signalled that training AI models on news content in India can fall within fair dealing, treating ingestion as a technical use rather than straightforward infringement, at least on an interim view.
  2. This stance offers short term reassurance to AI developers and investors but creates unease for publishers who see a broader exception being read into a statute that was originally drafted for limited research and reporting uses.
  3. The ruling places India in the middle of a global debate where the United States, Europe and others are still drawing lines between permissive training and uses that cause market harm or rely on unlawful sources.
  4. The case leaves major questions open about how scale, commercial intent and output similarity should be weighed in future disputes, which means businesses should treat the current position as a guidance signal rather than a final settlement.
  5. Over the next few years, expect a mix of judicial refinement, possible legislative or regulatory clarification, and new licensing and technical practices as both AI firms and publishers adjust to a world where large scale data use is increasingly central to innovation but no longer assumed to be legally unproblematic.
You May Also Like

More Than 200 Economists and Technology Leaders Call for Stronger AI Regulation

Sweeping demands from over 200 economists and tech leaders for stronger AI regulation reveal urgent concerns that could reshape the future of technology worldwide.

AI Agent Testing Crisis: Why Enterprise Autonomy Is Outpacing Safety Evaluations

Powerful AI agents are autonomously making critical enterprise decisions, but the safety evaluations meant to govern them are dangerously falling behind.

Universities Drop AI Detectors Over False Results Reddit

I uncover why universities are abandoning flawed AI detectors after false accusations, and how Reddit debates hint at the next big integrity crisis.

Anthropic Lobbying Spending Surpasses Nvidia as AI Regulation Battle Intensifies

Hurtling past Nvidia in DC spending, Anthropic reshapes the AI rulebook—and the next move in this high-stakes regulatory fight isn’t theirs.