An Indian judge has ruled that when Open AI used content owned by news agency ANI to train ChatGPT, without getting any permission, it was covered by a ‘private research’ exception in Indian copyright law and therefore is not liable for copyright infringement.
Although Judge Amit Bansal was ruling on ANI’s bid for an interim injunction against Open AI, his judgement is another significant decision regarding the copyright obligations of AI companies in different countries around the world.
And it’s a win for those AI companies that claim they can train models with existing content without getting permission from creators and rightsholders because of exceptions in copyright law.
A ruling in relation to journalism and the written word does not necessarily set a precedent that would apply to all kinds of copyright-protected works, including music. However, Indian record industry trade group IMI did make a submission to the court as part of this legal dispute, clearly recognising that any ruling on copyright and AI has the potential to impact on the wider copyright industries.
In the US, tech companies claim that AI training is covered by the ‘fair use’ principle under American copyright law - meaning they don’t need to get permission from copyright owners.
However, in most other countries, including India, if an AI company wants to train a model on existing content without getting rightsholder permission, they need to identify a relevant specific copyright exception. And if they can’t find one, they need to lobby lawmakers to introduce one.
AI companies usually start by looking for an exception that covers ‘text and data mining’ - often referred to as a ‘TDM exception’ - based on the argument that an AI model is simply extracting meaning from text when it copies existing content. However, Indian copyright law doesn’t have a specific TDM exception.
So when ANI sued Open AI for copyright infringement, the AI company’s lawyers honed in on the Indian copyright exception covering “private or personal use, including research”.
Which, from the perspective of many Indian copyright owners, including those in the music industry, seemed somewhat ambitious, and a stretch of the meaning of that exception. But Judge Bansal disagreed and decided that that exception does indeed apply in this case.
ANI and its supporters argued that the ‘private research’ exception should only really apply to non-commercial uses of copyright protected works and should only really benefit private individuals, rather than one of the most valuable companies in the world. Plus there was also the question as to whether AI training was really “research”.
However, none of those arguments impressed Bansal. The equivalent copyright exception in UK law explicitly says it only applies to ‘non-commercial’ research, but Indian copyright law doesn’t include that restriction. And the word ‘private’, the judge concluded, doesn’t just mean the private individual, but rather whether the research is conducted privately
When Open AI took ANI’s content, it was stored “in a closed space without access to the public”, Bansal wrote, and was “not publicly available to any human entity either for access or for download”. Therefore, he added, “in my opinion, the use amounts to being purely private”.
And AI training which “involves machine learning of the stored literary works by screening and organising them”, and then analysing the data and “making extractions from the literary works and converting them into machine-readable training inputs”, sounds, he continued, like a form of research.
With all that in mind, Bansal did not grant ANI the interim injunction it had asked for against Open AI, which would have prevented the AI company from using ANI’s content in its training moving forwards. To what extent that decision has ramifications on the wider AI and copyright debate in India is currently less clear.
Given Open AI actually trained ChatGPT in the US, it would actually argue that its training processes are only really subject to US copyright law. Whether or not AI training is actually fair use under American law is at the centre of dozens of lawsuits and the copyright obligations of AI companies currently remains uncertain within the US.
But if ultimately the US courts decide AI training is not fair use, this ruling in India that AI training there is covered by the research exception under Indian copyright law, may provide an incentive for more AI companies to formally base their training operations in India.
Of course, the copyright industries - like the music industry - argue that where an AI model is actually trained is only one factor, and where the model is commercially exploited is just as important. Under that logic, if an AI model is trained without rightsholder permission in a country where that is allowed, it can’t then be exploited in countries where permission would have been required for the training.
However, that position is disputed in many countries and the copyright industries may need to get copyright rules rewritten for that requirement to be clear in law.
And even if that happens, rules of that kind might be easy to enforce when a model is trained for use as part of an end-to-end ecosystem like ChatGPT - where training, access, inference and payment fall under the same company - but would be harder to enforce for the increasing number of ‘open-weight’ models, including those trained in China, where users can download and run an AI model on their own infrastructure.
All of which means that, for now at least, rulings like this one involving Open AI in India are potentially important to all copyright owners all over the world.