Categories
Article

Does the Use of Copyrighted Content to Train AI Models Violate the Law?

This is a guest blog post written by Nouf Al-Hoqani, Rayan Al-Hasni, Zahra Al-Balushi, and Zayd Al-Harrasi as part of their Decree Fellowship group project in July 2026.

Artificial intelligence has advanced significantly in recent years. AI systems can now generate text, compose music, and create images and other works. These advancements have also raised significant legal concerns, particularly regarding intellectual property rights. AI developers often use copyrighted materials to train AI models, raising the risk that these models will create counterfeit or derivative works belonging to other people. This issue has not yet been addressed under Omani law: there is no specific legal provision or regulation governing the relationship between AI and intellectual property. By contrast, jurisdictions such as the United Kingdom, the United States, and the European Union have already begun to grapple with the problem and to explore potential solutions. Oman should therefore begin addressing this issue, drawing on the experience of these jurisdictions to inform its own approach.

The connection between artificial intelligence and intellectual property arises from AI’s growing capacity to create works similar to those made by human artists, such as images and music. This capacity makes it easier to reproduce the works of well-known artists and imitate their distinctive styles, raising the question of whether such conduct amounts to intellectual property theft or should instead be regarded as merely drawing inspiration from existing works.

The Omani Legal Framework

When examining Omani law at the intersection of artificial intelligence and intellectual property, the most prominent issue is the infringement of intellectual property rights by AI, and how such infringement, though contrary to law and ethics, has become so simplified and widely accessible that it is now available to anyone with the click of a button. This section examines how Omani law treats AI activities that rely on imitating or using copyrighted materials, guided by a single question: does Omani law treat the training of AI systems on copyrighted material as a violation of the law, or as a permissible exception?

Training an AI system involves compiling a large database of the data on which the model is trained, including words, letters, shapes, patterns, and colours, drawn from sources such as websites, books, articles, images, videos, music, and other works. Once this data is collected, it is presented to the system, which learns to recognise statistical patterns in it, analyse them, and generate similar patterns or predict the most likely next one through repeated exposure. Since this process requires assembling a dedicated database, the developer must first identify the data to include, download or copy it, and then store that copy electronically for later use in generating derivative outputs.

These steps matter because the Copyright and Neighbouring Rights Law, reserves the economic rights in a work to its author. Article 6 grants the author the right to reproduce the work, one of the most significant rights the law protects, as well as the right to adapt it into other forms, create derivative works, and dispose of the work in both its original and copied forms. Article 1 defines reproduction broadly as making one or more copies identical to the original, whether directly or indirectly, “by any means such as printing, photocopying, recording, or permanent or temporary electronic storage.” This definition captures precisely what AI training involves: the electronic storage of copyrighted works, a step that is fundamental to building a training database but is reserved exclusively to the author unless the author grants that right to another party by agreement.

There are, in principle, two ways an AI developer might avoid infringing copyright in this process. The first is to obtain the author’s consent to use the work for training purposes; although the law does not address this scenario explicitly, such permission would, as in other contexts, allow the work to be used lawfully. The second is for the training process to fall within Chapter Five of the law, which sets out the free uses of works. Article 20 lists uses that do not require the author’s consent, including use for explanation or critique, educational and informational purposes, copying by archives or public libraries, and use to illustrate a concept in a study, provided that certain conditions relating to the quantity used, the manner of use, and the absence of any direct or indirect financial gain are met. Article 20 makes no mention of AI or its training on copyrighted material. We therefore conclude that training an AI system without the author’s permission, and without relying on works in the public domain, constitutes a clear violation of copyright under Omani law. This gap, the complete absence of any law or regulation addressing AI’s use of copyrighted material, creates considerable uncertainty about how Omani law will respond to these issues as the technology continues to advance.

Comparison with Other Jurisdictions

The European Union offers a significant comparative model for Oman, having been the first jurisdiction to establish a comprehensive legal framework for artificial intelligence through the Artificial Intelligence Act (Regulation (EU) 2024/1689). The Act creates a framework intended to build trust in AI technology while protecting human rights and safety. Although it permits text and data mining for the training of general-purpose AI (GPAI) systems, this mechanism is subject to strict conditions and does not give AI developers a free pass to use copyrighted data. The Act requires generative AI providers to comply with existing EU copyright law, imposes transparency obligations regarding the content used to train AI models, and gives rights holders the option to opt out of having their content used for training.

Because GPAI providers require large datasets that may contain copyrighted material, questions arise over whether such use might constitute infringement. The EU addresses this largely through Directive (EU) 2019/790 on Copyright in the Digital Single Market, which introduced text and data mining exceptions under Articles 3 and 4. Article 3 permits research organisations to use protected content lawfully but excludes commercial or industrial uses. Article 4 allows other institutions to reproduce and extract data, subject to an opt-out mechanism that allows rights holders to exclude their work by ‘machine-readable’ means. What qualifies as machine-readable has been contested, notably in the German case Kneschke v LAION, in which a non-profit organisation used Kneschke’s copyrighted content to build an AI training dataset. The court rejected the copyright claim on the basis that the use fell within the text and data mining exception for scientific research, and held that a reservation expressed only in ordinary language was not sufficiently machine-readable. An appeal is pending.

Article 53(1)(d) of the Act further requires generative AI providers to publish a sufficiently detailed summary of the content used to train their models, and Recital 107 explains that this transparency requirement is intended to support copyright holders in exercising their rights. Even so, uncertainty remains over how far the existing text and data mining exceptions extend to AI training, and EU member states continue to debate whether the current framework adequately addresses the scale and complexity of the practice.

Most member states nonetheless favour monitoring and clarifying the existing framework rather than introducing new legislation immediately, given the continued novelty of generative AI. This cautious approach is instructive for Oman, which may similarly benefit from clarifying and monitoring its existing copyright principles rather than enacting an entirely new framework at this stage.

The United States has not enacted a comprehensive federal AI statute; regulation instead derives from a mix of executive orders, existing sectoral laws applied to AI, and state legislation. In Thomson Reuters v Ross Intelligence (2025), the court rejected a fair use defence where Ross had engaged a third party, LegalEase, to produce training data that substantially copied headnotes from Thomson Reuters’ Westlaw platform. The court found that Ross had directly copied thousands of these headnotes and rejected fair use primarily because Ross intended to use the resulting AI tool to compete directly with Westlaw, a factor the court held weighed decisively against fair use.

By contrast, in Bartz v Anthropic (2025), Anthropic had trained its Claude models using a mix of purchased and pirated books to build a permanent digital library, arguing that the books were essential to training its models. The court found that Anthropic’s use of purchased books constituted fair use, but that its use of pirated copies did not.

The US approach is therefore highly fact-specific, with outcomes varying case by case, and indicates that fair use may apply, but only within certain limits. The EU, by contrast, takes a legislative approach through the text and data mining exception in Directive (EU) 2019/790.

Jurisdictions aside from the EU and the USA have taken different approaches. The United Kingdom, for example, has no broad copyright exception permitting commercial AI training on protected works, although proposals for a text and data mining exception remain under discussion.

These divergent approaches show that there is no settled international consensus on whether copyrighted content may be used to train AI systems. Oman therefore has no single international model to follow and should instead weigh the interests of copyright holders against the goal of supporting AI development in determining its own approach.

Recommendations

The existing exceptions under Article 20 are tied to non-commercial, educational, or family contexts. We recommend amending Article 20 to introduce a new AI training clause permitting commercial entities to use copyrighted content for AI training, provided they have lawful access to that content, whether through licensing, subscription, or other authorised means. This carve-out is necessary because AI development in Oman is largely a commercial activity; without it, a company would remain excluded from the exception even where it has lawful access to the content it seeks to use. At the same time, original creators may face heightened competitive risk, as AI-generated content trained on their work could saturate the market with similar output and reduce demand for their future work.

Oman may also wish to adopt a gradual approach to regulating AI training on copyrighted content, rather than introducing a comprehensive AI-IP framework immediately, given that the technology remains relatively new and not yet fully understood. Instead, Oman should clarify the existing copyright law to specify the circumstances under which AI training can use protected work without infringing copyright, potentially through a text and data mining exception modelled on the EU approach, paired with an effective opt-out mechanism allowing copyright holders to reserve their work from AI training. Oman should also impose transparency obligations requiring AI developers to disclose the sources and content used to develop their systems, strengthening copyright holders’ ability to identify and enforce their rights without imposing an outright prohibition on the use of their content for AI training. Given the rapid development of generative AI and the current uncertainty in copyright law, these recommendations would allow Oman to protect copyright holders’ interests while continuing to encourage technological innovation.

These reforms also carry risks. A broad text and data mining exception could weaken copyright protection by allowing developers to use large quantities of copyrighted material, reducing copyright holders’ control over their work and its economic value. An opt-out mechanism may be difficult to enforce where ownership is unclear or content originates outside the country. Strict transparency requirements could impose high compliance costs on AI providers, potentially discouraging international companies from operating in Oman. There is also a risk that legislating before international approaches have stabilised could produce requirements that quickly become outdated.

Conclusion

Despite the risks of reform, the absence of regulation carries the greater risk: legal uncertainty. Without clear governance over whether copyrighted content can be used to train AI, both copyright holders and system developers face uncertainty in determining whether their conduct is lawful, which in turn complicates innovation in Oman. A carefully defined text and data mining exception would not eliminate copyright protection; rather, it would establish predictable circumstances in which works may be used, while preserving authors’ right to opt out. The goal should not be to eliminate all risk, an unrealistic aim, but to regulate use and provide greater certainty while respecting copyright holders’ rights. This is particularly important if Oman seeks to attract AI investment and build a competitive digital economy.

Authors
Nouf Al-Hoqani
University of Manchester, United Kingdom

Rayan Al-Hasni
Sultan Qaboos University, Oman

Zahra Al-Balushi
Modern College of Business and Science, Oman

Zayd Al-Harrasi
Nottingham Trent University, United Kingdom