https://doi.org/10.25547/BE2X-P388
This observation was written by Alan Colin-Arce with thanks to Brittany Amell, Simon van Bellen, and Jack Young for their comments and review.
At a Glance
| Topic | Enclosure and Commonswashing: Clarivate announces “Research Commons database” |
| Key Participants | Clarivate, OpenAlex, Crossref, Creative Commons, Wikimedia |
| Date | 2025 |
| Keywords | Digital Commons / Commun numérique, AI bots / Robots d’indexation IA, bibliodiversity / bibliodiversité, commonswashing |
Summary
This observation describes Clarivate’s introduction of the “Research Commons” database, which draws from data from OpenAlex and Crossref and adds it to a separate database in the Web of Science. This announcement is connected to recent trends seen in open scholarship, such as the corporate appropriation of knowledge commons and commonswashing, an appropriation of the language of the commons without contributing to them.
Introduction
On September 4, 2025, Clarivate announced the incorporation of a new database they are referring to as the “Research Commons” into the Web of Science platform. Research Commons is a “separate, comprehensive collection of journal output from open metadata sources” (Goldfinger, 2025), namely OpenAlex and Crossref. The addition of this collection is part of Web of Science’s expansion strategy to increase the global and disciplinary breadth of the Web of Science platform while keeping it separate from the core collection, a more restrictive database that only accepts journals that meet Clarivate’s editorial selection process. So far, Clarivate has incorporated 32 million records from the last 10 years into the “Research Commons” database.
The “Research Commons” database is a response to criticisms accusing Web of Science and its competitor Scopus of being biased against “research produced in non-Western countries, non-English language research, and research from the arts, humanities, and social sciences” (Tennant, 2020, p. 1). Therefore, drawing data from OpenAlex and Crossref expands the geographic, linguistic and disciplinary coverage of the Web of Science, as recent studies have found that OpenAlex offers more inclusive and comprehensive indexing in all of these areas (Céspedes et al., 2025; Maddi et al., 2025).
Researchers and stakeholders consulted by Clarivate praised the move for increasing the inclusion of researchers and journals in the Global South, promoting bibliodiversity, and expanding the Web of Science (Goldfinger, 2025). However, accessing the Research Commons database requires a subscription to the Web of Science platform, an expensive platform that charges up to 50,000 euros (around 81,600 CAD) in subscription fees (Clavey, 2024). Some universities worldwide, including Sorbonne University, Utrecht University, and Dalhousie University have cancelled their subscriptions to Web of Science due to its high costs or to support alternative open databases, such as OpenAlex.
Corporate Appropriation of the Commons
Despite the name, the “Research Commons” database does not align with widespread definitions of digital commons, which are understood as “a subset of the commons, where the resources are data, information, culture and knowledge which are created and/or maintained online. They are shared in ways that avoid their enclosure and allow everyone to access and build upon them” (Rosnay & Stalder, 2020, p. 2). The “Research Commons” database, in contrast, prevents equitable access to its data due to the high cost of both consulting Web of Science and accessing in a large scale the metadata associated with the indexed articles.
Other open scholarship initiatives that function as digital commons, such as Knowledge Commons, the Humanities and Social Sciences Commons, Wikimedia Commons, and Creative Commons meet this criteria of horizontal governance and equitable access provision. Therefore, the announcement of the Research Commons database is part of a broader trend of companies appropriating and profiting from the discourse and work of open scholarship initiatives, and commons initiatives in particular, without contributing to or supporting them.
The “Research Commons” database contributes to the enclosure of the commons by appropriating and privatizing data that was originally shared in open infrastructures such as OpenAlex and Crossref. This appropriation furthers Clarivate’s transition from a publisher or research platform provider to a data brokering company, as argued by Sarah Lamdan in the book Data Cartels. According to Lamdan (2022), data brokering companies monetize the research process by creating data products that draw insights from academic and other types of personal data. Following this argument, the introduction of the Research Commons database is a way of expanding Clarivate’s available data to analyze and then sell the insights back to its customers.
Other examples of corporate appropriation and enclosure of the commons can be seen in the increasing extraction of works with (and without) Creative Commons licences to train artificial intelligence models (a recent OSPO post discussed how this affects open access repositories in particular). As argued in a Creative Commons report, “large AI models are benefiting from the labor, money, and care that goes into creating and maintaining the commons while also threatening to ‘bleed it dry’” (Hardinges et al., 2025, p. 4). In response, Creative Commons is working on new licenses, called CC Signals to “offer a new way for stewards of large collections of content to indicate their preferences as to how machines (and the humans controlling them) should contribute back to the commons when they reuse and benefit from using the content” (Hardinges et al., 2025, p. 4). These licenses do not prevent AI development, but they seek to incentivize actions that maintain and support digital commons when using materials with Creative Commons licenses.
Wikimedia faced a similar issue of AI bots straining its servers to extract contents from Wikimedia projects, such as Wikipedia, Wikimedia Commons, and Wikibooks (Edwards, 2025). To reduce this strain, Wikipedia released a dataset based on Wikipedia formatted specifically for AI developers (Maxwell, 2025). This move was praised for promoting ethical data use and democratizing access to AI training data. However, Wikipedia users have also expressed concerns about the exploitation of the open web to create proprietary AI models (Woodcock, 2023).
Commonswashing
These cases demonstrate the recurring concerns around the enclosure and corporate capture of digital commons. Clarivate’s introduction of the “Research Commons” database is another case of a corporation appropriating the data of open infrastructures for profit. As Simon van Bellen (Senior Research Advisor, Érudit) states in his comments on the piece:
What may be bothering in this particular case is that a major company builds on the commons without any reciprocity. There is no reciprocity with Crossref or OpenAlex, perhaps not even a proper attribution as both Crossref and OpenAlex make their data available in the public domain with a CC0 license.
However, connecting this database with the concept of the commons is also a process of commonswashing, a term coined by Mélanie Dulong de Rosnay to describe how for-profit companies frame their activities under the umbrella of the commons to benefit from public sympathy, appropriating the message of the commons for commercial purposes without endorsing its principles nor keeping value within the academic community (Dulong de Rosnay, 2020).
Dulong de Rosnay (2020) lists some features shared among commons products or services: common ownership of the means of production, technical architecture and design governed by the community, social governance of work patterns, and common ownership of resources. In the case of “Research Commons”, the database is owned by Clarivate, a for-profit corporation, the technical design and the governance are centralized and managed by Clarivate (not the academic community), and the ownership of the database rests within Clarivate. Therefore, this new database cannot be considered a commons at all.
Commonswashing is detrimental to knowledge commons, as well as the commons movement more generally, because it is not only an appropriation of common resources, but of the social benefits and the imaginaries of the actual commons (Dulong de Rosnay, 2020). Therefore, commonswashing makes it harder to recognize open scholarship initiatives that are in fact guided by principles of the commons, such as Knowledge Commons, the Humanities and Social Sciences Commons, and some publishers from the Radical Open Access Collective.
Key Questions and Considerations
Clarivate announced plans to introduce new indicators for publications available across all Web of Science databases, including “Research Commons”. This proposed initiative might provide metrics to a broader selection of publications in research assessment, but it also risks furthering the outsourcing of research evaluation to Web of Science (Pinfield, 2024), increasing the dependence on a single database, and commodifying an increasing number of publications that were not previously quantified in the Web of Science.
Some additional concerns on the commercialization of open data and open infrastructures were raised by Simon van Bellen:
We need to reflect on how Clarivate, as a for-profit entity, manages to introduce this database with a relative assurance that it will be commercially fruitful. As mentioned in this post, this is possible because Clarivate sells data as a package: it incorporates a free primary resource and connects it to its main product. This is not a new approach: commercial publishers have done this increasingly in recent years, often based on AI applied to their own publications, and perhaps even on contents copyrighted elsewhere. These companies extract trends from research output and create “prediction products” to be sold (Pooley, 2024). In this case, however, we assist at the flagrant integration of a public resource, built by public funds.
We can hardly blame a commercial entity taking such an approach, but perhaps we should, again, speak out against the commercialization of what was built by the community, with public funds. Open initiatives, which are becoming increasingly abundant, should be prioritized. These resources are already accessible at no cost and therefore, the only possible remaining barrier is a technical one. This is why we need to strengthen the communities’ technical ability to use these resources.
Another issue is that the metadata in OpenAlex (one of the databases from which Research Commons draws from) can be inaccurate in fields such as language, affiliation, and reference information (Alperin et al., 2024; Céspedes et al., 2025). Clarivate announced plans to enrich the metadata available on Research Commons. However, according to a Research Commons webinar from October 2025, they will do so by unifying organization names using Web of Science preferred names, and classifying the records into Web of Science subject categories or research areas. These are not widely adopted persistent identifiers nor classification systems.
OpenAlex also has plans to improve its own data and metadata by adding new works, and improving its topics and keywords algorithms. Similarly, Crossref renewed their partnerships with the Public Knowledge Project (PKP) and the Directory of Open Access Journals (DOAJ) to support more inclusive and equitable metadata in Crossref records.
Responses from the INKE Partnership
Jack Young (Research Impact & Bibliometrics Librarian, McMaster University):
For-profit publishers adopting the language of Open Science as a marketing strategy is nothing new, but the most recent example of this practice in Clarivate’s new “Research Commons” product is particularly striking.
Leveraging both the data and the terminology of the truly open data sources on which its entire existence depends (namely, OpenAlex and CrossRef), the “Research Commons” product proceeds to undercut the core ideals of the Commons by placing this data behind a costly subscription and limiting access instead of opening it up. While entirely permissible from a legal perspective, this naming practice is unquestionably misleading and requires Libraries to draw meaningful distinctions between tools that truly support the ideals of Open Science and those that are “open” in name only.
One area in which this work could have significant impact is in the push for more responsible research assessment practices, spurred by movements like the San Francisco Declaration on Research Assessment (DORA). DORA’s focus on transparency and openness appear in opposition to products like Clarivate’s “Research Commons”. One of DORA’s core recommendations for responsible research assessment is that metrics-supplying organizations, like Clarivate “[p]rovide the data under a licence that allows unrestricted reuse, and provide computational access to data, where possible” (Cagan, 2013). While it’s true that many popular research assessment products don’t currently meet this criteria, Clarivate’s “Research Commons”, in its unfortunate decision to adopt the language of the Commons without embodying its ideals, stands out as being particularly misaligned with DORA’s core principles.
I recognize that research assessment is only one of many potential use cases for this tool, and that meaningful change toward more transparent processes will need to come from multiple directions. Still, as more organizations sign on to DORA and conversations shift to the practical infrastructure required for truly responsible research assessment, I hope the industry will begin to exert greater pressure on database publishers and tool producers to open up their data and share-alike, or risk exclusion from research assessment workflows.
In the meantime, it feels vital for Libraries to be increasingly diligent in how such tools are evaluated, promoted, and used within our own institutions, ensuring that claims of openness are accurate and that practices align with the values we seek to uphold.
References
Alperin, J. P., Portenoy, J., Demes, K., Larivière, V., & Haustein, S. (2024). An analysis of the suitability of OpenAlex for bibliometric analyses (No. arXiv:2404.17663). arXiv. https://doi.org/10.48550/arXiv.2404.17663
Cagan R. (2013). The San Francisco Declaration on Research Assessment. Disease models & mechanisms, 6(4), 869–870. https://doi.org/10.1242/dmm.012955
Céspedes, L., Kozlowski, D., Pradier, C., Sainte-Marie, M. H., Shokida, N. S., Benz, P., Poitras, C., Ninkov, A. B., Ebrahimy, S., Ayeni, P., Filali, S., Li, B., & Larivière, V. (2025). Evaluating the linguistic coverage of OpenAlex: An assessment of metadata accuracy and completeness. Journal of the Association for Information Science and Technology. https://doi.org/10.1002/asi.24979
Clavey, M. (2024, March 4). La recherche française parie sur OpenAlex pour briser l’emprise d’Elsevier et Clarivate. Next. https://next.ink/129485/la-recherche-francaise-parie-sur-openalex-pour-briser-lemprise-delsevier-et-clarivate/
Dulong de Rosnay, M. (2020). Commonswashing – A Political Communication Struggle. Global Cooperation Research – A Quarterly Magazine, Vol. 2, No. 3, pp. 11-13. ISSN 2628-5142 (print). ISSN 2629-3080 (online). https://hal.science/hal-02986722
Edwards, B. (2025, April 2). AI bots strain Wikimedia as bandwidth surges 50%. Ars Technica. https://arstechnica.com/information-technology/2025/04/ai-bots-strain-wikimedia-as-bandwidth-surges-50/
Goldfinger, E. (2025, September 4). Expanding the Web of Science platform with Research Commons. https://clarivate.com/academia-government/blog/expanding-the-web-of-science-platform-with-research-commons/
Hardinges, J., Pearson, S., & Ross, R. (2025). From Human Content to Machine Data. Introducing CC Signals. Creative Commons. https://creativecommons.org/wp-content/uploads/2025/06/Human-Content-to-Machine-Data_Final.pdf
Lamdan, S. (2022). Data Cartels: The Companies That Control and Monopolize Our Information. Stanford University Press. https://doi.org/10.1515/9781503633728
Maddi, A., Maisonobe, M., & Boukacem-Zeghmouri, C. (2025). Geographical and disciplinary coverage of open access journals: OpenAlex, Scopus, and WoS. PLOS ONE, 20(4), e0320347. https://doi.org/10.1371/journal.pone.0320347
Maxwell, T. (2025, April 17). Wikipedia Is Making a Dataset for Training AI Because It’s Overwhelmed by Bots. Gizmodo. https://gizmodo.com/wikipedia-is-making-a-dataset-for-training-ai-because-its-overwhelmed-by-bots-2000590704
Pinfield, S. (2024). Achieving Global Open Access: The Need for Scientific, Epistemic and Participatory Openness (1st ed.). Routledge. https://doi.org/10.4324/9781032679259
Pooley, J. (2024). Large language publishing: The scholarly publishing oligopoly’s bet on ai. KULA: Knowledge Creation, Dissemination, and Preservation Studies, 7(1), 1–11. https://doi.org/10.18357/kula.291
Rosnay, M. D. de, & Stalder, F. (2020). Digital commons. Internet Policy Review, 9(4). https://doi.org/10.14763/2020.4.1530
Tennant, J. P. (2020). Web of Science and Scopus are not global databases of knowledge. European Science Editing, 46, e51987. https://doi.org/10.3897/ese.2020.e51987
Woodcock, C. (2023, May 2). AI Is Tearing Wikipedia Apart. VICE. https://www.vice.com/en/article/ai-is-tearing-wikipedia-apart/
