https://doi.org/10.25547/ZSTN-3Q98
This insight and signals report was written by Kyle Dase, Alan Colin-Arce, Faraz Forghan-Parast, and Ray Siemens.
| Topic / Titre | AI in Peer Review: A Targeted Research Scan |
| Key Participants / Participants clés | Scholarly Journals |
| Date / Période | 2023 – present |
| Keywords / Mots-clés | peer review / critique des pairs, publishing / édition |
Summary
This research scan is the first part of a series that will examine how artificial intelligence (AI) technologies are impacting scholarly communication. This post surveys the current discourse surrounding GenAI’s place within the peer review process. Primary academic concerns at the advent of GenAI revolved around ideas of authorship, plagiarism, and academic transparency, and these issues still drive a passionate conversation. However, as GenAI’s presence in society has regularized, academics have come to understand how this technology might aid in tasks beyond composition, driving a renewed discussion about GenAI’s role within the peer review workflow as an editorial aid. For instance, with proper oversight, GenAI stands to support under-resourced journals in administrative tasks (e.g., matching articles to reviewers and first-round copyediting), giving editors more time for those tasks that require human judgement and creativity. Of course, the introduction of GenAI into these workflows raises its own issues that warrant their own nuanced discussion. Below, we outline core considerations in this discussion, drawing on the latest research on the subject; we also make recommendations for how journals might approach implementing AI in peer review and provide annotations of key sources for further reading.
Themes and Considerations
Many readers will be familiar with how the use of GenAI in writing an article “raises ethics issues on authorship, authenticity, credibility and accountability of academic work, with subsequent legal implications on copyright,” but there are ethical questions of equal import on the editorial side of this issue. For instance, “automated screening tools” can evaluate the appropriateness of an article for a journal, judge whether it conforms to a journal’s editorial standards, and even suggest potential reviewers (Mollaki 2024), but this raises ethical questions about just how much agency editors ought to give these AI systems, how to best incorporate them into the editorial workflow, and how to support the process of Peer Review without compromising the human judgement that makes that process valuable. The following subsections identify key tensions and potential for AI in Peer Review.
AI and Intellectual Ownership
Intellectual property rights are integral to any discussion of AI and Peer Review: just as authors using GenAI risk submitting work that may not be wholly theirs, reviewers’ attempts to detect the use of AI may result in the use of unsanctioned AI-detection software that harvests the data of essays they upload (violating IP rights in the process). In addition, most publishers prohibit reviewers from submitting manuscripts under review to AI technologies to preserve the confidentiality and rights of the original submission. In order for journals to work effectively in the current climate, publishers need to revisit submission guidelines, considering how GenAI’s use affects authors, reviewers, and readers; editors must also re-clarify what is expected of reviewers (including what constitutes acceptable use of AI and AI detection software) and develop clear procedures for when a reviewer believes submission guidelines have not been followed.
AI in the Editorial Workflow
Journals are receiving submissions at an accelerated rate, placing greater stress on the community of volunteer reviewers (Chen et al. 2026; Gartenberg et al. 2026), and editors and reviewers are being asked to do more than ever. GenAI is uniquely positioned to lighten that workload if journals can implement the technology appropriately. GenAI has shown aptitude at appropriately pairing peer reviewers to submissions and performing base-level copyediting and formatting to journal guidelines. Even in terms of a first or last round of revisions, AI review has seen some success in increasing clarity and readability, if not qualitative research; for instance, ICLR, a machine-learning conference, used AI to suggest revisions to peer reviewers (Chen et al. 2026). Hosseini and Horbach have also posited that reviewers might use GenAI to summarize notes into proper reviews, thereby reducing reviewer workload, but advocate for “mandatory disclosure of LLM use in reviews, human accountability for all AI-generated content, training in bias recognition, and institutional policies prohibiting upload of confidential materials,” to ensure that reviewers’ use of AI does not compromise epistemic judgment (Hosseini and Horbach 2023). When partnered with human oversight, each of these roles GenAI performs stands to reduce the current workload of editors and reviewers tremendously through LLM support. However, in order for these tools to be effective, they must be accompanied by clear, discipline-specific guidelines that detail permissible use for reviewers especially; such guidelines are more likely to be absent from humanities and social sciences journals than their STEM counterparts (Wang and Gong 2026).
AI and Readership and Citability
In addition to aiding editors in placing submissions into the right reviewer’s hands and providing authors with an extra round of revisions to improve clarity and readability, GenAI also stands to help scholars anticipate the impact of their research. Recent studies suggest that AI tools like ChatGPT-4 can predict readership, altmetric scores, and even gauge potential citability with greater accuracy than traditional textual and stylometric analysis by evaluating certain criteria (e.g., Accessibility and Understandability, Novelty and Engagement, etc.), essentially giving editors insight into how submissions might perform (de Winter 2024). While there are risks in allowing such predictions to drive what research gets accepted, identifying features of scholarship that will resonate with audiences gives editors a better understanding of the landscape of scholarly communication and the opportunity to suggest improvements/revisions that will benefit their readers.
AI and the Integrity of Peer Review
It may be tempting to address the recent increase in submissions by incorporating AI into peer review wholesale. However, recent studies have shown that the implementation of GenAI into review could render it vulnerable to manipulation and abuse at every stage of the process (Wang et al. 2026). Alternatively, while many editors are concerned over this influx of AI-assisted submissions, the subsequent stress on the community of reviewers, and the potential abuse of AI tools within the review process, scholars have found another, positive use for AI technology that actually strengthens the integrity of peer review. The Institute of Physics Publishing has developed an AI tool that compares peer reviews: in a pilot study of about 500,000 peer reviews between 2020 and 2025, the tool found 89 exact duplicates, over 750 with 80% or more overlap with another report, and just under 2,500 reports with 60% or more overlap (Naddaf 2026). Although this is a limited case study, these tools may help expose vulnerabilities in peer review in the future.
The Need for Specificity
A final consideration: the majority of sources that discuss AI in Peer Review treat scholarship as a monolithic entity, and they make little distinction between the kinds of scholarship that researchers produce. However, it could be of benefit for smaller communities to have their own conversations about the use of AI in knowledge production. How does, for instance, the composition of a quantitative survey in linguistics differ from an archaeological field report or an essay discussing the form of a poem? Do some forms and fields warrant the integration of AI more than others? These discussions are particularly pressing in the humanities and social sciences, as these disciplines are less likely to have AI policies for peer review compared to STEM journals (Wang and Gong, 2026). Even within disciplines, different mindsets and methodologies can take distinct and even conflicting approaches towards AI while each serves that discipline; there is room for many approaches.
Recommendations
Several core recommendations for AI and Peer Review, “Establish collaborative and iterative AI Governance”, “Preserve Human Responsibility in Scholarly Judgment”, “Build Inclusive Processes for Ongoing Evaluation and Response”, and “Require Transparency in AI-Mediated Workflows”, echo our suggestions from the AI Scan developed by the INKE Partnership (available here). In addition, there are a few specific recommendations for the issues on AI use and disclosure in peer review:
Increase transparency and accountability in the peer review process: Concerns about AI use in peer review are based in part on a lack of transparency in the process. Exploring different models that hold reviewers accountable and increase the transparency in peer review might improve review quality and prevent irresponsible AI use. For example, open peer review, reviewer badges, or awards for best reviews might encourage responsible and transparent peer review, even if AI tools were used in the process.
Employ AI in Controlled and Locally-Run Spaces: In the long term, publishers and journals can develop tools to run AI models locally to avoid uploading original manuscripts to external tools not controlled by the publisher (Pilatti et al., 2026). There are also some emerging tools that use locally-run LLMs to check if the references in a paper exist and whether their claims are represented properly in a manuscript (Samsonau, 2026). These spaces preserve the privacy of the review process and give the academic community more control in how AI tools can complement peer review. Some of these tools can also automate the initial editorial labour of formatting and reference checking, which may accelerate editorial decisions.
Establish procedures to investigate misuse of AI tools in peer review: Policies regarding AI use in peer review should consider how to address cases of AI misuse and violations of the policies. One suggestion is to exclude reviewers from conducting future reviews for all journals by the same publisher. However, more research and community consultation is needed regarding the establishment of these procedures.
Recommended Sources
Gartenberg, Claudine, Sharique Hassan, Alex Murray, and Lamar Pierce. (2026). “More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review.” Organization Science 0(0). https://doi.org/10.1287/orsc.2026.ed.v37.n3
Gartenberg et al., as the AI Task Force for Organization Science (a social and behavioural sciences journal), document the development of a growing issue in editing: the advent of GenAI has led to an increase in submission volume alongside a decrease in submission quality, which correlates with an increase in AI-assisted writing. They note that Organization Science has seen a 42% increase in submissions, more than twice the increase experienced during the Covid-19 Pandemic. By using several measures as part of their quantitative analysis, the authors also identify a downward trend in writing quality across multiple scores (e.g., Flesch Reading Ease, FOG Index, Passive Voice, etc.) and attribute this to AI-assisted authorship. For instance, submissions that use AI in 30% or more of their paper are also 30% more likely to receive desk rejections, and those submissions that have “high AI-use” (more than 70%) have a revise & resubmit rate of just 3.2%. Acceptance and AI detection were conducted separately in Organization Science, so “editors, likely without recognizing the writing as AI-generated, consistently judged these manuscripts as lower quality and unworthy of reviewers’ time.” The study also found that a higher use of AI correlated with issues in quality of content, not just writing quality. In other words, articles that employ AI-assisted writing were also more likely to contain “underlying quality issues” related to research. Gartenberg et al. conclude by asserting that AI has the potential to change the field for the better, but individual authors need to consider ways to employ AI that helps them perform higher-quality research rather than simply writing more articles.
Hosseini, Mohammad, and Serge P. J. M. Horbach. 2023. “Fighting Reviewer Fatigue or Amplifying Bias? Considerations and Recommendations for Use of ChatGPT and Other Large Language Models in Scholarly Peer Review.” Research Integrity and Peer Review 8 (1): 15. DOI: https://doi.org/10.21203/rs.3.rs-2587766/v1
Published early in the ChatGPT era, Hosseini and Horbach’s analysis provides a prescient examination of how LLM integration transforms both the pragmatics and ethics of peer review labor. The authors document LLMs’ potential to combat reviewer fatigue by automating time-consuming tasks—transforming informal reviewer notes into polished reports, generating structured feedback on manuscript sections, identifying linguistic or formatting issues—thereby potentially expanding the pool of contributors who can participate effectively despite language barriers or time constraints. However, the article’s core contribution lies in its still-useful systematic identification of risks that emerge when operational assistance crosses into epistemic territory: LLMs trained on existing literature may amplify disciplinary biases (favoring established paradigms over novel approaches), geographic biases (privileging research contexts well-represented in training data), or methodological orthodoxies (flagging unconventional designs as errors rather than innovations). The authors demonstrate through examples how ChatGPT can generate cynical or biased reviews that violate Mertonian norms of universalism, and how confidentiality breaches can occur when reviewers input manuscript excerpts into external AI platforms without institutional safeguards. Hosseini and Horbach’s recommendations—mandatory disclosure of LLM use in reviews, human accountability for all AI-generated content, training in bias recognition, and institutional policies prohibiting upload of confidential materials—have influenced subsequent journal guidelines and provide foundational ethical framework for understanding why guardrails against leakage and disclosure norms become operational necessities rather than optional best practices in AI-entangled peer review systems.
Kousha, Kayvan, and Mike Thelwall. 2024. “Artificial Intelligence to Support Publishing and Peer Review: A Summary and Review.” Learned Publishing 37 (1): 4–12. DOI: https://doi.org/10.1002/leap.1570
Kousha and Thelwall provide a systematic mapping of AI tool deployment across the scholarly publishing pipeline, distinguishing demonstrated capabilities from promotional claims. For journal recommendation, the review documents AI-powered systems—including Springer Nature Journal Suggester, Wiley Journal Finder, IEEE Publication Recommender, and JANE (Journal/Author Name Estimator)—that analyze text similarity with previously published articles and reports high accuracy rates for appropriate journal matching. The analysis of initial quality control covers a diverse toolkit for plagiarism detection, robot author detection, methods checking, automated statistical verification, transparency and reproducibility checking and manuscript structure validation. Commercial systems draw on databases to suggest appropriate reviewers; the Natural Science Foundation of China’s AI-assisted reviewer recommender for grant applications reports approximately 80% accuracy. However, the review identifies a critical boundary: while AI proves effective for finding reviewers and conducting initial quality checks, its value in performing the actual substantive review process “has not been clearly demonstrated.” The synthesis reveals that substantial efficiency improvements are achievable in labor-intensive administrative tasks, while core intellectual evaluation functions remain resistant to automation—a distinction essential for understanding where AI integration in publishing will proceed incrementally versus face fundamental obstacles.
Liang, Weixin, Tara Iyer, Marianna Zhang, Zachary Lipton, and James Zou. 2024. “Evaluating Science: A Comparison of Human and AI Reviewers.” Judgment and Decision Making 19: e24. DOI: https://doi.org/10.1017/jdm.2024.24
Liang and colleagues’ large-scale field experiment comparing GPT-4 with human reviewers on conference abstracts provides empirical evidence about AI’s capabilities and limitations in evaluative judgment. The study demonstrates that while AI can approximate human performance on certain classification tasks—identifying “very best” abstracts shows moderate alignment—detailed evaluative assessments reveal persistent gaps, with human-AI agreement comparable to human-human variability, suggesting AI does not systematically outperform baseline reviewer disagreement. Critically, the research shows humans substantially outperform AI at detecting AI-generated versus human-written content, with detection tools like GPTZero exhibiting higher accuracy than GPT-4 itself when evaluating authorship. This finding has direct implications for peer review integrity when AI-generated manuscripts enter the submission pipeline. The authors position AI as effective for prescreening—rapidly filtering submissions for basic quality thresholds, identifying obvious errors, flagging compliance issues—while demonstrating it lacks the contextual understanding necessary for nuanced scientific judgment about significance, impact, or methodological soundness. The paper’s methodological rigor in isolating human-versus-AI performance dimensions provides context for understanding where operational assistance (screening) legitimately ends and where human evaluative authority (substantive assessment) must begin, directly addressing the guiding principle that human judgment remains central even in AI-augmented workflows.
Moffatt, B., & Hall, A. (2025). “Is AI my co-author? The ethics of using artificial intelligence in scientific publishing .” Accountability in Research, 32(8), 1313–1329. https://doi.org/10.1080/08989621.2024.2386285
Moffatt and Hall (2024) argue that AI systems, particularly Large Language Models (LLMs), need to be excluded from scientific authorship for multiple reasons: inability to take responsibility, lack of intentionality and persistent identity, failure to credit properly, and detrimental effects on the publication ecosystem. The authors draw a distinction between generating text and writing with the intention of publishing and share ethical concerns on author’s intentionality and why AI generated text should be prohibited in publishing. They expand this argument through “The fair credit” and state that opaque training data causes AI to properly attribute the sources of its generated text. The authors convincingly demonstrate that widespread AI-assisted publication would strain an already overburdened peer review system, create competitive disadvantages for researchers who choose not to use AI, and incentivize quantity over quality. The observation that AI adoption could disadvantage those unwilling or unable to use it raises important equity concerns as well. The authors elaborately give reasons why widespread AI-assisted publication would add pressure to already struggling peer review systems, create competitive disadvantages for researchers who do not use AI, and incentivize quantity over quality. They also share their equity concerns on how AI adoption could disadvantage those unable to use it.
Mollaki, V. (2024). “Death of a reviewer or death of peer review integrity? the challenges of using AI tools in peer reviewing and the need to go beyond publishing policies.” Research Ethics, 20(2), 239-250. https://doi.org/10.1177/17470161231224552.
Mollaki discusses the implications of the lack of policies regarding the use of AI tools in the peer review process among the 10 largest commercial publishers in the world (as of early 2024). Analyzing the publisher websites, the author found that only Elsevier and Taylor & Francis had a specific policy for reviewers regarding the use of AI tools. Both publishers banned it due to confidentiality and copyright concerns. However, broader risks to this lack of policies around AI use in peer review include a reduction of trust in the peer review process and decision-making in published manuscripts, as well as an increase in disputes due to a lack of transparency in the review process. Mollaki suggests implementing clear policies regarding the use of AI tools in peer review, including procedures to investigate their misuse and potentially excluding reviewers from the publishing process if they violate the policies.
Pilatti, Luiz Alberto, José Roberto Herrera Cantorani, and Fabiana Fátima Do Prado Sedelak Pinheiro. 2026. “AI-Assisted Peer Review: A Scoping Review of Governance, Ethical-Behavioral Risks, and Integrity.” Ethics & Behavior 0 (0): 1–17. https://doi.org/10.1080/10508422.2026.2660125.
Drawing from a literature review, Pilatti et al. argue that AI-assisted peer review poses three risks to the process: effort outsourcing and accountability laundering, limited detectability of AI-generated reviews, and exploitation of AI tools’ vulnerabilities. In response, the authors propose an operational model called Auditable Hybrid Review (AHR) that ensures the confidentiality, accountability, and verifiability of peer review supported by AI technologies.
The AHR model suggests that uses of generative AI should 1) happen in environments controlled by the journal or publisher; 2) require specific disclosure of the tasks and content processed by AI tools; 3) require substantive critiques anchored in specific sections of the manuscript; and 4) process the original manuscript to prevent exploitation of the AI tools’ configurations and biases (also see Wang et al., 2026). This model is based on the assumption that AI-assisted peer review is not an individual issue, but a problem of institutional governance, so addressing it should go beyond AI bans or textual policing. The model can also be adapted depending on the publishers’ size, capacity, and risk of AI use by reviewers.
Samsonau, Sergey V. 2026. “Sciwrite-Lint: Verification Infrastructure for the Age of Science Vibe-Writing.” arXiv:2604.08501. Preprint, arXiv, April 9. https://doi.org/10.48550/arXiv.2604.08501.
Samsonau introduces Sciwrite-Lint, a locally run LLM tool that assesses the accuracy of the claims made in a manuscript and verifies whether the cited references are not hallucinated and support these claims. It measures the proportion of accurate and unretracted references, the consistency of the numbers and statistics mentioned, the context in which citations appear, the accuracy of the reference lists of the cited papers, and the contribution of the paper. After conducting these assessments, the tool assigns a SciLint Score, which aggregates the structural quality of the evidence and arguments in a paper, as well as its contribution (calculated based on frameworks for assessing arguments in the philosophy of science). Samsonau suggests that this tool could help journals to assess papers before sending them for review and reduce reviewers’ time spent verifying reference lists. It can also assist preprint servers to provide a quality indicator for their papers and give authors concrete feedback on how to strengthen the evidence supporting their claims.
Shiping, Chen, Shu Zhong, Duncan P. Brumby, and Anna L. Cox. (2026). “What Happens When Reviewers Receive AI Feedback in Their Reviews?” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. 1-19. https://doi.org/10.48550/arXiv.2602.13817.
Chen et al. conduct a survey of respondents from presenters who submitted to the International Conference on Learning Representations (ICLR), “one of the top three Machine Learning conferences.” ICLR’s submission rate has grown exponentially from around 3500 to over 11000 submissions in just three years. To accommodate this growth, ICLR implemented an “AI feedback tool” for the 2025 review cycle, providing feedback to reviewers to “improve clarity, tone and specificity”; reviewers were then given the option to revise their reviews according to this feedback if they wished. Chen et al. survey 51 reviewers who participated in peer review for ICLR’s 2025 conference, asking three key research questions: how do reviewers perceive AI feedback? How do they respond to that feedback and what actions do they take? And “what benefits and drawbacks do reviewers identify” in this “AI-assisted peer review”?
The authors conclude by identifying a paradox in reviewer perceptions of AI-generated feedback: reviewers consistently found that “the tool improved their reviews” and yet at the same time did not find it particularly helpful. At the root of this judgment is the notion that the AI was merely “polishing prose” rather than providing the judgment and intellectual contributions that make a review valuable. Chen et al. argue that the inclusion of AI in this process may help shift perceptions of reviewer roles for the better, encouraging them towards “greater elaboration and clarity.”
Sun, Zhuanlan. 2025. “Large Language Models in Peer Review: Challenges and Opportunities.” Scientometrics 60 (3): 1683-1706. DOI: https://doi.org/10.1007/s11192-025-05440-w
Sun provides essential mapping of LLM applications across the peer review lifecycle, categorizing five distinct operational roles: checklist assistants for formatting and protocol adherence, reviewer selection aids that match manuscripts with appropriate expertise, feedback generators for preliminary assessments, bias detectors that flag methodological or statistical irregularities, and agents that coordinate multi-stage review workflows. The article systematically examines technical approaches including prompt engineering strategies, model evaluation frameworks, and architectural designs for integrating LLMs into editorial management systems. The author argues that while LLMs excel at standardized operational tasks—checking reference integrity, identifying duplicated content across submission databases, verifying compliance with reporting guidelines—they remain fundamentally limited in assessing research novelty, theoretical contributions, and domain-specific methodological rigor. The analysis emphasizes that current LLM limitations (inadequate scientific validation, domain knowledge gaps, inability to analyze complex datasets, ethical concerns around bias perpetuation) position them as supportive tools within human-led processes rather than autonomous decision-makers. This work is useful in understanding operational versus epistemic divisions of labor in AI-assisted scholarship service and documenting the technical infrastructure through which AI becomes entangled with peer review operations while maintaining clear boundaries around human evaluative judgment.
Tang, Y., Kang, Y., Wu, S., Zhang, R., & Sun, Z. (2026). “Can large language models assess the quality of peer review? An empirical study.” Scientometrics 131(4), 2237–2259. https://doi.org/10.1007/s11192-026-05622-0
This study examines whether large language models (LLMs) can reliably evaluate the quality of scholarly peer reviews. The authors compare the performance of three leading LLMs (ChatGPT-4o, Claude 3.5 Sonnet, and Gemini-EXP-120) using six established review quality instruments (RQIs) applied to peer review reports from 180 biomedical papers published in eLife. Two prompting strategies are used, Zero-Shot and Few-Shot prompting, across three evaluation scenarios: direct scoring, simulated double-blind review, and acting as domain experts. The authors situate their research within the growing use of artificial intelligence in scholarly publishing, emphasizing that traditional assessment of review quality is subjective, labor-intensive, and dependent on human judgment. They argue that LLMs, trained on extensive textual datasets and capable of producing human-like responses, may improve review of evaluation processes. While researchers have examined LLMs in manuscript evaluation and review generation, the authors note that their use for assessing review quality remains largely unexplored. Methodologically, the study uses a representative sample of biomedical peer review reports selected according to citation impact and reviewer h-index to ensure diversity and relevance. Among the tested models, Claude 3.5 Sonnet showed the strongest overall performance in terms of repetition consistency and similarity to human scoring. Overall, the article contributes to emerging research on AI-assisted scholarly publishing by demonstrating both the promise and limitations of LLMs in peer review quality assessment. The study concludes that although LLMs may assist in evaluating review quality, they are not yet reliable enough to replace human evaluators and should be used cautiously alongside human oversight.
Wang, Jialiang, Yuchen Liu, Hang Xu, et al. 2026. “When AI Reviews Science: Can We Trust the Referee?” The Innovation Informatics 2 (1): 100030–106. https://doi.org/10.59717/j.xinn-inform.2026.100030.
Wang et al. classify the security vulnerabilities at all stages of AI peer review: training and data retrieval of AI models, desk review, deep review, rebuttal, and broader system-level unreliability. Through a review of the literature and an AI peer review experiment using GPT 5.1 and Gemini 2.5, the authors found that AI peer review can be compromised at every stage of the review process by several attack methods that exploit biases in the training data and the models’ configuration. These methods include contaminating a model’s training data, drafting papers with hidden prompts or exaggerated claims to mislead the AI into providing a positive review, challenging the AI reviewer’s initial assessment without evidence, and exploiting the models’ bias towards prestigious institutions and references.
Therefore, AI peer review systems are susceptible to penalizing papers using cautious language, accepting authoritative but unsubstantiated rebuttals, and favoring prestigious institutions over lesser-known ones. These vulnerabilities highlight the need for further research into preventing such attacks, as well as mechanisms for addressing them, such as human-in-the-loop processes, audit trails, and bias reviews.
Wang, Zhongshi, and Mengyue Gong. 2026. “A Cross-Disciplinary Analysis of AI Policies in Academic Peer Review.” Learned Publishing 39 (1): e2035. https://doi.org/10.1002/leap.2035.
Wang and Gong analyzed the editorial policies of 802 journals regarding the use of AI in peer review. Their results found that 83% of journals with high impact factors and 75% of journals with middle impact factors had established AI policies, although the presence and content of the policies varied by discipline. Fewer journals in the humanities and social sciences had AI review policies in place and they were less prohibitive regarding AI use, while STEM journals had established more policies and were more likely to completely restrict AI use in peer review. To conclude, the authors suggest establishing AI policies for specific disciplines, increasing the transparency of peer review processes to prevent misconduct, and developing measures for addressing AI misuse.
