Skip to main navigation Skip to search Skip to main content

Reviewer Experience Detecting and Judging Human Versus Artificial Intelligence Content: The Stroke Journal Essay Contest

  • Gisele S. Silva
  • , Rohan Khera
  • , Lee H. Schwamm
  • , Maurizio Acampa
  • , Eric E. Adelman
  • , Johannes Boltze
  • , Joseph P. Broderick
  • , Amy Brodtmann
  • , Hanne Christensen
  • , Lachlan Dalli
  • , Kelsey Rose Duncan
  • , Islam Y. Elgendy
  • , Adviye Ergul
  • , Larry B. Goldstein
  • , Janice L. Hinkle
  • , Michelle C. Johansen
  • , Katarina Jood
  • , Scott E. Kasner
  • , Steven R. Levine
  • , Zixiao Li
  • Gregory Lip, Elisabeth B. Marsh, Keith W. Muir, Johanna Maria Ospel, Joanna Pera, Terence J. Quinn, Silja Räty, Anna Ranta, Lorie Gage Richards, Jose Rafael Romero, Joshua Z. Willey, Argye E. Hillis, Janne M. Veerbeek

Research output: Contribution to journalReview articlepeer-review

9 Scopus citations

Abstract

Artificial intelligence (AI) large language models (LLMs) now produce human-like general text and images. LLMs' ability to generate persuasive scientific essays that undergo evaluation under traditional peer review has not been systematically studied. To measure perceptions of quality and the nature of authorship, we conducted a competitive essay contest in 2024 with both human and AI participants. Human authors and 4 distinct LLMs generated essays on controversial topics in stroke care and outcomes research. A panel of Stroke Editorial Board members (mostly vascular neurologists), blinded to author identity and with varying levels of AI expertise, rated the essays for quality, persuasiveness, best in topic, and author type. Among 34 submissions (22 human and 12 LLM) scored by 38 reviewers, human and AI essays received mostly similar ratings, though AI essays were rated higher for composition quality. Author type was accurately identified only 50% of the time, with prior LLM experience associated with improved accuracy. In multivariable analyses adjusted for author attributes and essay quality, only persuasiveness was independently associated with odds of a reviewer assigning AI as author type (adjusted odds ratio, 1.53 [95% CI, 1.09-2.16]; P=0.01). In conclusion, a group of experienced editorial board members struggled to distinguish human versus AI authorship, with a bias against best in topic for essays judged to be AI generated. Scientific journals may benefit from educating reviewers on the types and uses of AI in scientific writing and developing thoughtful policies on the appropriate use of AI in authoring manuscripts.

Original languageEnglish
Pages (from-to)2573-2578
Number of pages6
JournalStroke
Volume55
Issue number10
DOIs
StatePublished - Oct 1 2024

Bibliographical note

Publisher Copyright:
© 2024 American Heart Association, Inc.

Funding

Dr Silva reports employment by Universidade Federal de São Paulo, compensation from ISchemaView for consultant services, compensation from Bayer for consultant services, compensation from Boehringer Ingelheim for consultant services, employment by Sociedade Beneficente Israelita Brasileira Albert Einstein, and compensation from Pfizer for other services. Dr Khera reports grants from BridgeBio; a provisional patent for methods of generating digital twin-based data sets; an ownership stake in Ensight-AI, Inc; employment by Yale School of Medicine; grants from the National Institutes of Health; a patent pending for Methods for Neighborhood Phenomapping for Clinical Trials, No. 63/177117; a provisional patent for format independent detection of cardiovascular disease from printed ECG images with deep learning licensed to Ensight-AI, Inc; grants from Novo Nordisk; grants from Blavatnik Family Foundation; a provisional patent for machine learning method for adaptive trial enrichment; an ownership stake in Evidence2Health; grants from Doris Duke Charitable Foundation; a provisional patent for a multimodal video-based progression score for aortic stenosis using artificial intelligence; a provisional patent for artificial intelligence–guided screening of underrecognized cardiomyopathies adapted for POCUS; grants from Bristol Myers Squibb; stock holdings in Evidence2Health; a provisional patent for biometric contrastive learning for data-efficient deep learning from electrocardiographic images licensed to Ensight-AI, Inc; and a provisional patent for articles and methods for detecting hidden cardiovascular disease from portable electrocardiography. Dr Schwamm reports compensation from Medtronic for consultant services.

FundersFunder number
Doris Duke Charitable Foundation
Blavatnik Family Foundation
Novo Nordisk Data Science
Yale University, School of Medicine
Bristol-Myers Squibb
National Institutes of Health (NIH)63/177117

    Keywords

    • artificial intelligence
    • neurologists
    • peer review
    • stroke
    • writing

    ASJC Scopus subject areas

    • Clinical Neurology
    • Cardiology and Cardiovascular Medicine
    • Advanced and Specialized Nursing

    Fingerprint

    Dive into the research topics of 'Reviewer Experience Detecting and Judging Human Versus Artificial Intelligence Content: The Stroke Journal Essay Contest'. Together they form a unique fingerprint.

    Cite this