Download PDF
Artificial intelligence is increasingly acting as a first interpreter of biomedical research, shaping how evidence is applied to patient care. Simultaneously, science is reaching broader, non-specialist human audiences. In both cases, interpretive errors and generalization bias can skew clinical decision-making and endanger public health. Rather than waiting passively for more advanced AI to solve these problems, we contend that science itself can adapt by modernizing its reporting standards
Subjects
Introduction
Scientific information is increasingly being consumed in AI-summarized form, and the first “reader” and interpreter of a study is commonly becoming a machine (e.g., ChatGPT, OpenEvidence)1,2,3,4. Though these technologies offer promising new opportunities, they are also prone to interpretive inaccuracies which can color public perceptions and undermine sound research, treatment decisions, clinical guideline development, public health decision-making, and global health equity5,6,7. A prevailing assumption is that these challenges will be solved primarily by more advanced AI; we contend that science itself can also adapt by developing clearer, more intentional ways of communicating results to the AI systems that increasingly interpret it.
Similarly, scientific findings are also reaching broader and more heterogeneous audiences, including journalists, policymakers, and social media. While this increases visibility and potential impact, it also introduces risk. Nonspecialist readers may be less equipped to interpret methodological nuance and more likely to perceive tenuous associations as authoritative. These challenges are often treated as unavoidable, but we posit that science should proactively adapt how it communicates results to a more diverse human readership.
We propose that, just as reporting standards such as CONSORT and STROBE formalize reproducibility in study design, distinct but analogous standards are needed to preserve accuracy and reproducibility in evidence interpretation. To illustrate this point, we outline two mechanisms for interpretive framing—SIMcard and FACTS—that anchor human and machine interpreters to explanations verified by peer-review, thereby reducing the risk that clinicians, patients, and policymakers act on downstream interpretations untethered from what the data show. But the larger takeaway is collective: broad and collaborative efforts are needed to reimagine how future reporting standards can ensure discovery survives translation through rapidly diversifying human and machine interpreters.
Vulnerabilities within scholarship and evidence-based medicine
Every stage of knowledge production is becoming filtered through AI. Researchers are using generative AI models to conduct studies and draft manuscripts1, peer reviewers are using AI to assess quality2, and readers increasingly rely on AI tools to summarize published articles3,4,5. This can result in several layers of opaque filtering that shape how evidence is understood. Put more plainly, we are moving toward a future where machines will increasingly communicate science to and through other machines before information reaches a human user.
On one hand, the use of AI interpreters has the potential to offer important benefits. Patient chart summaries produced by generative AI, for example, have been shown to mirror those authored by physicians in terms of completeness and correctness8. On the other hand, outsourcing interpretive responsibility at any level creates vulnerability to potential disinformation. A 2025 study compared nearly 5000 AI-generated summaries across 10 prominent LLMs (e.g., ChatGPT, Claude, DeepSeek) and found that, even when explicitly prompted for accuracy, LLMs produced more extreme generalizations than justified by the original study5. Even AI designed specifically to help clinicians interpret evidence have shown concerningly poor results in early studies3,4. Similarly, AI used for literature review consistently miss key references and are, at best, only partially accurate1. As a concerning example, more than twenty peer-reviewed publications have uncritically repeated the term “vegetative electron microscopy,” a nonsensical phrase introduced by an erroneous AI summary and propagated through citation. This resulted in at least one contested retraction in a Springer Nature Journal9.
A key consideration is that AI interpretations are susceptible to multiple forms of bias and can be inconsistent between users and AI platforms. The same article may be summarized with different conclusions depending on how a prompt is phrased, language used, AI model, training data, or hallucination5,10,11,12. Several publicly-available AI are trained on internet content that may be susceptible to inaccuracies, and as models become saturated with public data, individual human-AI dialogs are even being repurposed as training data12. This feedback loop risks amplifying user-confirmation bias10. An AI asked, “Does diet soda cause stroke?” may give a substantively different answer than if prompted with, “What associations and potential confounders were reported in the 2017 Pase et al. study?”13,14 Thus, even among experts, people may no longer be “reading” the same paper in the same way. This can confuse the responsible use of evidence at every level, from high-level guidelines to individual clinical decision making.
Vulnerabilities across broader audiences
The same forces reshaping expert discourse are also transforming how the public encounters scientific evidence, especially as the wide accessibility of AI animates public engagement with peer-reviewed research5,6. However, non-specialist readers often lack the training to recognize important methodological limitations when interpreting and debating studies, which can result in mistaking weak or conditional associations for strong evidence. For example, in the United States, observational studies on prenatal acetaminophen exposure and autism have recently been widely misinterpreted as proof of causality, despite authors’ explicit caveats15. This has contributed to public concerns about the safety of acetaminophen during pregnancy that have been difficult to assuage and may compromise clinical care16.
At the broadest level, the interpretive unreliability of AI becomes a matter of global health equity. A disproportionate amount of research is conducted in high-income countries and published in English, creating algorithmic bias in the training data for AI. As AI systems then increasingly mediate how biomedical research is summarized and applied to patient care, readers in non-English speaking and low-resource settings may receive more biased or mistranslated outputs. Furthermore, AI adoption will occur at different rates in different settings based on demand and infrastructure constraints. When misinformation spreads unevenly across regions, it can undermine the consistent implementation of evidence-based guidelines and drive inconsistent adoption of effective therapies, further reinforcing geographic inequities. Therefore, the reporting standards that guide how we communicate scientific findings must evolve with the systems and audiences that now interpret them.
Potential strategies: interpretive framing
Early and accurate interpretive framing can be an effective strategy to mitigate downstream bias and inaccurate or exaggerated narratives. Current efforts to improve foundational AI model performance, for example, increasingly emphasize training data quality and structure as one of the most effective paths to proper alignment and reliability17. Similarly, among human readers, pre-emptive psychological priming against online misinformation has been shown to effectively increase resilience to misinformation across a wide variety of covariates18. As a corollary, inaccurate framing can be a driver of misinformation. A 2016 study found that exaggerations in academic press releases predict exaggerations in news media, whereas press releases with caveats predict corresponding caveats in news media about causal relationships19. These findings highlight the importance of interpretive framing at the point of publication, rather than scrambling to correct or retract errant narratives after they have spread.
As an example of what one such strategy might look like, we propose that dissemination of biomedical research might benefit from a layer of peer-reviewed Structured Interpretive Metadata (a “SIMcard”) for machines (Fig. 1) and a corresponding Framing Appropriate Claims & Their Scope Label (a FACTS label) (Table 1) for humans. These would be authored by the investigators and evaluated by reviewers alongside the main text. They would complement the traditional abstract by ensuring that the essential claims and limitations of the study are clearly stated and pre-emptively anchored to subsequent layers of AI and public interpretation. We describe these approaches to open a wider conversation on how future reporting standards might better promote accurate interpretation across increasingly diverse human and machine audiences.
Figure 1 shows an example template for structured interpretive metadata (a “SIMcard”) for an observational study using XML. For purposes of illustration, this template is populated for the Pase et al. 2017 study on diet soda consumption and stroke/dementia risk13, which was widely misinterpreted at the time of publication as evidence that diet soda consumption caused strokes. A SIMcard will vary by study design and discipline, but in all cases, it is structured and simple. Authors unfamiliar with XML or JSON can still easily populate each field based on their study and reviewers can still easily assess accuracy.
Structured interpretive metadata for machine readers
The SIMcard is structured in that it follows a clear, standardized template (or “card”) that authors fill out based on their study. This includes controlled vocabulary (e.g., observational, randomized, etc.) to ensure easy and consistent categorization. It is interpretive in that it makes the limits of inference explicit, putting constraints on how an AI can and cannot interpret the study. Finally, as metadata, it is encoded in a machine-readable format (e.g., XML or JSON, which are widely used standards for structured data exchange and are compatible with modern AI and web-based systems). Practically, structured XML tags or JSON keys are retrievable by AI systems during document ingestion and could be used to tether a machine’s interpretive process to verified boundaries rather than probabilistic inference.
SIMcard represents a practical extension of existing metadata processes. Journals already generate metadata for PubMed and Crossref indexing; however, this metadata is typically only used to capture bibliographic information. Similarly, the FAIR data principles (Findable, Accessible, Interoperable, and Reusable) and efforts toward machine-actionable metadata have shown that publications can be structured so that computers can extract and reason over study attributes such as design, sample, and variables20. Frameworks such as RO-Crate already use this logic: each research article or dataset is packaged with a small companion file that describes what the object is and how it connects to related data and code. This makes it possible for automated systems to find and reuse research outputs across platforms. Studies further suggest tagging scientific literature with metadata improves algorithmic accuracy in identifying study characteristics21. Taken togther, current metadata frameworks enable machines to locate, organize, and extract information about studies, but they do not yet shape how those findings are interpreted once extracted. SIMcard addresses this gap. In this sense, existing metadata work is enabling for SIM, but non-overlapping (Table 2).
Similarly, SIMcard complements existing frameworks for AI in research, such as CONSORT-AI/SPIRIT-AI22,23, DECIDE-AI24, MI-CLAIM25, TRIPOD + AI26, PROBAST + AI27, and STARD-AI28 (Table 2). These frameworks mainly focus on transparency, evaluation, bias, and reporting quality for studies that develop and validate AI systems. By contrast, SIMcard addresses a different downstream problem: how any biomedical study, including non-AI studies, is interpreted by AI systems after publication. In this sense, existing standards improve the reporting and conduct of research about AI, whereas SIMcard aims to improve research interpreted by AI. The overlap is that all are concerned with transparency and reproducibility; the distinction is that they act at different stages of the research life cycle (e.g., before/while a study is executed vs. after a study is published) and attempt to solve fundamentally different problems.
FACTS label for general readers
For human readers, the FACTS label would be written in plain text and include:
- 1.
A one-sentence primary claim written in neutral, non-speculative language;
- 2.
A statement clarifying causal or associational status;
- 3.
A description of the study context (e.g., population, setting, implementation conditions studied) and a statement about degree of generalizability;
- 4.
A concise summary of uncertainties and limitations;
- 5.
A prohibited-inference clause anticipating and flagging potential misreadings; and
- 6.
The traditional abstract summarizes the full text in a way familiar to researchers, but it does not take any further steps to guide accurate interpretation for the nonspecialist reader. Likewise, many journals also publish compressed summaries, such as “Key Points” in JAMA or the “Research in Context” boxes in The Lancet, but these are not a universal academic standard. They often assume shared research literacy and function primarily to convey importance and relevance. They do not have a core mandate to distill findings in plain language for broad audiences in a way that resists over-generalization or naïve misunderstanding.
A FACTS label, by contrast, is defined for this purpose and audience. In particular, the “context where results apply,” “key limitations,” and “prohibited inference” sections clearly delineate the limits of inference, which are usually left implicit under the assumption of a reader with shared research literacy. Similarly, the “recommended interpretation” section orients the reader to a simple and scientifically-sound way to understand the findings. In these ways, the FACTS label complements the traditional abstract by ensuring that the right ideas are communicated to the right audience from the start. It does not replace the traditional abstract, however, as it does not contain the necessary conceptual and methodological rigor to meaningfully communicate the findings to a trained researcher in that field.
Treatment of causal inference
Causal inference is often based on an accumulation of evidence at the field-level rather than a single study. For example, we do not have randomized trials that show smoking causes cancer because it would be unethical to assign people to a smoking group. As a result of the accumulation of observational cohort studies, however, we can confidently say it does. A single study may support different levels of inference depending on its design, methods, consistency with prior evidence, biological plausibility, temporality, dose-response patterns, and risk of bias.
SIMcard and FACTS do not treat causation as binary (i.e., permitted vs. prohibited). Rather, they are designed to describe the level of causal inference supported by the study itself, while allowing that stronger causal confidence may emerge from the accumulated body of evidence. For example, a study may be best characterized as hypothesis-generating, associational (as depicted in Fig. 1), supportive of limited causal inference, or consistent with a causal interpretation when integrated with prior evidence. This distinction is critical because SIMcard and FACTS are not intended to decide whether a causal relationship is true at the field level. More modestly, they are intended to prevent a single paper from being interpreted as stronger evidence than it actually is.
Path to implementation
Implementation will require careful attention to workflow. SIMcard and FACTS should not require authors to hand-code XML or reviewers to adjudicate metadata syntax. Instead, they should be built into journal submission and production systems as structured templates. Authors would draft the scientific content: the primary claim, study population, outcomes, causal or associational status, limitations, uncertainty, and inappropriate inferences. Reviewers and editors would assess whether these interpretive statements fairly reflect the manuscript, much as they already assess abstracts, methods, conclusions, and limitations. Publishers should manage formatting and machine-readability. Coordination through established frameworks such as the International Committee of Medical Journal Editors (ICMJE) or the National Library of Medicine (NLM) could further ensure consistency across journals and indexing services.
Adoption will depend on incentives. Authors and publishers are unlikely to embrace SIMcard or FACTS if they are experienced only as additional administrative work. The value proposition is that these tools may protect authors from having erroneous claims attached to their work and allow journals to signal responsible publication practices in an AI-mediated information environment. Framed this way, SIMcard and FACTS are not “extra paperwork.” They are a form of quality control that directly serves author and editor interests. If further shown to be effective by empirical validation, it would be perplexing for authors and publishers not to want SIMCard and FACTS in their published findings.
Effectiveness will require further study. Adversarial prompting provides a realistic framework to evaluate whether SIMcards improve interpretive fidelity, particularly in scenarios where models are pressed toward causal overinterpretation or hallucination. A pilot study design could compare LLM outputs under different conditions (e.g., with and without SIM metadata, and with retrieval-augmented access to SIM), using both neutral and adversarial prompts. Outcomes could be evaluated based on interpretive accuracy, consistency across prompts, and violations of prohibited inferences. This would be an important direction for future empirical work. Even then, in practice, utility will depend on how it is integrated into AI systems.
AI integration will occur in stages. As currently configured, LLMs do not inherently enforce metadata constraints when operating in a purely generative mode. Without explicit retrieval or instruction, LLM outputs remain driven by probabilistic patterns learned during training. Therefore, long-term success of interpretive guardrails will depend on alignment between journals, indexing services (e.g., PubMed), and AI developers. Journals, as the central arbiter of new scientific knowledge, should act as the initial catalyst for this process by proactively adding the metadata layer at publication, creating an authoritative signal that AI systems can access.
In the early stages, individual users could directly instruct AI systems to prioritize <SIM> tags. Scientists and journalists who use AI to query scholarly articles presumably have an innate interest in contextually accurate summaries. Thus, the barrier to adoption at this level is simply user awareness. However, for these signals to be recognized consistently and automatically across all users, AI developers can design retrieval-augmented generation (RAG) and agentic systems to preferentially query and reason over SIM. They could even incorporate SIM-like metadata into the pretraining and fine-tuning of foundation models.
Rivalry between AI companies and the competitive advantage of producing high-fidelity, contextually accurate summaries would create market incentives for doing so, and emerging decisions already point in this direction. For example, in April 2026, an AI literature review tool called Scite launched a “connector” with Anthropic’s Claude, allowing Claude’s responses to be grounded in citation-backed analysis from more than 250 million peer-reviewed full-text scientific articles rather than relying solely on training data29. So, developer interest in interpretive guardrails and accurately summarizing peer-reviewed science already exists. Over time, AI systems designed from the outset to prioritize interpretive metadata could become an industry standard. This would open an entirely new avenue of research into how interpretive metadata can be optimized to produce the most reliable interpretations, especially as AI retrieval architecture evolves.
Finally, there is the question of governance. SIMcard and FACTS labels could introduce new risks if poorly implemented. Authors may overgeneralize their findings or frame limitations strategically. Reviewers may treat the process as another checkbox. Responsibility should therefore be distributed across the publication workflow: authors are responsible for scientific accuracy; reviewers and editors are responsible for adjudicating whether those statements are justified by the manuscript; and publishers are responsible for formatting and correction mechanisms. As with the other content of a scholarly publication, SIMcard and FACTS labels should be reviewable and amendable when errors are identified.
Conclusion
Most readers of biomedical science will soon encounter evidence only after it has passed through an AI interpreter. The challenge for 21st-century medicine is no longer to merely report results accurately, but to ensure that those results remain accurate as they pass through increasingly opaque systems of interpretation. We have proposed that, rather than waiting passively for more sophisticated AI, science itself can adapt to communicate its results more clearly for its new interpreters. Thus, interpretive framing, optimized separately for humans and machines, offers one potential strategy in this new landscape. But the larger takeaway is collective: we as a scientific community must devote greater attention to the processes by which we communicate complex results to a rapidly evolving audience.
The use of AI holds great promise to positively impact humanity by advancing biomedical inquiry and clinical practice. The future of successful science, therefore, depends not only on what we discover, but on whether our machines—and human readers—can understand it the same way
Data availability
No datasets were generated or analysed during the current study
References
Sollini, M. et al. Human researchers are superior to large language models in writing a medical systematic review in a comparative multitask assessment. Sci. Rep.16, 173 (2025)
Perlis, R. H. et al. Artificial intelligence in peer review. JAMA334, 1520–1522 (2025)
Low, Y. S. et al. Answering real-world clinical questions using large language models, retrieval-augmented generation, and agentic systems. Digit Health11, 20552076251348850 (2025)
Jagarapu, J., Babata, K., Chamarthi, S., Hoyt R. The accuracy and repeatability of OpenEvidence on complex medical subspecialty scenarios: A pilot study. medRxiv. 2025-11 (2025)
Peters, U. & Chin-Yee, B. Generalization bias in large language model summarization of scientific research. R. Soc. Open Sci.12, 241776 (2025)
Handler, R., Sharma, S. & Hernandez-Boussard, T. The fragile intelligence of GPT-5 in medicine. Nat. Med.31, 3968–3970 (2025)
Peoples, N., Østbye, T. & Yan, L. L. Burden of proof: Combating inaccurate citation in biomedical literature. BMJ. 383 (2023)
Van Veen, D. et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med.30, 1134–1142 (2024)
Snowswell, A., Witzenberger, K. & El-Masri, R. A strange phrase keeps turning up in scientific papers, but why? Science Alert. Accessed. Retrievable from: https://www.sciencealert.com/a-strange-phrase-keeps-turning-up-in-scientific-papers-but-why
Lu, J. G., Song, L. L. & Zhang, L. D. Cultural tendencies in generative AI. Nat. Human Behav.9, 2360–2369 (2025)
Xu, Z., Jain, S. & Kankanhalli, M. Hallucination is inevitable: An innate limitation of large language models. arXiv preprint arXiv:2401.11817. (2024)
Xu, R. et al. The earth is flat because…: Investigating llms’ belief towards misinformationeeting of the Association for Computational Linguistics (Volume 1: Long Papers) 2024 (pp. 16259–16303)
Siden, R. et al. A typology of physician input approaches to using AI chatbots for clinical decision-making. npj Digital Med.9, 14 (2025)
Pase, M. P. et al. Sugar-and artificially sweetened beverages and the risks of incident stroke and dementia: A prospective cohort study. Stroke48, 1139–1146 (2017)
Schweitzer K. Acetaminophen use in pregnancy—study author explains the data. JAMA. 1499–1501 (2025)
Faust, J. S. & Barnett, M. L. Changes in paracetamol and leucovorin use after a White House briefing. Lancet407, 1051–1053 (2026)
Liu, Y., Cao, J., Liu, C., Ding, K. & Jin, L. Datasets for large language models: A comprehensive survey. Artif. Intell. Rev.58, 403 (2025)
Roozenbeek, J., Van Der Linden, S., Goldberg, B., Rathje, S. & Lewandowsky, S. Psychological inoculation improves resilience against misinformation on social media. Sci. Adv.8, eabo6254 (2022)
Sumner, P. et al. Exaggerations and caveats in press releases and health-related science news. PloS one11, e0168217 (2016)
Batista, D., Gonzalez-Beltran, A., Sansone, S. A. & Rocca-Serra, P. Machine actionable metadata models. Sci. Data9, 592 (2022)
Zhang, Y., Jin, B., Zhu, Q., Meng, Y., Han, J. The effect of metadata on scientific literature tagging: A cross-field cross-model study. In Proceedings of the ACM Web Conference 2023 1626–1637 (2023)
Liu, X. et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Lancet Digit. Health2, e537–e548 (2020)
Rivera, S. C. et al. Guidelines for clinical trial protocols for interventions involving artificial intelligence: The SPIRIT-AI extension. Lancet Digit. Health2, e549–e560 (2020)
Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ28, 924–933 (2022)
CAS
Google ScholarNorgeot, B. et al. Minimum information about clinical artificial intelligence modeling: the MI-CLAIM checklist. Nat. Med.26, 1320–1324 (2020)
Collins, G. S. et al. TRIPOD + AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ385 (2024)
Moons, K. G. et al.PROBAST+ AI: an updated quality, risk of bias, and applicability assessment tool for prediction modelsusing regression or artifi cial intelligence methods. BMJ388 (2025)
Sounderajah, V. et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat. Med.31, 3283–3289 (2025)
Scite is now a Claude connector. Accessible from: https://scite.ai/blog/scite-claude-connector (Accessed, 2026)
Acknowledgements
The authors would like to acknowledge Dr. Chris Carpenter from Mayo Clinic for helpful feedback on an early draft
Author information
Authors and Affiliations
Department of Emergency Medicine, Massachusetts General Hospital and Brigham and Women’s Hospital, Boston, MA, USA
Nicholas Peoples
Centre for Evidence-Based Medicine, University of Oxford, Oxford, UK
Ken Milne
Department of Medicine, Western University, London, ON, Canada
Ken Milne
Digital Innovation Research Center, Duke Kunshan University, Kunshan, Jiangsu, China
Bing Luo
Global Health Research Center, Duke Kunshan University, Kunshan, Jiangsu, China
Lijing L. Yan
Global Health Research Center, Duke University, Durham, NC, USA
Lijing L. Yan
Authors
- Nicholas PeoplesView author publications
Search author on:PubMed Google Scholar
- Ken MilneView author publications
Search author on:PubMed Google Scholar
- Bing LuoView author publications
Search author on:PubMed Google Scholar
- Lijing L. YanView author publications
Search author on:PubMed Google Scholar
Contributions
N.P. conceived of the manuscript, including SIMcard and FACTS, and drafted the manuscript. K.M. and L.L.Y. contributed critical input, review, and revision. B.L. as a computer scientist, contributed technical expertise on the feasibility and optimization of the proposal
Ethics declarations
Competing interests
The authors declare no competing interests
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
About this article
Cite this article
Peoples, N., Milne, K., Luo, B. et al. When machines misread science: creating guardrails for human and AI interpretation of biomedical research.
npj Digit. Med.9, 648 (2026). https://doi.org/10.1038/s41746-026-03160-w
Received:27 December 2025
Accepted:12 August 2026
Published:24 August 2026
Version of record:24 August 2026
DOI
:https://doi.org/10.1038/s41746-026-03160-w


