Artificial intelligence is now deeply embedded in how science is written, searched, and summarized, yet the systems doing this work are quietly warping the scholarly record. Large audits of the literature show that fabricated and misdirected citations are no longer rare glitches but a structural problem that is already undermining research reliability in fast-moving and high-stakes fields. A recent forensic audit of 50 AI-assisted survey papers found a 17% phantom rate in 5,514 citations published between late 2024 and early 2026, quantifying how aggressively invalid references are now seeping into AI-written surveys.
AI is quietly warping the scholarly record, turning citation hallucinations into a structural integrity crisis
The quiet flood of fake citations
Over the past few years, generative models have moved from experimental tools to everyday assistants in drafting papers, reviews, and grant applications. That shift coincided with a measurable surge in fake citations that look perfectly plausible yet point to nothing real in the scholarly world. A major audit led by researchers at Cornell and the University of California examined about 111 million references across roughly 2.5 million papers deposited in arXiv, bioRxiv, SSRN, and PubMed Central between 2020 and 2025. They estimated that 146,932 citations in material published during 2025 alone were hallucinated, meaning the referenced article, journal, or digital object did not exist in any major index. This issue has prompted federal safety evaluations to ensure accuracy in AI-generated content.
These are not edge cases hiding in obscure venues. The study found fake citations distributed across leading repositories that together anchor large parts of physics, biology, social science, and biomedicine. Nature’s coverage of the work highlighted that SSRN, a core social science preprint server, showed the highest hallucination rate at roughly 1.91 percent of citations, almost five times the rate observed in any other repository.
ArXiv followed with about 0.39 percent hallucinated citations, PubMed Central with 0.27 percent, and bioRxiv with 0.21 percent.
On a monthly basis, that translates into thousands of bogus scholarly pointers smuggled into the record. One synthesis estimates that in August 2025 alone, PubMed Central hosted about 8,140 hallucinated citations, with thousands more spread across the other platforms. When those papers are later cited by others, the fake references propagate, seeding a trail of misleading authority that is hard to unwind.
How generative models create phantom references
The audits distinguish several patterns of failure rather than treating all hallucinations as the same. There is complete fabrication, where a reference combines realistic author names, titles, and journal labels into an object that simply does not exist.
There are partial fabrications, where some attributes are correct but key elements such as the year, volume, or identifier are wrong enough to make the reference untraceable.
Another category is identifier hijacking. In these cases, the model returns a real-looking digital object identifier that either points to an unrelated paper or does not exist in registries such as CrossRef at all. These patterns collectively erode trust because the references look legitimate at a glance and often pass superficial checks, especially when readers are scanning long bibliographies under time pressure.
A dedicated cross-model audit evaluated ten widely used language models across four academic domains and generated more than 69,000 citation instances that were then checked against multiple scholarly databases. Reported hallucination rates ranged from about 11.4 percent up to 56.8 percent depending on the model, domain, and the way prompts were framed.
The work underscores that citation fabrication is not limited to fringe systems and that prompt design can either mitigate or aggravate the risk.
Clinical and domain-specific citation tests
The problem becomes more serious when the citations are used to support clinical decisions, guidelines, or regulatory policies. In medicine, several groups have stress-tested general-purpose chatbots on targeted referencing tasks and found strikingly low accuracy.
One study that looked at randomized clinical trials in a specific therapeutic area found that Google Bard correctly cited only 13.6 percent of the requested trials, while ChatGPT and Chatsonic achieved 2.4 percent and zero percent respectively. Most references provided by these systems either pointed to unrelated articles or to digital object identifiers that did not resolve to any existing study in major databases.
Other evaluations of ChatGPT-generated medical articles reported that out of 115 references, about 47 percent were fabricated, 46 percent were authentic but inaccurate, and only 7 percent were both authentic and accurate. A specialty-focused assessment in ear, nose, and throat disciplines concluded that a newer model version improved reliability compared with its predecessor, but both variants still produced a mix of erroneous and non-existent references that would mislead readers relying on the citations as a shortcut to the evidence base.
Taken together, these tests show that apparently precise bibliographies can mask very unreliable underlying retrieval, especially when authors assume the model has already done the painstaking verification work.
Structural blind spots in AI literature search
Even when models avoid outright fabrication, their ability to perform comprehensive literature searches is constrained by the data and infrastructure they depend on. Library teams and repository managers have begun issuing guidance that no current AI-powered discovery tool can guarantee full coverage of the relevant literature.
There are multiple reasons for this. Databases are incomplete and unevenly indexed; journals differ in how quickly new articles are deposited; and fast-moving fields produce preprints and conference contributions that may take months to surface in major indexes.
Models tied to static training snapshots or opaque knowledge cutoffs inherit these gaps, which means recent or controversial work often remains invisible even if it is directly relevant to a user’s query.
Specialized or historical material presents another blind spot. Technical standards, older monographs, and niche conference proceedings sometimes sit in separate collections or local archives that are not ingested into widely used training corpora. A model can respond fluently with general explanations while completely missing essential mechanistic or methodological studies that would change the conclusions of a systematic review or meta-analysis.
Citation unfaithfulness and the role of retraction
Beyond missing entire regions of the literature, there is the subtler problem of citation unfaithfulness. Audits of AI-assisted academic writing show that models frequently blur the line between primary experimental work and secondary review or commentary.
They often condense diverse bodies of evidence into broad claims anchored to a small number of high-profile papers, sidelining foundational but less cited studies that actually carry the methodological weight.
Analyses in research integrity journals have argued that what is sometimes described as hallucination is more accurately understood as fabrication and falsification when it comes to references, because the model is generating specific scholarly claims that do not correspond to real work.
When that output is incorporated into grant proposals, doctoral theses, or policy briefs without verification, it shifts the evidentiary base in ways that are hard to detect later.
Retractions add another layer of risk. General-purpose chatbots are usually not wired into real-time retraction databases or journal notices, and they rely heavily on historical training data that can include withdrawn or corrected studies.
Without explicit integration of retraction feeds and robust filtering, these systems may continue to cite invalidated work as if it were fully current, allowing discredited findings to persist in AI-assisted evidence syntheses and summaries.
How repositories and institutions are responding
Major repositories are beginning to take the issue seriously. ArXiv, for example, has tightened its policy on hallucinated references and issued guidance that explicitly warns about ghost or fabricated citations.
The policy describes ghost references as pointers to works that do not exist or that combine authentic-looking metadata into false bibliographic objects, and signals that moderators will pay closer attention to suspicious reference lists.
Complementary research has started to ask a pointed question of the published record itself. One recent study uses citation verifiability to ask how often accepted and peer-reviewed papers contain hallucinated references.
The authors define a citation as hallucinated when the pointer fails at the level of identity, either because no corresponding work can be found or because the best available match has a substantially different author list.
These institutional moves matter because they shift the incentive landscape. When journals and repositories treat fake citations as a serious integrity issue rather than a minor inconvenience, authors and tool builders are more likely to invest in robust verification workflows.
Implications for technology, business, and society
For AI developers, the findings are a clear signal that citation handling is now a core performance dimension, not a peripheral feature.
Systems that can generate impressive text but systematically distort the underlying references are not ready for unsupervised use in scientific or clinical contexts. The audits suggest there is meaningful variance between models and prompts, which means citation evaluation should become part of standard benchmarking alongside accuracy, fairness, and safety metrics.
For businesses building products on top of general models, the risk is reputational as well as operational. Tools marketed as research copilots or evidence assistants that quietly insert fabricated references can erode user trust and expose organizations to legal or regulatory scrutiny, especially in regulated industries such as healthcare and finance where documentation requirements are strict.
Societally, there is a deeper concern about the dilution of expertise. If policymakers, journalists, and practitioners rely on AI-generated summaries that appear well-referenced but in fact draw on a patchy and occasionally imaginary evidence base, public discourse around science can drift away from the best available knowledge.
This matters for urgent topics such as pandemic response, climate policy, and AI governance itself, where decisions depend on careful reading of complex and evolving literatures.
At the same time, there are genuine opportunities if the technology is handled with care. Models can already help researchers navigate unfamiliar fields, surface adjacent work, and draft initial outlines that are later refined by humans.
With the right guardrails, including explicit citation verification and integration with curated databases, AI could reduce the friction of exploratory search while leaving final evidentiary judgments to human experts.
Using AI citation tools safely
Several practical steps are emerging from the research and policy discussions. First, treat citations from general chatbots as hypotheses rather than facts.
Every reference should be checked in trusted databases such as CrossRef, PubMed, or domain-specific indexes before it enters a manuscript or report.
Second, prompt design matters. Studies that model chatbot behavior show that specific and well-structured prompts aimed at reviews or focused questions tend to yield more accurate references than vague general requests.
Framing tasks clearly and asking models to separate primary trials from secondary commentary can reduce unfaithful mixing of evidence.
Third, combine AI tools with established reference managers and discovery services rather than using them as standalone literature search engines.
Automated cross-checking against verified indexes and journal feeds can catch many fabricated or hijacked identifiers before they slip into final documents.
Finally, institutions should update training and supervision practices. Graduate programs, research groups, and clinical organizations need to explicitly teach how to use AI assistants responsibly, including the limitations revealed by recent citation audits and the importance of manual verification.
Key takeaways and what to watch next
The emerging picture is clear. Generative models have introduced a new layer of fragility into the research ecosystem by making it easy to produce authoritative-looking text tied to references that are incomplete, inaccurate, or entirely fabricated.
Large audits across multiple repositories show that these problems are no longer theoretical; tens of thousands of hallucinated citations have already entered the scientific record and many have survived peer review.
Over the next few years, the critical questions will be whether model developers can build systems that are citation aware by design, whether repositories and journals can enforce stronger verification standards without slowing legitimate research, and whether researchers themselves can adapt practices to keep the benefits of AI assistance while guarding the integrity of the scholarly record.
The stakes are high because trust in science depends not only on what is discovered but on how those discoveries are documented and connected.
Getting AI citation behavior right is now part of that trust infrastructure, and it will shape how comfortably society can rely on machine-assisted research in the decade ahead.
Conclusion
Scientific discovery is entering a new phase where models billed as AI scientists can read millions of papers, generate hypotheses, and even propose experiments, yet still miss some of the most important studies in the record. This matters right now because research teams and funders are starting to treat these systems as trusted partners in science, often without realizing how blind their evidence base can be.
From paper deluge to AI scientist pipeline
Over the past decade, the sheer volume of scientific publishing has exploded, and traditional literature reviews have struggled to keep up. That gap opened the door for AI systems that promise to ingest vast corpora and surface the most relevant findings, sometimes framed as autonomous or semi autonomous AI scientists.
The flagship paper on this emerging pipeline describes AI scientist systems that repeatedly select, narrate, and evaluate evidence from the literature to drive new work. In this setup the published record becomes the substrate from which models learn which hypotheses look promising, which findings seem robust, and what directions appear most fruitful. Perplexity Sonar and its Deep Research features are part of this broader trend, offering exhaustive searches across hundreds of sources and long form synthesis that go far beyond classic keyword search.
On paper this looks like a dream for overworked researchers. In practice it turns out the pipeline is built on a deeply uneven foundation.
How publication bias becomes a systems failure
The core problem is not the models themselves but the evidence they are trained and evaluated on. The literature is heavily filtered before it ever reaches an AI system. Positive and striking results are more likely to be written up, more likely to be accepted, and more likely to be cited, while null results, failed replications, and falsified hypotheses often disappear into what has long been called the file drawer.
A classic file drawer study, the TESS project, found strong results were roughly forty percentage points more likely to be published than null findings and sixty percentage points more likely to be written up in the first place. The AI scientist paper formalizes this distortion as the null result gap, the difference between how successful hypotheses look in the corpus and how often they actually survive empirical tests. When this gap is positive, the literature misrepresents scientific reality and any model that learns from it inherits that skew.
The authors argue that AI scientist systems transform publication bias from a slow epistemic tax into a fast systems failure. Instead of bias accumulating gradually across many human careers, it can be amplified instantly when a system trained on skewed corpora is tasked with guiding entire research programs. The risk is that these systems repeatedly select the same flattering evidence, overestimate the reliability of popular claims, and systematically overlook critical studies that challenge dominant narratives.
This pattern does not only affect null results. There is growing evidence that AI tools can miss methodologically rigorous critiques while privileging prestigious consensus reports, especially when asked to rank or summarize literature at scale. In one case study, multiple leading models favored an institutional consensus document over a dissertation that carefully documented statistical assumption violations and measurement problems in the consensus work. Only when explicitly pushed to think about methodology did the systems reverse their judgment, underscoring how easily subtle but crucial critiques can be ignored.
What Perplexity Sonar Deep Research reveals about AI research assistants
Perplexity Sonar and Deep Research are designed to address exactly this challenge of navigating overwhelming literature. Sonar now runs on a large modern model and uses a high performance inference stack to pull from many sources quickly, with Deep Research orchestrating multiple searches and synthesizing detailed reports. Benchmarking suggests these tools can match or exceed competing systems on accuracy and readability for complex research tasks.
At the same time, independent evaluations show that retrieval design and corpus selection strongly influence answer quality. A study of web retrieval assisted language models in neurology found that restricting searches to curated guideline domains improved correctness by roughly eight to eighteen percentage points and cut output variance in half, with the effect most pronounced in a smaller Sonar tier. These results echo the central message of the AI scientist pipeline work. Better evidence improves AI performance, but better evidence requires structural changes in how results are captured and exposed.
The AI scientist paper proposes a three layer governance framework focused on the evidence substrate, evaluation incentives, and publication norms. At the corpus level, it argues that scientific infrastructure should include structured databases of null results and failed replications, with machine readable metadata on hypotheses, protocols, effect sizes, and provenance, so that AI systems can see not only successes but also dead ends. At the evaluation layer, it recommends retraction aware benchmarks that penalize systems for drawing on retracted or unreliable studies. Finally, it calls for training corpus disclosure for AI assisted and AI generated papers, including information about source coverage, null result inclusion, retraction filtering, and knowledge cutoffs, so that reviewers and readers can audit the boundaries of what the system knows.
These proposals align with broader work on AI governance, which shows that high risk post deployment contexts remain underrepresented in safety and reliability research, especially in corporate settings. If AI scientists are trained primarily on pre deployment studies and glossy success stories, they will struggle to anticipate real world failures once systems are deployed.
Comparing human and AI judgment in scientific review
One reassuring thread in the literature is that AI support can sometimes help reduce certain human biases, such as undue focus on author prestige or buzzwords, when used carefully in peer review. Some experiments suggest that AI assistance can push reviewers toward more consistent treatment of null findings and lessen the influence of famous names or fashionable topics.
Yet other studies highlight how AI reviewers can inflate quality ratings and fail to differentiate between human written and AI generated abstracts, indicating a tendency to over trust fluent text regardless of underlying rigor. Combined with the case studies on missed methodological critiques, this paints a mixed picture. AI tools can counter some biases while amplifying others, especially when they are not grounded in well curated corpora and are not evaluated against robust, retraction aware benchmarks.
The net effect in many current pipelines is a narrowing of scientific exploration. Systems trained on over represented success stories tend to rediscover that same literature, prioritize popular directions, and under explore fragile or minority viewpoints. Instead of widening the lens, they risk reinforcing existing biases and blind spots.
Implications for technology, businesses, and society
For technology companies building AI scientist platforms, the message is clear. Sophisticated models and high throughput search are not enough if the underlying evidence is skewed. Competitive advantage will increasingly depend on access to richer, more balanced corpora that capture failures, contested claims, and messy post deployment realities, not just elegant success stories.
For businesses in areas like drug discovery, climate modeling, or financial risk, the stakes are tangible. If AI powered literature pipelines underrepresent null results and critical safety work, they can nudge teams toward overconfident bets, underestimation of side effects, and blind spots in risk management. The danger is not that the models are malicious, but that they are selectively ignorant, steering decisions with an overly flattering picture of the state of the art.
For society and regulators, this research underscores why AI governance cannot stop at model cards and accuracy benchmarks. The evidence base itself needs oversight. Regulators may need to ask not only how a system performs on standard tasks, but also what corpora it relies on, how it handles retractions, and whether null result databases and post deployment case studies are part of its training or retrieval substrate. Without that, audits risk certifying systems that look trustworthy while systematically missing the very studies that warn of their limitations.
There are opportunities as well. AI systems with deep research capabilities can help surface neglected work once the infrastructure to store and expose it exists. If null result repositories and replication registries become standard, models like Sonar could make them more visible to everyday researchers, pushing the community toward more realistic expectations of effect sizes and success rates.
Looking ahead actionable guardrails for AI driven science
The emerging evidence points toward a balanced conclusion. AI scientist systems are powerful accelerators of discovery, but they inherit and can amplify the biases of the scientific record unless that record is deliberately repaired. Strong governance, diversified training corpora, and meaningful human oversight are not optional add ons. They are necessary conditions for trustworthy AI driven science.
In practice this means several concrete shifts. Scientific communities will need to invest in infrastructure that treats null results, failed replications, and critical post deployment case studies as first class citizens rather than background noise. Publishers and conferences can require corpus disclosure for AI assisted submissions and encourage the use of retraction aware, null sensitive benchmarks in evaluation. Research labs can combine tools like Perplexity Sonar Deep Research with explicit practices for checking methodological critiques and minority perspectives, rather than relying solely on whichever studies rise to the top of automated rankings.
Most importantly, teams should treat AI scientists as powerful colleagues, not infallible oracles. When a model returns a neat narrative built from millions of papers, the responsible question is not only whether the story is internally coherent, but also which worlds of evidence are missing from view. The next phase of scientific progress will depend on whether the community can build systems that accelerate discovery without erasing the uncomfortable, inconvenient, but essential parts of scientific knowledge. reddit








