Elicit: AI Citation Accuracy in Literature Reviews
Citation and Source-Attribution Accuracy in Generic AI Tools: Evidence Across Generations
Executive summary
Empirical evaluations consistently find that generic, non-retrieval LLM outputs can contain a large fraction of fabricated or invalid references, and that the magnitude of the problem depends strongly on the model family, domain, and prompting setup. Early studies of ChatGPT-3.5 reported majority-fabricated bibliographies in both medical QA (69% fabricated references) and literature-review-style writing (55% fabricated works).
Newer generations show meaningful improvements in some settings, but not a universal "fix." In a large 2026 cross-provider audit (69,557 citation instances), observed citation hallucination rates still spanned a wide range (11.4%–56.8%), and within the GPT family specifically the hallucination rate dropped from 45.3% (GPT-4o-mini) to 11.4% (GPT-5-mini). At the same time, a 2026 large-scale, multi-domain analysis (GHOST CITE) verified only 49.71% of extracted citations as valid overall (50.29% invalid), and reported that all tested models hallucinated citations with per-model rates ranging from 14.23% to 94.93%.
For “GPT-5+, Claude Opus 4+, Gemini 2.5/3,” the clearest trend in the provided evidence is that tool use and verification infrastructure are increasingly important for controlling reference validity, but content-level grounding and link validity can remain weak. On HALLUHARD, adding web search reduced reference-grounding failures for GPT-5.2-thinking from 28.1% to 6.4% and for Opus-4.5 from 38.6% to 7.0% (research-questions domain), yet content-grounding failures remained substantial (e.g., 51.6% for GPT-5.2-thinking + web search).
Overall, the evidence supports a nuanced conclusion: newer models can reduce outright fabricated references in some tasks (including large improvements within a model family), but citation reliability remains variable, domain-sensitive, and not guaranteed by “newer model” status alone.
How citation accuracy is measured
The studies in this evidence set operationalize “citation accuracy” in several distinct ways, which is crucial because different metrics penalize different failure modes (made-up sources vs wrong metadata vs wrong attribution). One common approach is reference validity checking, where model-produced citations are extracted and then verified against scholarly databases, with outputs labeled as verified/valid vs hallucinated/invalid at scale. For example, a large 2026 audit verified 69,557 citation instances against CrossRef, OpenAlex, and Semantic Scholar and reported hallucination rates across models and conditions.
Similarly, GHOST CITE extracted 331,809 citations and verified 164,933 (49.71%) as valid while labeling 166,876 (50.29%) as invalid, and also reported per-model hallucination rates spanning 14.23%–94.93% across domains.
A second approach is component-level bibliographic correctness, where a reference is considered correct only if key metadata fields match the real paper. In one medical evaluation, “accurate” required that all four bibliographic elements matched exactly. A related variant is BibTeX-specific grading: a 2026 study scored generated BibTeX entries at the field level (correct, missing, fabricated, partially correct, substituted), reporting 83.6% correct fields overall and 1.8% fabricated fields.
A third family of evaluations targets attribution/grounding, separating whether a citation exists from whether it actually supports the claim it is attached to. HALLUHARD explicitly reports distinct failure rates for “reference” vs “content grounding,” and shows that these can diverge strongly.
Finally, some work evaluates link-level validity for cited URLs in research agents, distinguishing non-resolving links from hallucinated URLs and reports substantial cross-model variation even when models are configured as search-enabled agents.
Baseline
Across 2023–2024 studies, fabricated references were frequently the norm rather than the exception when models were prompted to produce citation-bearing answers or literature-review-style bibliographies without explicit retrieval constraints. In a benchmark analyzing 471 references, hallucination rates were 39.6% (55/139) for GPT-3.5 and 28.6% (34/119) for GPT-4. In the same evaluation, Bard reached 91.4% hallucination.
Work on literature synthesis in 2024 reinforced that even capable general models could be extremely unreliable when asked to cite “up-to-date” literature without retrieval: GPT-4o fabricated citations in 78%–90% of cases across fields like computer science and biomedicine, while retrieval-augmented systems achieved citation accuracy “on par with human experts.”
Practical implications
For literature review tasks, the dominant practical takeaway is that AI-generated bibliographies and in-text citations should be treated as hypotheses to verify, not as references you can safely propagate into a manuscript. Multiple studies show that fabricated references can constitute a large fraction of outputs (e.g., 69% fabricated in a medical QA setting, 55% fabricated works in GPT-3.5 literature-review-style papers).
When tools are available, prefer workflows that explicitly ground citations in retrieval or web search, because the evidence shows large reductions in reference-grounding failures with web search. Operationally, the evidence supports a verification checklist centered on existence, metadata correctness, and attribution:
- Existence and metadata: check that the cited work exists in authoritative indexes and that key bibliographic elements match exactly.
- Attribution and support: check claim–citation alignment, because even systems with lower fabrication can show nontrivial misattribution.
- Links and traceability: validate that links resolve and are not hallucinated, as hallucinated URL rates can reach 13.3% and non-resolving rates 18.5% in evaluated agents/models.
Finally, interpret model-to-model improvements cautiously: while large gains are observed in some comparisons, other large-scale evaluations still show high invalid-citation prevalence overall.