Monday, August 3, 2026

AI & Models

GPTZero finds hallucinated citations in NeurIPS research papers

GPTZero identified 100 hallucinated citations across 51 NeurIPS papers, a finding considered not statistically significant, though organizers note the research is not necessarily invalidated.

GPTZero finds hallucinated citations in NeurIPS research papers
Photo: GPTZero

AI detection startup GPTZero scanned 4,841 papers accepted by the Conference on Neural Information Processing Systems (NeurIPS), a major AI research conference, which took place last month in San Diego. From this scan, the company identified 100 hallucinated citations across 51 papers. While having a paper accepted by NeurIPS is a highly regarded achievement for AI researchers, the findings suggest that even leading experts may be relying on large language models (LLMs) to generate references.

However, caveats abound with this finding. The 100 confirmed hallucinated citations across 51 papers represent 1.1% of the papers, which is not statistically significant given that each paper contains dozens of references. Fortune first reported on the research by GPTZero. In response to the findings, NeurIPS conference organizers emphasized that inaccurate references do not automatically undermine the scientific contributions of the work. As NeurIPS told Fortune, “Even if 1.1% of the papers have one or more incorrect references due to the use of LLMs, the content of the papers themselves [is] not necessarily invalidated.”

Despite the low statistical significance, the presence of AI-fabricated citations highlights a growing challenge for academic publishing. Peer reviewers are instructed to flag hallucinations, but the sheer volume of submissions makes manual verification difficult. GPTZero noted that the goal of its analysis was to provide data on how AI-generated content enters academic literature through a “submission tsunami” that has strained conference review pipelines. The startup pointed to a May 2025 paper titled “The AI Conference Peer Review Crisis,” which previously detailed how these review pipelines at premier conferences, including NeurIPS, are being pushed to their limits by the influx of AI-generated content.

Why it matters

The findings highlight the irony of leading AI experts using LLMs to generate citations, while also underscoring the strain on peer review pipelines due to a “submission tsunami” of AI-generated content.