Skip to content

arXiv clarifies penalty for uncorrected LLM hallucinations: 1-year ban + peer-review requirement

Tom Dietterich (an arXiv moderator) posted a 5-tweet thread clarifying arXiv’s enforcement of its Code of Conduct on author responsibility for LLM-generated content. By signing as an author, each author takes full responsibility for the paper’s contents regardless of how they were generated. When a submission contains “incontrovertible evidence” that authors did not check LLM output, arXiv now applies a 1-year ban followed by a requirement that subsequent arXiv submissions first be accepted at a reputable peer-reviewed venue. Examples of incontrovertible evidence given: hallucinated references and stray meta-comments from the LLM left in the text (e.g. “here is a 200 word summary; would you like me to make any changes?”, “fill in with the real numbers from your experiments”).

  • Each author of an arXiv submission takes full responsibility for the paper’s contents, irrespective of how the contents were generated [tweet 1/5].
  • Inappropriate language, plagiarism, biased content, errors, incorrect references, and misleading content from generative AI are author responsibility once included in scientific works [tweet 2/5].
  • The penalty for incontrovertible evidence of unchecked LLM output is a 1-year arXiv ban [tweet 4/5].
  • After the ban, the author must first pass peer review at a reputable venue before submitting to arXiv again [tweet 4/5].
  • “Incontrovertible evidence” is illustrated with two example classes: hallucinated references, and meta-comments from the LLM left inline in the manuscript [tweet 5/5].

Not applicable — this is a policy statement from an arXiv moderator, posted as a Twitter/X thread. The thread is the artifact; no underlying paper or blog post is linked. The replies and surrounding context on Dietterich’s timeline (also fetched) are unrelated topics (AGI/Turing critique, conformal prediction, sentience taxonomy) and not part of this announcement.

No empirical results. The thread does not quote enforcement statistics — only the policy clarification and the two illustrative evidence categories.

This is operational signal for any Luma submission that uses LLM assistance for drafting, formatting, or reference management — the failure modes arXiv is flagging (hallucinated references, leftover assistant meta-text) are exactly the failure modes a vibe-coded writing pipeline produces if no one re-reads the output. The framing — author signature implies responsibility regardless of generation method — also matters as a reference point for downstream venues. It connects loosely to Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity in that both touch on the question of what an LLM “knows” vs. fabricates, but here the concern is purely procedural rather than capability-evaluative. Distinct from the misalignment-from-narrow-training thread in Training large language models on narrow tasks can lead to broad misalignment — that’s about model behavior; this is about author hygiene.