Saturday, August 22, 2026

ChatGPT identified 142 statistical "issues" in problematic paper

Today the Bayesian statistician Andrew Gelman posted a rather pointed critique of a paper on his blog stating: "Wow! This paper is an absolute clinic in bad quantitative social science research. I’ll leave it as an exercise for the reader to count up all the problems."

Well, ChatGPT-5.6 Sol (Pro) compiled a list of 142 "issues" with the paper far exceeding my paltry efforts. For comparison, Claude (Fable 5 Max) found 16, and Gemini 3.1 Pro was satisfied with 7. Admittedly the ChatGPT list contains a fair amount of repetition, i.e. same error but described in different terms, and nit-picking. 

After prompting from me, ChatGPT grouped the 142 problems into 9 categories.

After more prompting, it mapped common statistical fallacy/bias names on to its 9 categories.

Finally, ChatGPT formulated its own study design. which it contrasted with the methodology in the paper concluding:

"Thus, the paper compared selected deceased athletes with an aggregate population benchmark, whereas a rigorous study would compare defined groups of living and deceased athletes over observed person-time properly handle censoring and truncation, balance confounders, and estimate standardized differences in mortality or survival. The latter could support carefully qualified sport-specific associations; the paper’s method cannot validly estimate how many years a sport adds to or subtracts from life."

I am not qualified to assess the responses from ChatGPT, but I trust that they would be a good starting point for an expert human referee reviewing the paper whose first task would be whittling down the list to remove minor and duplicated errors.

At the very least, it is a valuable statistics lesson for me from my AI statistics tutor.

ChatGPT identified 142 statistical "issues" in problematic paper

Today the Bayesian statistician Andrew Gelman posted a rather pointed critique of a paper on his blog stating: "Wow! This paper is an ...