Scientific article
OA Policy
English

Can large language models accurately compute descriptive statistics from structured datasets ? A comparative evaluation of ChatGPT and Claude

ContributorsSeboe, Paul; Wang, Ting
Published inAmerican heart journal plus, vol. 70, 100877
Publication date2026-10
First online date2026-08-27
Abstract

Background: Large language models (LLMs) are increasingly used to support statistical analyses in biomedical research. However, their ability to accurately and reproducibly compute descriptive statistics directly from datasets has received limited evaluation.

Objective: To compare the accuracy, within-modality repeatability, and between-modality consistency of ChatGPT and Claude in generating descriptive statistics from structured datasets provided through different input modalities.

Methods: Two publicly available Stata datasets were evaluated: 'auto.dta' (74 observations, 12 variables) and 'citytemp.dta' (956 observations, 6 variables). Original variable names were replaced with generic labels. Each dataset was analyzed using three input modalities (copy-paste, Word, and Excel), with two independent repetitions per modality. Stata served as the reference standard. Accuracy was assessed for counts of missing and non-missing observations, minima, maxima, means, standard deviations, medians, quartiles, frequencies, and percentages.

Results: ChatGPT and Claude produced identical results across all analyses. For each model, 624 categories of descriptive statistics were evaluated. Exact agreement with the Stata reference standard was observed for 576 of 624 categories (92.3%). The only deviations involved first and third quartiles (Q1-Q3) for eight variables. Post hoc analyses demonstrated that these differences were entirely attributable to the use of a different, but mathematically valid, quartile definition based on linear interpolation rather than computational errors. Within-modality repeatability and between-modality consistency were complete across the evaluated analyses.

Conclusions: ChatGPT and Claude demonstrated excellent accuracy and consistent results across repeated analyses and input modalities for the datasets and descriptive statistics evaluated. After accounting for differences in quartile definitions, no calculation errors were identified across the 624 evaluated categories per model.

Keywords
  • AI
  • Artificial intelligence
  • ChatGPT
  • Claude
  • Descriptive statistic
  • LLM
  • Large language model
  • Research
  • Statistical analysis
Citation (ISO format)
SEBOE, Paul, WANG, Ting. Can large language models accurately compute descriptive statistics from structured datasets ? A comparative evaluation of ChatGPT and Claude. In: American heart journal plus, 2026, vol. 70, p. 100877. doi: 10.1016/j.ahjo.2026.100877
Main files (1)
Article (Published version)
Secondary files (1)
Supplemental data
accessLevelPublic
Identifiers
Journal ISSN2666-6022
4views
12downloads

Technical informations

Creation01/09/2026 07:22:02
First validation18/09/2026 09:19:21
Update18/09/2026 09:19:21
Status update18/09/2026 09:19:21
Last indexation18/09/2026 09:19:23
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack