ChatGPT Raises Grades; Causal-Reasoning Training Raises Originality

August 28, 2026news

Researchers at Bocconi University, working with OpenAI Economic Research, published results on August 27, 2026 from a randomized experiment involving more than 1,000 first-year undergraduate students. The design isolated four conditions: ChatGPT (GPT‑4o) access alone, causal-reasoning training alone, both interventions together, and a control group receiving neither. Students were assigned by class period rather than self-selected, which sidesteps the confounding that plagues most observational AI-in-education research.

The task was a real-world business case: developing marketing recommendations for Bocconi's own merchandise store. Evaluations ran on two parallel tracks — trained human graders using a five-point rubric, and automated text analysis measuring idea count, idea variety, causal-reasoning signals, and submission similarity to three expert benchmarks. The rubric and the text analysis diverged in instructive ways.

GPT-4o Lifted Rubric Scores and Expert Alignment

Students with ChatGPT access scored almost a full point higher on the five-point grading rubric. Their submissions contained more ideas, exhibited clearer logical flow, and clustered more closely to the expert-written recommendations in text-similarity analysis. Students still had to formulate queries, evaluate model outputs, and decide what to include — the model compressed the expertise gap rather than replacing the student's role. The similarity-to-expert metric operationalizes this directly: GPT‑4o access moved student submissions structurally closer to what domain experts produced, without students having prior marketing training.

Causal-Reasoning Training Produced Orthogonal Gains

The critical-thinking intervention — a standalone exercise teaching causal reasoning through a game, worked examples, questions, and feedback, with no AI component — did not raise rubric scores. The rubric measured only how well recommendations addressed two standard marketing goals (awareness and store usage), and on those dimensions the training moved nothing. The text analysis told a different story: students who completed the exercise produced submissions with higher idea variety and greater distinctiveness relative to peer submissions. They also showed stronger evidence of explaining why an idea might work and identifying conditions under which it might fail — counterfactual reasoning that a conventional rubric built around two fixed outcome categories will not capture.

Combined Condition: Where the Effects Stack

The group that received both GPT‑4o access and the causal-reasoning training showed gains across the widest range of measured dimensions, and the effects did not cancel or suppress each other.

Condition Rubric Score Gain Idea Count Idea Variety / Distinctiveness Expert Similarity Causal Reasoning Signals
Control (neither) Baseline Baseline Baseline Baseline Baseline
ChatGPT access only ~1 point higher Higher Not elevated Higher Not elevated
Causal-reasoning training only No change No change Higher / more distinctive Not elevated Higher
Both interventions Similar to ChatGPT-only Similar to ChatGPT-only Matched training-only group Higher Stronger logical coherence

Idea variety in the combined group matched the training-only group; rubric scores and idea count matched the ChatGPT-only group. The combined group additionally showed stronger logical coherence and more assumption-questioning behavior than either single-intervention group alone. The additive pattern holds because the two interventions target distinct cognitive operations: the model handles production fluency and structural polish, while the causal-reasoning exercise builds capacity to generate and stress-test original hypotheses.

The Assessment Gap This Exposes

The divergence between rubric scores and text-analysis metrics surfaces a measurement problem with direct implications for how AI capabilities get verified in high-stakes settings. A rubric tuned to two fixed marketing dimensions will reward a well-structured GPT‑4o-assisted answer and miss entirely whether the student produced a genuinely novel idea. As model access becomes more widespread, evaluation instruments calibrated only to polish and coverage will increasingly conflate model capability with student reasoning.

The immediate empirical signal is that LLM access and structured critical-thinking instruction are not substitutes — they address different cognitive deficits and produce measurably different output characteristics. Institutions treating AI access as a threat to originality, or treating critical-thinking curricula as redundant given AI, are misreading what this data shows. Both the rubric gains and the distinctiveness gains are real, they are separable, and they compound.