ChatGPT Raises Grades; Causal-Reasoning Training Raises Originality
In this article
Researchers at Bocconi University, working with OpenAI Economic Research, published results on August 27, 2026 from a randomized experiment involving more than 1,000 first-year undergraduate students. The design isolated four conditions: ChatGPT (GPT‑4o) access alone, causal-reasoning training alone, both interventions together, and a control group receiving neither. Students were assigned by class period rather than self-selected, which sidesteps the confounding that plagues most observational AI-in-education research.
The task was a real-world business case: developing marketing recommendations for Bocconi's own merchandise store. Evaluations ran on two parallel tracks — trained human graders using a five-point rubric, and automated text analysis measuring idea count, idea variety, causal-reasoning signals, and submission similarity to three expert benchmarks. The rubric and the text analysis diverged in instructive ways.
GPT-4o Lifted Rubric Scores and Expert Alignment
Students with ChatGPT access scored almost a full point higher on the five-point grading rubric. Their submissions contained more ideas, exhibited clearer logical flow, and clustered more closely to the expert-written recommendations in text-similarity analysis. Students still had to formulate queries, evaluate model outputs, and decide what to include — the model compressed the expertise gap rather than replacing the student's role. The similarity-to-expert metric operationalizes this directly: GPT‑4o access moved student submissions structurally closer to what domain experts produced, without students having prior marketing training.
Causal-Reasoning Training Produced Orthogonal Gains
The critical-thinking intervention — a standalone exercise teaching causal reasoning through a game, worked examples, questions, and feedback, with no AI component — did not raise rubric scores. The rubric measured only how well recommendations addressed two standard marketing goals (awareness and store usage), and on those dimensions the training moved nothing. The text analysis told a different story: students who completed the exercise produced submissions with higher idea variety and greater distinctiveness relative to peer submissions. They also showed stronger evidence of explaining why an idea might work and identifying conditions under which it might fail — counterfactual reasoning that a conventional rubric built around two fixed outcome categories will not capture.
Combined Condition: Where the Effects Stack
The group that received both GPT‑4o access and the causal-reasoning training showed gains across the widest range of measured dimensions, and the effects did not cancel or suppress each other.
| Condition | Rubric Score Gain | Idea Count | Idea Variety / Distinctiveness | Expert Similarity | Causal Reasoning Signals |
|---|---|---|---|---|---|
| Control (neither) | Baseline | Baseline | Baseline | Baseline | Baseline |
| ChatGPT access only | ~1 point higher | Higher | Not elevated | Higher | Not elevated |
| Causal-reasoning training only | No change | No change | Higher / more distinctive | Not elevated | Higher |
| Both interventions | Similar to ChatGPT-only | Similar to ChatGPT-only | Matched training-only group | Higher | Stronger logical coherence |
Idea variety in the combined group matched the training-only group; rubric scores and idea count matched the ChatGPT-only group. The combined group additionally showed stronger logical coherence and more assumption-questioning behavior than either single-intervention group alone. The additive pattern holds because the two interventions target distinct cognitive operations: the model handles production fluency and structural polish, while the causal-reasoning exercise builds capacity to generate and stress-test original hypotheses.
The Assessment Gap This Exposes
The divergence between rubric scores and text-analysis metrics surfaces a measurement problem with direct implications for how AI capabilities get verified in high-stakes settings. A rubric tuned to two fixed marketing dimensions will reward a well-structured GPT‑4o-assisted answer and miss entirely whether the student produced a genuinely novel idea. As model access becomes more widespread, evaluation instruments calibrated only to polish and coverage will increasingly conflate model capability with student reasoning.
The immediate empirical signal is that LLM access and structured critical-thinking instruction are not substitutes — they address different cognitive deficits and produce measurably different output characteristics. Institutions treating AI access as a threat to originality, or treating critical-thinking curricula as redundant given AI, are misreading what this data shows. Both the rubric gains and the distinctiveness gains are real, they are separable, and they compound.