A randomized controlled experiment involving 1,053 university freshmen at Bocconi University revealed that access to GPT-4o significantly boosted student grades on a business assignment, according to a recent research paper. Students using GPT-4o scored nearly a full point higher on a 5-point scale, generating more logically coherent arguments and standard recommendations that aligned closely with expert consensus.
However, the study highlighted a disconnect between traditional grading systems and genuine learning. While a separate teaching intervention focused on causal reasoning pushed students toward novel, falsifiable, and highly diverse solutions, those papers often received lower traditional scores. Standard rubrics heavily rewarded polish and structure—traits easily generated by AI—rather than original or critical thought.
The researchers noted that there was no follow-up test to verify retained knowledge without AI assistance. The results suggest that educational institutions and enterprise evaluators must redesign scoring criteria if they aim to reward actual problem-solving and original thinking over AI-assisted polish.
Why it matters
Current evaluation frameworks often reward AI-generated polish and standard answers rather than critical thinking or true novelty.
EdTech founders and educators must build explicit incentives for originality and causality into assessment platforms.
AI tools enhance output coherence quickly, but explicit human instruction remains necessary to foster diverse problem-solving.
Source: the-decoder.com



