Online Math MCQs Are Obsolete: An LLM Benchmarking Framework and Academic Integrity Alternatives
The rapid advancement of LLMs has improved learning but raised concerns about academic integrity in online assessments. This study benchmarks LLM performance on an undergraduate Calculus I MCQ dataset from the University of Houston and examines implications for assessment design. A representative dataset of 331 questions was sampled from 4,900 MCQs and validated with a 0.43% margin of error. Models from OpenAI, Meta, Google, Anthropic, and Alibaba were evaluated using standardized prompts and metrics, including accuracy, precision, recall, and F1-score. Results show substantial capability growth, with accuracy rising from ~63% in 2023 (GPT-3.5) to ~99% across metrics by mid-2025. Most commercial models have reached near-perfect performance, while smaller open-weight models such as GPT-OSS 20B are achieving comparable results. These findings indicate that traditional unproctored online MCQ assessments are no longer reliable measures of student understanding. We argue this is a structural limitation of the format rather than solely misconduct. Alternatives include in-person proctored exams, open-ended questions, AI proctoring tools, and conversational assessments where students explain their reasoning. This work provides a replicable benchmarking framework and highlights the need for assessment redesign.
Presenters
-
Mansib Mursalin
Researcher,
University of Houston