Online Math MCQs Are Obsolete: An LLM Benchmarking Framework and Academic Integrity Alternatives

Wednesday, September 30, 2026 | 9:30AM–10:30AM MT
Session Type: Poster Session
Delivery Format: Poster Session
The rapid advancement of LLMs has improved learning but raised concerns about academic integrity in online assessments. This study benchmarks LLM performance on an undergraduate Calculus I MCQ dataset from the University of Houston and examines implications for assessment design. A representative dataset of 331 questions was sampled from 4,900 MCQs and validated with a 0.43% margin of error. Models from OpenAI, Meta, Google, Anthropic, and Alibaba were evaluated using standardized prompts and metrics, including accuracy, precision, recall, and F1-score. Results show substantial capability growth, with accuracy rising from ~63% in 2023 (GPT-3.5) to ~99% across metrics by mid-2025. Most commercial models have reached near-perfect performance, while smaller open-weight models such as GPT-OSS 20B are achieving comparable results. These findings indicate that traditional unproctored online MCQ assessments are no longer reliable measures of student understanding. We argue this is a structural limitation of the format rather than solely misconduct. Alternatives include in-person proctored exams, open-ended questions, AI proctoring tools, and conversational assessments where students explain their reasoning. This work provides a replicable benchmarking framework and highlights the need for assessment redesign.

Presenters

  • Mansib Mursalin

    Researcher, University of Houston