Evaluating AI Tools for Written Production Assessment in Moroccan Secondary Education: Benchmark and Validation

Authors

  • Souad Mhada UMP, FLSHO
  • Ahmed Oujak UMP, OUJDA

DOI:

https://doi.org/10.37870/mhrx7882

Keywords:

Automated Essay Scoring, Written Production, CEFR, Artificial Intelligence, Large language models, Reliability, validity, benchmark, Moroccan education system

Abstract

The automated scoring of written production by artificial intelligence (AI) tools is advancing rapidly; however, the literature consistently emphasizes the need to verify their reliability, validity, and fairness prior to any deployment. In Moroccan secondary education, teachers rely on the Common European Framework of Reference for Languages (CEFR), while AI tools operate according to heterogeneous logics. We conducted a rigorous benchmark study on 60 authentic student essays (general secondary, 1st and 2nd year baccalaureate), scored using a CEFR-aligned rubric (global score /20 + six criteria rated 0–4), comparing four tools — two multilingual large language models (GPT-4o, Claude 3.7 Sonnet) and two French-language grammar checkers (Grammalecte, LanguageTool) — against a human reference. The primary metric is the Mean Absolute Error (MAE) on the global score; secondary metrics include Quadratic Weighted Kappa (QWK) and 95% bootstrap confidence intervals. Results show that GPT-4o achieves the lowest MAE (3.77) but exhibits a systematic under-scoring bias (−2.68), while Grammalecte offers near-zero centering bias (≈ −0.02). At the criterion level, LLMs converge more strongly on formal dimensions (grammar/spelling) than on discursive ones (cohesion, register), thereby supporting the exploratory hypothesis. The study proposes a reproducible protocol (fixed prompts, JSON schema, CEFR rubric) and evidence-based recommendations for the responsible integration of AI in educational assessment.

Author Biography

  • Ahmed Oujak, UMP, OUJDA

    Professeur Habilité, ENCG Oujda, Université Mohammed Premier

References

AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net

AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net

Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater® V.2. Journal of Technology, Learning, and Assessment, 4(3). https://files.eric.ed.gov/fulltext/EJ843852.pdf

Bennett, R. E. (2004). Moving the field forward: Some thoughts on validity and automated scoring (ETS Research Memorandum RM-04-01). Educational Testing Service. https://www.ets.org/Media/Research/pdf/RM-04-01.pdf

Conseil de l'Europe. (2020). Common European Framework of Reference for Languages: Companion volume. Council of Europe Publishing. https://rm.coe.int/common-european-framework-of-reference-for-languages-learning-teaching/16809ea0d4

CSEFRS. (2015). Pour une école de l'équité, de la qualité et de la promotion : Vision stratégique de la réforme 2015–2030. https://iro.umi.ac.ma/wp-content/uploads/2021/10/ODD-4-A4-Vision-stratégique-de-la-réforme-2015-2030.pdf

Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7, 1–30. https://jmlr.org/papers/volume7/demsar06a/demsar06a.pdf

Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552

Ke, Z., & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI-19) (pp. 6300–6308). https://www.ijcai.org/proceedings/2019/0879.pdf

Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://pmc.ncbi.nlm.nih.gov/articles/PMC4913118/

Lee, S., Cai, Y., Meng, D., Wang, Z., & Wu, Y. (2024). Unleashing large language models' proficiency in zero-shot essay scoring. Findings of the Association for Computational Linguistics: EMNLP 2024, 181–198. https://doi.org/10.18653/v1/2024.findings-emnlp.10

Litman, D., Zhang, H., Correnti, R., Matsumura, L. C., & Wang, E. (2021). A fairness evaluation of automated methods for scoring text evidence usage in writing. In I. Roll et al. (Eds.), Artificial Intelligence in Education (LNAI 12748, pp. 255–267). Springer. https://doi.org/10.1007/978-3-030-78292-4_21

McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282. https://www.biochemia-medica.com/en/journal/22/3/10.11613/BM.2012.031

OCDE. (2024). L'évaluation de la performance des établissements scolaires au Maroc. Organisation de Coopération et de Développement Économiques. https://www.oecd.org/content/dam/oecd/fr/publications/reports/2024/03/l-evaluation-de-la-performance-des-etablissements-scolaires-au-maroc_9031fa8c/4f59bfc1-fr.pdf

Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers & Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234

Powers, D. E., Burstein, J., Chodorow, M., Fowles, M. E., & Kukich, K. (2001). Stumping e-rater: Challenging the validity of automated essay scoring (ETS Research Report RR-01-03). Educational Testing Service. https://www.ets.org/Media/Research/pdf/RR-01-03-Powers.pdf

Royaume du Maroc. (2019). Loi-cadre n° 51-17 relative au système d'éducation, de formation et de recherche scientifique (9 août 2019). https://planipolis.iiep.unesco.org/sites/default/files/ressources/morocco_la-loi-cadre-17-51-fr.pdf

Schaller, N.-J., Ding, Y., Horbach, A., Meyer, J., & Jansen, T. (2024). Fairness in automated essay scoring: A comparative analysis of algorithms on German learner essays from secondary education. In Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA) (pp. 210–221). https://aclanthology.org/2024.bea-1.18.pdf

Zesch, T., Wojatzki, M., & Scholten-Akoun, D. (2015). Task-independent features for automated essay grading. In Proceedings of the 10th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 224–232). https://aclanthology.org/W15-0626/

Published

2026-08-30

How to Cite

Mhada, S., & Oujak, A. (2026). Evaluating AI Tools for Written Production Assessment in Moroccan Secondary Education: Benchmark and Validation. The Journal of Quality in Education, 16(28), 92-101. https://doi.org/10.37870/mhrx7882

Similar Articles

1-10 of 90

You may also start an advanced similarity search for this article.