Evaluating AI Tools for Written Production Assessment in Moroccan Secondary Education: Benchmark and Validation
DOI:
https://doi.org/10.37870/mhrx7882Keywords:
Automated Essay Scoring, Written Production, CEFR, Artificial Intelligence, Large language models, Reliability, validity, benchmark, Moroccan education systemAbstract
The automated scoring of written production by artificial intelligence (AI) tools is advancing rapidly; however, the literature consistently emphasizes the need to verify their reliability, validity, and fairness prior to any deployment. In Moroccan secondary education, teachers rely on the Common European Framework of Reference for Languages (CEFR), while AI tools operate according to heterogeneous logics. We conducted a rigorous benchmark study on 60 authentic student essays (general secondary, 1st and 2nd year baccalaureate), scored using a CEFR-aligned rubric (global score /20 + six criteria rated 0–4), comparing four tools — two multilingual large language models (GPT-4o, Claude 3.7 Sonnet) and two French-language grammar checkers (Grammalecte, LanguageTool) — against a human reference. The primary metric is the Mean Absolute Error (MAE) on the global score; secondary metrics include Quadratic Weighted Kappa (QWK) and 95% bootstrap confidence intervals. Results show that GPT-4o achieves the lowest MAE (3.77) but exhibits a systematic under-scoring bias (−2.68), while Grammalecte offers near-zero centering bias (≈ −0.02). At the criterion level, LLMs converge more strongly on formal dimensions (grammar/spelling) than on discursive ones (cohesion, register), thereby supporting the exploratory hypothesis. The study proposes a reproducible protocol (fixed prompts, JSON schema, CEFR rubric) and evidence-based recommendations for the responsible integration of AI in educational assessment.
References
AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net
AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net
Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater® V.2. Journal of Technology, Learning, and Assessment, 4(3). https://files.eric.ed.gov/fulltext/EJ843852.pdf
Bennett, R. E. (2004). Moving the field forward: Some thoughts on validity and automated scoring (ETS Research Memorandum RM-04-01). Educational Testing Service. https://www.ets.org/Media/Research/pdf/RM-04-01.pdf
Conseil de l'Europe. (2020). Common European Framework of Reference for Languages: Companion volume. Council of Europe Publishing. https://rm.coe.int/common-european-framework-of-reference-for-languages-learning-teaching/16809ea0d4
CSEFRS. (2015). Pour une école de l'équité, de la qualité et de la promotion : Vision stratégique de la réforme 2015–2030. https://iro.umi.ac.ma/wp-content/uploads/2021/10/ODD-4-A4-Vision-stratégique-de-la-réforme-2015-2030.pdf
Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7, 1–30. https://jmlr.org/papers/volume7/demsar06a/demsar06a.pdf
Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552
Ke, Z., & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI-19) (pp. 6300–6308). https://www.ijcai.org/proceedings/2019/0879.pdf
Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://pmc.ncbi.nlm.nih.gov/articles/PMC4913118/
Lee, S., Cai, Y., Meng, D., Wang, Z., & Wu, Y. (2024). Unleashing large language models' proficiency in zero-shot essay scoring. Findings of the Association for Computational Linguistics: EMNLP 2024, 181–198. https://doi.org/10.18653/v1/2024.findings-emnlp.10
Litman, D., Zhang, H., Correnti, R., Matsumura, L. C., & Wang, E. (2021). A fairness evaluation of automated methods for scoring text evidence usage in writing. In I. Roll et al. (Eds.), Artificial Intelligence in Education (LNAI 12748, pp. 255–267). Springer. https://doi.org/10.1007/978-3-030-78292-4_21
McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282. https://www.biochemia-medica.com/en/journal/22/3/10.11613/BM.2012.031
OCDE. (2024). L'évaluation de la performance des établissements scolaires au Maroc. Organisation de Coopération et de Développement Économiques. https://www.oecd.org/content/dam/oecd/fr/publications/reports/2024/03/l-evaluation-de-la-performance-des-etablissements-scolaires-au-maroc_9031fa8c/4f59bfc1-fr.pdf
Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers & Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234
Powers, D. E., Burstein, J., Chodorow, M., Fowles, M. E., & Kukich, K. (2001). Stumping e-rater: Challenging the validity of automated essay scoring (ETS Research Report RR-01-03). Educational Testing Service. https://www.ets.org/Media/Research/pdf/RR-01-03-Powers.pdf
Royaume du Maroc. (2019). Loi-cadre n° 51-17 relative au système d'éducation, de formation et de recherche scientifique (9 août 2019). https://planipolis.iiep.unesco.org/sites/default/files/ressources/morocco_la-loi-cadre-17-51-fr.pdf
Schaller, N.-J., Ding, Y., Horbach, A., Meyer, J., & Jansen, T. (2024). Fairness in automated essay scoring: A comparative analysis of algorithms on German learner essays from secondary education. In Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA) (pp. 210–221). https://aclanthology.org/2024.bea-1.18.pdf
Zesch, T., Wojatzki, M., & Scholten-Akoun, D. (2015). Task-independent features for automated essay grading. In Proceedings of the 10th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 224–232). https://aclanthology.org/W15-0626/
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 The Journal of Quality in Education

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
