Évaluation des outils d'IA pour la production écrite au lycée marocain : Benchmark et validation.
DOI :
https://doi.org/10.37870/mhrx7882Mots-clés :
Evaluation automatisée, Production écrite, Lycée, CECRL, Intelligence artificielle, Fiabilité, Validité, Benchmark, LLMRésumé
La correction automatisée des productions écrites par l'intelligence artificielle (IA) progresse rapidement, mais la littérature souligne la nécessité d'en vérifier la fiabilité, la validité et l'équité avant tout déploiement. Dans le secondaire marocain, les enseignants s'appuient sur le référentiel CECRL, alors que les outils IA mobilisent des logiques hétérogènes. Nous avons conduit un benchmark rigoureux sur 60 copies de lycéens (tronc commun, 1ʳᵉ et 2ᵉ BAC), notées selon une grille CECRL (score global /20 + six critères 0–4), afin de comparer quatre outils — deux LLM multilingues (GPT-4o, Claude 3.7 Sonnet) et deux correcteurs francophones (Grammalecte, LanguageTool) — à une référence humaine. Le critère principal est la MAE (Mean Absolute Error) sur la note globale ; les critères secondaires mobilisent le kappa pondéré quadratique (QWK) et des intervalles de confiance à 95 % par bootstrap. Les résultats montrent que GPT-4o affiche la MAE la plus basse (3,77) mais présente un biais de sous-notation (−2,68), tandis que Grammalecte offre un centrage quasi nul (biais ≈ −0,02). Sur les critères, les LLM convergent davantage sur les dimensions formelles (grammaire/orthographe) que sur les dimensions discursives (cohésion, registre), validant ainsi l'hypothèse exploratoire. L'étude propose un protocole reproductible (prompts figés, schéma JSON, grille CECRL) et des recommandations d'usage pour une intégration responsable de l'IA dans l'évaluation.
Références
AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net
AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.testingstandards.net
Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater® V.2. Journal of Technology, Learning, and Assessment, 4(3). https://files.eric.ed.gov/fulltext/EJ843852.pdf
Bennett, R. E. (2004). Moving the field forward: Some thoughts on validity and automated scoring (ETS Research Memorandum RM-04-01). Educational Testing Service. https://www.ets.org/Media/Research/pdf/RM-04-01.pdf
Conseil de l'Europe. (2020). Common European Framework of Reference for Languages: Companion volume. Council of Europe Publishing. https://rm.coe.int/common-european-framework-of-reference-for-languages-learning-teaching/16809ea0d4
CSEFRS. (2015). Pour une école de l'équité, de la qualité et de la promotion : Vision stratégique de la réforme 2015–2030. https://iro.umi.ac.ma/wp-content/uploads/2021/10/ODD-4-A4-Vision-stratégique-de-la-réforme-2015-2030.pdf
Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7, 1–30. https://jmlr.org/papers/volume7/demsar06a/demsar06a.pdf
Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552
Ke, Z., & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI-19) (pp. 6300–6308). https://www.ijcai.org/proceedings/2019/0879.pdf
Koo, T. K., & Li, M. Y. (2016). A guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15(2), 155–163. https://pmc.ncbi.nlm.nih.gov/articles/PMC4913118/
Lee, S., Cai, Y., Meng, D., Wang, Z., & Wu, Y. (2024). Unleashing large language models' proficiency in zero-shot essay scoring. Findings of the Association for Computational Linguistics: EMNLP 2024, 181–198. https://doi.org/10.18653/v1/2024.findings-emnlp.10
Litman, D., Zhang, H., Correnti, R., Matsumura, L. C., & Wang, E. (2021). A fairness evaluation of automated methods for scoring text evidence usage in writing. In I. Roll et al. (Eds.), Artificial Intelligence in Education (LNAI 12748, pp. 255–267). Springer. https://doi.org/10.1007/978-3-030-78292-4_21
McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282. https://www.biochemia-medica.com/en/journal/22/3/10.11613/BM.2012.031
OCDE. (2024). L'évaluation de la performance des établissements scolaires au Maroc. Organisation de Coopération et de Développement Économiques. https://www.oecd.org/content/dam/oecd/fr/publications/reports/2024/03/l-evaluation-de-la-performance-des-etablissements-scolaires-au-maroc_9031fa8c/4f59bfc1-fr.pdf
Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers & Education: Artificial Intelligence, 6, 100234. https://doi.org/10.1016/j.caeai.2024.100234
Powers, D. E., Burstein, J., Chodorow, M., Fowles, M. E., & Kukich, K. (2001). Stumping e-rater: Challenging the validity of automated essay scoring (ETS Research Report RR-01-03). Educational Testing Service. https://www.ets.org/Media/Research/pdf/RR-01-03-Powers.pdf
Royaume du Maroc. (2019). Loi-cadre n° 51-17 relative au système d'éducation, de formation et de recherche scientifique (9 août 2019). https://planipolis.iiep.unesco.org/sites/default/files/ressources/morocco_la-loi-cadre-17-51-fr.pdf
Schaller, N.-J., Ding, Y., Horbach, A., Meyer, J., & Jansen, T. (2024). Fairness in automated essay scoring: A comparative analysis of algorithms on German learner essays from secondary education. In Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA) (pp. 210–221). https://aclanthology.org/2024.bea-1.18.pdf
Zesch, T., Wojatzki, M., & Scholten-Akoun, D. (2015). Task-independent features for automated essay grading. In Proceedings of the 10th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 224–232). https://aclanthology.org/W15-0626/
Téléchargements
Publié
Numéro
Rubrique
Catégories
Licence
© La Revue de la Qualité en Education 2026

Cette œuvre est sous licence Creative Commons Attribution - Pas d'Utilisation Commerciale - Partage dans les Mêmes Conditions 4.0 International.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
