Journal of Arabic Language and Literature

تقييم أداء النماذج اللغوية الضخمة (LLMs)

في كشف وتصنيف الأخطاء الكتابية لدى متعلمي اللغة العربية لغة ثانية

Evaluating the Performance of Large Language Models (LLMs) in Detecting and Classifying Writing Errors among Arabic-as-a-Second-Language Learners

د. خليوي سامر خليوي, العياضي

, ,

أستاذ تعليم اللغة العربية المشارك بمعهد تعليم, اللغة العربية بالجامعة الإسلامية

, ,                                                               البريد الإلكتروني:khlewe@gmail.com

Keywords: الكلمات المفتاحية: النماذج اللغوية الضخمة، تعليم اللغة العربية لغة ثانية، تحليل الأخطاء الكتابية، تقييم الأداء، هندسة الموجّهات.

Major: تعليم اللغة العربية للناطقين بغيرها

Sub Major: التقنيات الحديثة في تعليم اللغة

DOI:10.36046/2356-000-019-019
DownloadPDF
Abstract

الملخص

هدفت هذه الدراسة إلى تقييم أداء النماذج اللغوية الضخمة في كشف وتصنيف الأخطاء  الكتابية لدى متعلمي اللغة العربية لغة ثانية، ولتحقيق ذلك، استخدمت الدراسة المنهج الوصفي التحليلي، واختيرت أربعة نماذج من أجل تقييم أدائها وهي :Gpt-4o  و Gemini 2.5 pro   و DeepSeek v3.2  و ALLaM 34B ، وقد أسفرت نتائج الدراسة عن وجود قصور واضح في قدرة النماذج اللغوية الضخمة على كشف وتصنيف الأخطاء  الكتابية بدقة كافية، حيث بلغ أعلى قيمة أداء محقق  لدى نموذج Gemini 2.5 pro 0.5399 في مقياس  F1-Score، وهي نسبة لا يمكن الاعتماد عليها في سياق التعليم، وفي ضوء هذه النتائج. أوصت الدراسة بضرورة توخي الحذر عند توظيف هذه النماذج في ميدان التعليم بشكل عام وتعليم اللغة بشكل خاص، مع أهمية الإشراف المباشر من المعلم. كما أوصت بالعمل على بناء معايير تقييم دقيقة للنماذج اللغوية الضخمة في ميدان التعليم، وتحسين تصميم الموجّهات (Prompts)  من خلال تجربة أساليب متقدمة مثل سلسلة الأفكار (CoT)؛ طلبا لرفع دقة النماذج وتقليل نسبة الهلوسة.

الكلمات المفتاحية: النماذج اللغوية الضخمة، تعليم اللغة العربية لغة ثانية، تحليل الأخطاء الكتابية، تقييم الأداء، هندسة الموجّهات.


Abstract

This study aimed to evaluate the performance of Large Language Models (LLMs) in identifying written errors made by non-native learners of Arabic. To achieve this, the study employed a descriptive-analytical methodology, selecting four models for performance evaluation: Gpt-4o, Gemini 2.5 pro, DeepSeek v3.2, and ALLaM 34B. The findings revealed significant shortcomings in the ability of LLMs to identify written errors with sufficient accuracy. The highest performance was achieved by the Gemini 2.5 pro model with an F1-Score of 0.5399, a score considered insufficient for full reliance in an educational context. Considering these findings, the study recommended exercising caution when employing these models in the field of education in general and language teaching in particular, emphasizing the importance of direct supervision by the teacher. It also recommends the development of rigorous evaluation criteria for LLMs in the educational domain and improving prompt design by experimenting with advanced techniques such as Chain-of-Thought (CoT) to enhance model accuracy and reduce the incidence of hallucination.


Keywords: Large Language Models (LLMs), Arabic as a Second Language (ASL), Written Error Analysis, Performance Evaluation, Prompt Engineering.


References

المصادر والمراجع:

أولا- المراجع العربية:

الأفيوني، بشار مصطفى، وشالكينسكايا، أرينا. "تحليل الأخطاء اللغوية في الإنتاج الكتابي لدى متعلمي اللغة العربية الروس في جامعة موسكو الحكومية." مجلة الدراسات اللغوية والأدبية، س15، ع2، (2024م): 67–90.

أمين، مروة مصطفى السيد. "أخطاء الكتابة بين متعلمي العربية الناطقين بها والناطقين بغيرها في المرحلة الجامعية الأولى: كلية الألسن جامعة عين شمس نموذجًا." فيلولوجي: سلسلة في الدراسات الأدبية واللغوية، ع65، (2016م): 35–69.

بني عرابة، إخلاص بنت إبراهيم بن حمد، والكاف، فاطمة بنت محمد بن أحمد. "فاعلية بعض تطبيقات الذكاء الاصطناعي في تنمية مهارات القراءة الإبداعية وبقاء أثر التعلم لدى طالبات الصف السادس الأساسي." المجلة العربية للنشر العلمي، ع78، (2025م): 112–131.

بوبنديرة ، عبدالعزيز. "الامتحانات التقليدية ومشكلات التقييم." مجلة البحوث التربوية والتعليمية 13 (عدد خاص)، (2024م): 441–458.

طعيمة، رشدي أحمد. "تعليم العربية لغير الناطقين بها: مناهجه وأساليبه." المنظمة الإسلامية للتربية والعلوم والثقافة (الإيسيسكو)، (1989م)

عبد اللوي، محمد. "توظيف الذكاء الاصطناعي في تدريس اللغة العربية للناطقين بغيرها عن بعد." مجلة التطوير العلمي للدراسات والبحوث، ع17، (2024م): 402–417.

عنتر، دينا. "دمج الذكاء الاصطناعي في تعليم اللغة العربية في التعليم العالي: دراسة حالة على ChatGPT 4 Education." مجلة الأرائك للعلوم والإنسانيات، ع6، (2024م): 467–491.

محسني، فاطمة. "دمج الذكاء الاصطناعي في تعليم اللغة العربية للناطقين بغيرها: إستراتيجيات فعالة لتطوير المهارات اللغوية." مجلة الباحث، (2025م): 1–13.

الملحم، تركي عبد العزيز. "واقع استخدام تطبيقات الهواتف الذكية في تعليم اللغة العربية للناطقين بلغات أخرى في معهد تعليم اللغة العربية لغير الناطقين بها بالجامعة الإسلامية من وجهة نظر المعلمين." مجلة كلية التربية (أسيوط) 37، ع2، (2021م): 39–108.

الهذلول، علي بن هذلول علي. "أثر تطبيقات الذكاء الاصطناعي في تنمية مهارات الكتابة الأكاديمية لدى متعلمي اللغة العربية الناطقين بلغات أخرى." مجلة جامعة الملك عبد العزيز – العلوم التربوية والنفسية، مج4، ع2، (2025م): 50–69.

ثانيا- المراجع الأجنبية:

Al-Alami, Suhair E. "EFL Learners' Attitudes Towards Utilizing ChatGPT for Acquiring Writing Skills in Higher Education: A Case Study of Computing Students." Journal of Language Teaching & Research 15, iss. 4 (2024): 1029.

Al-khafaji, N. J., and B. K. Majeed. "Evaluating Large Language Models Using Arabic Prompts to Generate Python Codes." 2024 4th International Conference on Emerging Smart Technologies and Applications (eSmarTA) (2024.

Al-Khalifa, Shahad; Al-Khalifa, Hend. The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic. arXiv preprint arXiv:2407.00146, 2024.

Alahmadi, M. D., Alharbi, M., Tayeb, A., and M. Al-Shanqiti. "Evaluating Large Language Models' Proficiency in Answering Arabic GAT Exam Questions." Engineering, Technology & Applied Science Research 14, iss. 6 (2024): 17774–17780.

Alfaifi, Abdullah, and Eric Atwell. "Potential Uses of the Arabic Learner Corpus." Paper presented at the Leeds Language, Linguistics and Translation PGR Conference 2013, Leeds, UK (2013).

Al-Husain, A., and A. M. Azmi. "Beyond Event-Centric Narratives: Advancing Arabic Story Generation with Large Language Models and Beam Search." Mathematics 12, no. 10 (2024): 1548.

Almazrouei, E., Cojocaru, R., Baldo, M., Malartic, Q., Alobeidli, H., Mazzotta, D., Penedo, G., Campesan, G., Farooq, M., Alhammadi, M., Launay, J., and Noune, B. "AlGhafa Evaluation Benchmark for Arabic Language Models." Proceedings of the First Arabic Natural Language Processing Conference (ArabicNLP 2023) (2023): 244–275.

Al-Shammari, W., and S. Alhumoud. "TAQS: An Arabic Question Similarity System Using Transfer Learning of BERT with BiLSTM." IEEE Access 10 (2022): 91509–91522.

Ayodele, O. S., Aliu, J., Owoeye, F. O., Ajayi, E. A., and Sheidu, A. Y. "The Role of Artificial Intelligence in Curriculum Development and Management." Journal of Digital Innovations & Contemporary Research in Science, Engineering & Technology 11, no. 2 (2023): 37–46.

Bansal, P. "Prompt Engineering Importance and Applicability with Generative AI." Journal of Computer and Communications 12 (2024): 14–23.

Bari, M. Saiful, et al. Allam: Large language models for Arabic and English. arXiv preprint arXiv:2407.15390, 2024.

Ben Abacha, A., W.-w. Yim, Y. Fu, Z. Sun, M. Yetisgen, F. Xia, and T. Lin. "MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes (arXiv:2412.19260v1)." arXiv preprint (2024).

Biswas, Md. R., Mohsen, F., Shah, Z., and Zaghouani, W. "Potentials of ChatGPT for Annotating Vaccine-Related Tweets." 2023 Tenth International Conference on Social Networks Analysis, Management and Security (SNAMS) (2023): 1-6.

Bonner, Euan; Lege, Ryan; Frazier, Erin. Large Language Model-Based Artificial Intelligence in the Language Classroom: Practical Ideas for Teaching. Teaching English with Technology, )2023(, 23.1: 23-41.

Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., and Xie, X. "A Survey on Evaluation of Large Language Models." ACM Transactions on Intelligent Systems and Technology 15, no. 3 (2024): 1–45.

DeepSeek-AI. "DeepSeek-V3 Technical Report (arXiv:2412.19437v2)." DeepSeek-AI (2025)

Dong, Bingyu, et al. "Large Language Models in Education: A Systematic Review." 2024 6th International Conference on Computer Science and Technologies in Education (CSTE) (2024): 131–134.

Einieh, Y., Almansour, A., and Jamal, A. "Fine-Tuning an AraT5 Transformer for Arabic Abstractive Summarization." 2022 14th International Conference on Computational Intelligence and Communication Networks (CICN) (2022): 194–197.

Ejjami, R. "The Future of Learning: AI-Based Curriculum Development." International Journal for Multidisciplinary Research 6, no. 4 (2024): 1–31.

Fadel, A. S., O. A. Abulnaja, and M. E. Saleh. "Multi-Task Learning Model with Data Augmentation for Arabic Aspect-Based Sentiment Analysis." CMC-Computers, Materials & Continua 75, no. 2 (2023): 4419–4442.

Farraj, K. K. "Tahlil Al-Akhtha 'inda Muta'allimi Al-Lughah Al-'Arabiyah Li Al-Naathiqiina Bi Ghairiha." Arabiyat: Jurnal Pendidikan Bahasa Arab dan Kebahasaaraban 2, no. 1 (2015): 112–129.

Fincham, Naiyi Xie; Alvarez, Aitor Arronte. Using large language models (llms) to facilitate l2 proficiency development through personalized feedback and scaffolding: An empirical study. In: Proceedings of the International CALL Research Conference. 2024. p. 59-64.

Hamaniuk, Vita A. "The Potential of Large Language Models in Language Education." Educational Dimension 5 (2021): 208–210.

Hilal, Mahmud, and Ahmet ?smailo?lu. "Arapça ??retimi G?ren ??rencilerin Yayg?n Dilbilimsel Hatalar? ve C?zümler Uzerine Uygulamal? Alan Cal??mas?: K?r?kkale Universitesi K.K.U.FEF. Arapça M.T.B ?rne?i." K?r?kkale Universitesi Sosyal Bilimler Dergisi 10, no. 2 (July 2020): 731–742.

Jaashan, H. M. S., and A. A. Alashabi. "Using AI Large Language Model (LLM-ChatGPT) to Mitigate Spelling Errors of EFL Learners." Forum for Linguistic Studies 7, no. 3 (2025): 328–339.

Karajeh, O., M. N. Al-Kabi, and E. A. Fox. "Fusing AraBERT and Graph Neural Networks for Enhanced Arabic Text Classification." 2023 24th International Arab Conference on Information Technology (ACIT). IEEE (2023).

Khondaker, M. T. I., Naeem, N., Khan, F. L., Elmadany, A. A., and M. Abdul-Mageed. "Benchmarking LLaMA-3 on Arabic Language Generation Tasks." Proceedings of the Second Arabic Natural Language Processing Conference (2024): 283–297.

Koto, F., Li, H., Shatnawi, S., Doughman, J., Sadallah, A. B., Alraeesi, A., Almubarak, K., Alyafeai, Z., Sengupta, N., Shehata, S., Habash, N., Nakov, P., and T. Baldwin. "ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic." arXiv preprint arXiv:2402.12840 (2024).

Lamsiyah, S., Zeinalipour, K., El Amrany, S., Brust, M., Maggini, M., Bouvry, P., and Schommer, C. "ArabicSense: A Benchmark for Evaluating Commonsense Reasoning in Arabic with Large Language Models." Proceedings of the 4th Workshop on Arabic Corpus Linguistics (WACL-4) (2025): 1–11.

Mohseni, Fatima. “Integrating Artificial Intelligence into Teaching Arabic as a Foreign Language: Effective Strategies for Developing Linguistic Skills.” (in Arabic). Al?B??ith Journal, (2025): 1–13.

Mulyanto, D., Zaki, M., Ridho, A., and Fata, K. "Artificial Intelligence Utilization for Arabic Language Skills Development in Arabic Language Learning." An-Nidzam: Jurnal Manajemen Pendidikan dan Studi Islam 11, no. 1 (2024): 18–32.

Nehar, A., Bellaouar, S., Souffi, S., and Bouameur, M. "Dhati: A Fine-Tuned Large Language Model for Evaluating Subjectivity in Arabic Textual Data." Proceedings of the 2023 5th International Conference on Pattern Analysis and Intelligent Systems (PAIS) (2023): p. 1-7

Panagiotidis, P. "LLM-Based Chatbots in Language Learning." European Journal of Education 7, no. 1 (2024): 102–123.

Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and A. Chadha. "A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications." arXiv preprint (2024).

Van Huyssteen, G. B., E. R. Eiselen, and M. J. Puttkammer. "Re-Evaluating Evaluation Metrics for Spelling Checker Evaluations." In Proceedings of First Workshop on International Proofing Tools and Language Technologies, (2004): 91–99

White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and D. C. Schmidt. "A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT." arXiv preprint (2023).

Yavuz, F., Celik, ?., and G. Yava? Celik. "Utilizing Large Language Models for EFL Essay Grading: An Examination of Reliability and Validity in Rubric-Based Assessments." British Journal of Educational Technology 56 (2025): 150–166.

Zhang, Z., and X. Huang. "The Impact of Chatbots Based on Large Language Models on Second Language Vocabulary Acquisition." Heliyon 10 (2024): e25370.

Zheng, R. "AWE and LLMs in L2 Writing Feedback: An Exploration of EFL Learners’ Experiences." Academic Journal of Humanities & Social Sciences 8, no. 7 (2025): 60–69.

Zheng, Y., Li, T., Huang, H., Zeng, T., Lu, J., Chu, C., Huang, Y., Jiang, Z., Xiong, Q., Ge, Y., and M. Li. "Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models." arXiv preprint (2025).

Zuccon, G., and B. Koopman. "Dr ChatGPT, Tell Me What I Want to Hear: How Prompt Knowledge Impacts Health Answer Correctness." arXiv preprint (2023).

ثالثا- الموقع الالكتروني:

Google DeepMind. "Gemini Model Thinking Updates: Advanced Coding." Google Blog, March 2025. Retrieved October 19, 2025, from https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#advanced-coding.


Bibliography

Abdelloui, Muhammad. “Employing Artificial Intelligence in Teaching Arabic as a Foreign Language Remotely.” (in Arabic). Journal of Scientific Development for Studies and Research, iss. 17 (2024): 402–417.

Afyouni, Bashar Mustafa, & Shalkinskaya, Arina. “Linguistic Error Analysis in the Written Production of Russian Learners of Arabic at Moscow State University.” (in Arabic). Journal of Linguistic and Literary Studies, vol. 15, iss. 2 (2024): 67–90.

al-Hadhl?l, ?Al? ibn Hadhl?l ?Al?. “The Impact of Artificial Intelligence Applications on Developing Academic Writing Skills among Arabic Learners Who Speak Other Languages.” (in Arabic). King Abdulaziz University Journal – Educational and Psychological Sciences, vol. 4, iss. 2 (2025): 50–69.

al-Mul?im, Turk? ?Abd-al-?Az?z. “The Reality of Using Smartphone Applications in Teaching Arabic to Speakers of Other Languages at the Institute for Teaching Arabic to Non?Native Speakers at the Islamic University, from the Teachers’ Perspective.” (in Arabic). Assiut University, Faculty of Education Journal, vol. 37, iss. 2 (2021): 39–108.

Amin, Marwa Mustafa Al?Sayed. “Writing Errors among Native and Non?Native Learners of Arabic in the First University Stage: The Faculty of Al?Alsun, Ain Shams University as a case study.” (in Arabic). Philology: Series in Literary and Linguistic Studies, iss. 65 (2016): 35–69. 

?Antar, D?n?. “Integrating Artificial Intelligence into Teaching Arabic in Higher Education: A Case Study on ChatGPT?4 Education.” Al?Araek Journal for Sciences and Humanities, iss. 6 (2024): 467–491.

Ban? ?Ur?bah, Ikhl?? bint Ibr?h?m ibn ?amad and al-K?f, F??imah bint Mu?ammad ibn A?mad. “The Effectiveness of Selected Artificial Intelligence Applications in Developing Creative Reading Skills and Retention of Learning Among Sixth?Grade Female Students.” (in Arabic). Arab Journal for Scientific Publishing, isss. 78 (2025): 112–131.

Boubandirah, Abdelaziz. “Traditional Examinations and Assessment Problems.” (in Arabic). Journal of Educational and Pedagogical Research, vol. 13 (Special Issue), (2024): 441–458.

Tu‘aimah, Rushdi Ahmad. "Teaching Arabic to Non?Native Speakers: Its Curricula and Methods". (in Arabic). Islamic Educational, Scientific and Cultural Organization (ISESCO), 1989.