تقييم أداء النماذج اللغوية
الضخمة (LLMs)
في كشف وتصنيف الأخطاء الكتابية لدى متعلمي اللغة العربية لغة ثانية
Evaluating the Performance of Large Language Models (LLMs) in Detecting and Classifying Writing Errors among Arabic-as-a-Second-Language Learners
د. خليوي سامر خليوي, العياضي
أستاذ تعليم اللغة العربية المشارك بمعهد تعليم, اللغة العربية بالجامعة الإسلامية
الملخص
هدفت هذه
الدراسة إلى تقييم أداء النماذج اللغوية الضخمة في كشف وتصنيف الأخطاء الكتابية لدى متعلمي اللغة العربية لغة ثانية،
ولتحقيق ذلك، استخدمت الدراسة المنهج الوصفي التحليلي، واختيرت أربعة نماذج من أجل
تقييم أدائها وهي :Gpt-4o و Gemini 2.5 pro و DeepSeek
v3.2 و ALLaM
34B ، وقد أسفرت نتائج الدراسة
عن وجود قصور واضح في قدرة النماذج اللغوية الضخمة على كشف وتصنيف الأخطاء الكتابية بدقة كافية، حيث بلغ أعلى قيمة أداء
محقق لدى نموذج Gemini
2.5 pro 0.5399 في
مقياس F1-Score،
وهي نسبة لا يمكن الاعتماد عليها في سياق التعليم، وفي ضوء هذه النتائج. أوصت
الدراسة بضرورة توخي الحذر عند توظيف هذه النماذج في ميدان التعليم بشكل عام وتعليم
اللغة بشكل خاص، مع أهمية الإشراف المباشر من المعلم. كما أوصت بالعمل على بناء
معايير تقييم دقيقة للنماذج اللغوية الضخمة في ميدان التعليم، وتحسين تصميم الموجّهات (Prompts) من خلال تجربة أساليب متقدمة مثل سلسلة الأفكار (CoT)؛ طلبا
لرفع دقة النماذج وتقليل نسبة الهلوسة.
الكلمات المفتاحية: النماذج اللغوية الضخمة، تعليم اللغة العربية لغة ثانية، تحليل الأخطاء الكتابية، تقييم الأداء، هندسة الموجّهات.
Abstract
This study aimed to evaluate the performance of Large Language Models (LLMs) in identifying written errors made by non-native learners of Arabic. To achieve this, the study employed a descriptive-analytical methodology, selecting four models for performance evaluation: Gpt-4o, Gemini 2.5 pro, DeepSeek v3.2, and ALLaM 34B. The findings revealed significant shortcomings in the ability of LLMs to identify written errors with sufficient accuracy. The highest performance was achieved by the Gemini 2.5 pro model with an F1-Score of 0.5399, a score considered insufficient for full reliance in an educational context. Considering these findings, the study recommended exercising caution when employing these models in the field of education in general and language teaching in particular, emphasizing the importance of direct supervision by the teacher. It also recommends the development of rigorous evaluation criteria for LLMs in the educational domain and improving prompt design by experimenting with advanced techniques such as Chain-of-Thought (CoT) to enhance model accuracy and reduce the incidence of hallucination.
Keywords: Large Language Models (LLMs), Arabic as a Second Language (ASL), Written Error Analysis, Performance Evaluation, Prompt Engineering.
المصادر
والمراجع:
أولا-
المراجع العربية:
الأفيوني، بشار مصطفى، وشالكينسكايا، أرينا.
"تحليل الأخطاء اللغوية في الإنتاج الكتابي لدى متعلمي اللغة العربية الروس
في جامعة موسكو الحكومية." مجلة الدراسات اللغوية والأدبية، س15، ع2،
(2024م): 67–90.
أمين، مروة مصطفى السيد. "أخطاء الكتابة
بين متعلمي العربية الناطقين بها والناطقين بغيرها في المرحلة الجامعية الأولى:
كلية الألسن جامعة عين شمس نموذجًا." فيلولوجي: سلسلة في الدراسات الأدبية
واللغوية، ع65، (2016م): 35–69.
بني عرابة، إخلاص بنت إبراهيم بن حمد، والكاف،
فاطمة بنت محمد بن أحمد. "فاعلية بعض تطبيقات الذكاء الاصطناعي في تنمية
مهارات القراءة الإبداعية وبقاء أثر التعلم لدى طالبات الصف السادس الأساسي."
المجلة العربية للنشر العلمي، ع78، (2025م): 112–131.
بوبنديرة ، عبدالعزيز. "الامتحانات
التقليدية ومشكلات التقييم." مجلة البحوث التربوية والتعليمية 13 (عدد خاص)،
(2024م): 441–458.
طعيمة، رشدي أحمد. "تعليم العربية لغير
الناطقين بها: مناهجه وأساليبه." المنظمة الإسلامية للتربية والعلوم والثقافة
(الإيسيسكو)، (1989م)
عبد اللوي، محمد. "توظيف الذكاء الاصطناعي
في تدريس اللغة العربية للناطقين بغيرها عن بعد." مجلة التطوير العلمي
للدراسات والبحوث، ع17، (2024م): 402–417.
عنتر، دينا. "دمج الذكاء الاصطناعي في
تعليم اللغة العربية في التعليم العالي: دراسة حالة على ChatGPT 4 Education." مجلة
الأرائك للعلوم والإنسانيات، ع6، (2024م): 467–491.
محسني، فاطمة. "دمج الذكاء الاصطناعي في
تعليم اللغة العربية للناطقين بغيرها: إستراتيجيات فعالة لتطوير المهارات
اللغوية." مجلة الباحث، (2025م): 1–13.
الملحم، تركي عبد العزيز. "واقع استخدام
تطبيقات الهواتف الذكية في تعليم اللغة العربية للناطقين بلغات أخرى في معهد تعليم
اللغة العربية لغير الناطقين بها بالجامعة الإسلامية من وجهة نظر المعلمين."
مجلة كلية التربية (أسيوط) 37، ع2، (2021م): 39–108.
الهذلول، علي بن هذلول علي. "أثر تطبيقات
الذكاء الاصطناعي في تنمية مهارات الكتابة الأكاديمية لدى متعلمي اللغة العربية
الناطقين بلغات أخرى." مجلة جامعة الملك عبد العزيز – العلوم التربوية
والنفسية، مج4، ع2، (2025م): 50–69.
ثانيا-
المراجع الأجنبية:
Al-Alami, Suhair E. "EFL Learners' Attitudes Towards
Utilizing ChatGPT for Acquiring Writing Skills in Higher Education: A Case
Study of Computing Students." Journal of Language
Teaching & Research 15, iss. 4 (2024): 1029.
Al-khafaji, N. J., and B. K.
Majeed. "Evaluating Large Language Models Using Arabic Prompts to Generate
Python Codes." 2024 4th International Conference on
Emerging Smart Technologies and Applications (eSmarTA) (2024.
Al-Khalifa, Shahad; Al-Khalifa,
Hend. The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language
Understanding in Arabic. arXiv preprint arXiv:2407.00146, 2024.
Alahmadi, M. D., Alharbi, M., Tayeb, A., and M. Al-Shanqiti.
"Evaluating Large Language Models' Proficiency in Answering Arabic GAT
Exam Questions." Engineering, Technology & Applied
Science Research 14, iss. 6 (2024): 17774–17780.
Alfaifi, Abdullah, and Eric Atwell. "Potential Uses
of the Arabic Learner Corpus." Paper presented at the Leeds
Language, Linguistics and Translation PGR Conference 2013, Leeds,
UK (2013).
Al-Husain, A., and A. M. Azmi. "Beyond Event-Centric
Narratives: Advancing Arabic Story Generation with Large Language Models and
Beam Search." Mathematics 12, no. 10
(2024): 1548.
Almazrouei, E., Cojocaru, R., Baldo, M., Malartic, Q.,
Alobeidli, H., Mazzotta, D., Penedo, G., Campesan, G., Farooq, M., Alhammadi,
M., Launay, J., and Noune, B. "AlGhafa Evaluation Benchmark for Arabic
Language Models." Proceedings of the First Arabic Natural
Language Processing Conference (ArabicNLP 2023) (2023): 244–275.
Al-Shammari, W., and S. Alhumoud. "TAQS: An Arabic
Question Similarity System Using Transfer Learning of BERT with BiLSTM." IEEE
Access 10 (2022): 91509–91522.
Ayodele, O. S., Aliu, J., Owoeye, F. O., Ajayi, E. A., and
Sheidu, A. Y. "The Role of Artificial Intelligence in Curriculum
Development and Management." Journal of Digital
Innovations & Contemporary Research in Science, Engineering &
Technology 11, no. 2 (2023): 37–46.
Bansal, P. "Prompt
Engineering Importance and Applicability with Generative AI." Journal
of Computer and Communications 12 (2024): 14–23.
Bari, M.
Saiful, et al. Allam: Large language models for Arabic and English. arXiv
preprint arXiv:2407.15390, 2024.
Ben Abacha, A., W.-w. Yim, Y. Fu, Z. Sun, M. Yetisgen, F.
Xia, and T. Lin. "MEDEC: A Benchmark for Medical Error Detection and
Correction in Clinical Notes (arXiv:2412.19260v1)." arXiv
preprint (2024).
Biswas, Md. R., Mohsen, F., Shah, Z., and Zaghouani, W.
"Potentials of ChatGPT for Annotating Vaccine-Related Tweets." 2023
Tenth International Conference on Social Networks Analysis, Management and
Security (SNAMS) (2023):
1-6.
Bonner, Euan; Lege, Ryan; Frazier, Erin. Large Language
Model-Based Artificial Intelligence in the Language Classroom: Practical Ideas
for Teaching. Teaching English with Technology, )2023(, 23.1: 23-41.
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K.,
and Xie, X. "A Survey on Evaluation of Large Language Models." ACM
Transactions on Intelligent Systems and Technology
15, no. 3 (2024): 1–45.
DeepSeek-AI. "DeepSeek-V3 Technical Report
(arXiv:2412.19437v2)." DeepSeek-AI (2025)
Dong, Bingyu, et al. "Large Language Models in
Education: A Systematic Review." 2024 6th International
Conference on Computer Science and Technologies in Education (CSTE)
(2024): 131–134.
Einieh, Y., Almansour, A., and Jamal, A. "Fine-Tuning
an AraT5 Transformer for Arabic Abstractive Summarization." 2022
14th International Conference on Computational Intelligence and Communication
Networks (CICN) (2022): 194–197.
Ejjami, R. "The Future of Learning: AI-Based
Curriculum Development." International Journal for
Multidisciplinary Research 6, no. 4 (2024): 1–31.
Fadel, A. S., O. A. Abulnaja, and M. E. Saleh.
"Multi-Task Learning Model with Data Augmentation for Arabic Aspect-Based
Sentiment Analysis." CMC-Computers,
Materials & Continua 75, no. 2 (2023): 4419–4442.
Farraj, K. K.
"Tahlil Al-Akhtha 'inda Muta'allimi Al-Lughah Al-'Arabiyah Li
Al-Naathiqiina Bi Ghairiha." Arabiyat:
Jurnal Pendidikan Bahasa Arab dan Kebahasaaraban 2, no. 1 (2015): 112–129.
Fincham, Naiyi Xie; Alvarez, Aitor Arronte. Using large language
models (llms) to facilitate l2 proficiency development through personalized
feedback and scaffolding: An empirical study. In: Proceedings of the
International CALL Research Conference. 2024. p. 59-64.
Hamaniuk, Vita A. "The Potential of Large Language
Models in Language Education." Educational Dimension
5 (2021): 208–210.
Hilal, Mahmud, and Ahmet ?smailo?lu. "Arapça ??retimi
G?ren ??rencilerin Yayg?n Dilbilimsel Hatalar? ve C?zümler Uzerine Uygulamal?
Alan Cal??mas?: K?r?kkale Universitesi K.K.U.FEF. Arapça M.T.B ?rne?i." K?r?kkale
Universitesi Sosyal Bilimler Dergisi 10, no. 2 (July 2020):
731–742.
Jaashan, H. M. S., and A. A.
Alashabi. "Using AI Large Language Model (LLM-ChatGPT) to Mitigate
Spelling Errors of EFL Learners." Forum for Linguistic
Studies 7, no. 3 (2025): 328–339.
Karajeh, O., M. N. Al-Kabi, and E. A. Fox. "Fusing
AraBERT and Graph Neural Networks for Enhanced Arabic Text
Classification." 2023 24th International Arab Conference
on Information Technology (ACIT). IEEE (2023).
Khondaker, M. T. I., Naeem, N., Khan, F. L., Elmadany, A.
A., and M. Abdul-Mageed. "Benchmarking LLaMA-3 on Arabic Language
Generation Tasks." Proceedings of the Second Arabic Natural
Language Processing Conference (2024): 283–297.
Koto, F., Li, H., Shatnawi, S., Doughman, J., Sadallah, A.
B., Alraeesi, A., Almubarak, K., Alyafeai, Z., Sengupta, N., Shehata, S.,
Habash, N., Nakov, P., and T. Baldwin. "ArabicMMLU: Assessing Massive
Multitask Language Understanding in Arabic." arXiv preprint
arXiv:2402.12840 (2024).
Lamsiyah, S., Zeinalipour, K., El Amrany, S., Brust, M.,
Maggini, M., Bouvry, P., and Schommer, C. "ArabicSense: A Benchmark for
Evaluating Commonsense Reasoning in Arabic with Large Language Models." Proceedings
of the 4th Workshop on Arabic Corpus Linguistics (WACL-4) (2025):
1–11.
Mohseni, Fatima. “Integrating Artificial Intelligence into
Teaching Arabic as a Foreign Language: Effective Strategies for Developing
Linguistic Skills.” (in Arabic). Al?B??ith Journal, (2025): 1–13.
Mulyanto, D., Zaki, M., Ridho, A., and Fata, K.
"Artificial Intelligence Utilization for Arabic Language Skills
Development in Arabic Language Learning." An-Nidzam: Jurnal
Manajemen Pendidikan dan Studi Islam 11, no. 1 (2024): 18–32.
Nehar, A., Bellaouar, S., Souffi, S., and Bouameur, M.
"Dhati: A Fine-Tuned Large Language Model for Evaluating Subjectivity in
Arabic Textual Data." Proceedings of the 2023 5th International
Conference on Pattern Analysis and Intelligent Systems (PAIS)
(2023): p. 1-7
Panagiotidis, P. "LLM-Based Chatbots in Language
Learning." European Journal of Education
7, no. 1 (2024): 102–123.
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S.,
and A. Chadha. "A Systematic Survey of Prompt Engineering in Large
Language Models: Techniques and Applications." arXiv preprint
(2024).
Van Huyssteen, G. B., E. R. Eiselen, and M. J. Puttkammer.
"Re-Evaluating Evaluation Metrics for Spelling Checker Evaluations."
In Proceedings of First Workshop on International Proofing Tools
and Language Technologies, (2004):
91–99
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C.,
Gilbert, H., Elnashar, A., Spencer-Smith, J., and D. C. Schmidt. "A Prompt
Pattern Catalog to Enhance Prompt Engineering with ChatGPT." arXiv
preprint (2023).
Yavuz, F., Celik, ?., and G. Yava? Celik. "Utilizing
Large Language Models for EFL Essay Grading: An Examination of Reliability and
Validity in Rubric-Based Assessments." British Journal of
Educational Technology 56 (2025): 150–166.
Zhang, Z., and X. Huang. "The Impact of Chatbots
Based on Large Language Models on Second Language Vocabulary Acquisition."
Heliyon 10 (2024): e25370.
Zheng, R. "AWE and LLMs in L2 Writing Feedback: An
Exploration of EFL Learners’ Experiences." Academic Journal of
Humanities & Social Sciences 8, no. 7 (2025): 60–69.
Zheng, Y., Li, T., Huang, H., Zeng, T., Lu, J., Chu, C.,
Huang, Y., Jiang, Z., Xiong, Q., Ge, Y., and M. Li. "Are All Prompt
Components Value-Neutral? Understanding the Heterogeneous Adversarial
Robustness of Dissected Prompt in Large Language Models." arXiv
preprint (2025).
Zuccon, G., and B. Koopman. "Dr ChatGPT, Tell Me What
I Want to Hear: How Prompt Knowledge Impacts Health Answer Correctness." arXiv
preprint (2023).
ثالثا-
الموقع الالكتروني:
Google DeepMind. "Gemini Model Thinking Updates: Advanced Coding." Google Blog, March 2025. Retrieved October 19, 2025, from https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/#advanced-coding.
Bibliography
Abdelloui, Muhammad. “Employing Artificial Intelligence in Teaching Arabic as a Foreign Language Remotely.” (in Arabic). Journal of Scientific Development for Studies and Research, iss. 17 (2024): 402–417.
Afyouni, Bashar Mustafa, & Shalkinskaya, Arina. “Linguistic Error Analysis in the Written Production of Russian Learners of Arabic at Moscow State University.” (in Arabic). Journal of Linguistic and Literary Studies, vol. 15, iss. 2 (2024): 67–90.
al-Hadhl?l, ?Al? ibn Hadhl?l ?Al?. “The Impact of Artificial Intelligence Applications on Developing Academic Writing Skills among Arabic Learners Who Speak Other Languages.” (in Arabic). King Abdulaziz University Journal – Educational and Psychological Sciences, vol. 4, iss. 2 (2025): 50–69.
al-Mul?im, Turk? ?Abd-al-?Az?z. “The Reality of Using Smartphone Applications in Teaching Arabic to Speakers of Other Languages at the Institute for Teaching Arabic to Non?Native Speakers at the Islamic University, from the Teachers’ Perspective.” (in Arabic). Assiut University, Faculty of Education Journal, vol. 37, iss. 2 (2021): 39–108.
Amin, Marwa Mustafa Al?Sayed. “Writing Errors among Native and Non?Native Learners of Arabic in the First University Stage: The Faculty of Al?Alsun, Ain Shams University as a case study.” (in Arabic). Philology: Series in Literary and Linguistic Studies, iss. 65 (2016): 35–69.
?Antar, D?n?. “Integrating Artificial Intelligence into Teaching Arabic in Higher Education: A Case Study on ChatGPT?4 Education.” Al?Araek Journal for Sciences and Humanities, iss. 6 (2024): 467–491.
Ban? ?Ur?bah, Ikhl?? bint Ibr?h?m ibn ?amad and al-K?f, F??imah bint Mu?ammad ibn A?mad. “The Effectiveness of Selected Artificial Intelligence Applications in Developing Creative Reading Skills and Retention of Learning Among Sixth?Grade Female Students.” (in Arabic). Arab Journal for Scientific Publishing, isss. 78 (2025): 112–131.
Boubandirah, Abdelaziz. “Traditional Examinations and Assessment Problems.” (in Arabic). Journal of Educational and Pedagogical Research, vol. 13 (Special Issue), (2024): 441–458.
Tu‘aimah, Rushdi Ahmad. "Teaching Arabic to Non?Native Speakers: Its Curricula and Methods". (in Arabic). Islamic Educational, Scientific and Cultural Organization (ISESCO), 1989.