<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>انجمن حسابرسی فناوری اطلاعات ایران</PublisherName>
				<JournalTitle>حسابرسی سیستم‌ها و فناوری اطلاعات</JournalTitle>
				<Issn>3115-8773</Issn>
				<Volume>1</Volume>
				<Issue>2</Issue>
				<PubDate PubStatus="epublish">
					<Year>2025</Year>
					<Month>09</Month>
					<Day>23</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Evaluating the Performance of Large Language Models on Doctoral Accounting Exams: A Comparative Study of Six Generative AI Chatbots</ArticleTitle>
<VernacularTitle>ارزیابی عملکرد مدل‌های زبانی بزرگ در آزمون دکتری حسابداری: مطالعه‌ای مقایسه‌ای از شش چت‌بات هوش مصنوعی مولد</VernacularTitle>
			<FirstPage>57</FirstPage>
			<LastPage>91</LastPage>
			<ELocationID EIdType="pii">240947</ELocationID>
			
<ELocationID EIdType="doi">10.22034/jista.2026.568921.1077</ELocationID>
			
			<Language>FA</Language>
<AuthorList>
<Author>
					<FirstName>ساسان</FirstName>
					<LastName>خادمی</LastName>
<Affiliation>دانش آموخته دکتری حسابداری، بخش حسابداری، دانشکده اقتصاد مدیریت و علوم اجتماعی، دانشگاه شیراز، شیراز، ایران</Affiliation>
<Identifier Source="ORCID">0000-0003-0599-6719</Identifier>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2025</Year>
					<Month>11</Month>
					<Day>26</Day>
				</PubDate>
			</History>
		<Abstract>&lt;span style=&quot;font-family: &#039;Times New Roman&#039;,serif; mso-ascii-theme-font: major-bidi; mso-fareast-font-family: Calibri; mso-hansi-theme-font: major-bidi; mso-bidi-theme-font: major-bidi; color: black;&quot;&gt;The rapid advancement of large language models (LLMs) has drawn increasing attention from accounting education researchers to their performance on specialized questions and potential implications for learning and assessment. This study aims to evaluate and compare the performance of six LLMs (ChatGPT, Gemini, Perplexity, Grok, DeepSeek, and Qwen) on the Iranian PhD Accounting Examination and to assess their potential as educational support tools. The dataset comprises 300 official multiple-choice questions from three subjects (Auditing, Management Accounting, and Accounting Theory) administered between 2021 and 2025. Responses generated by each model were coded dichotomously (correct/incorrect) and evaluated against two reference levels, 0.25 (random performance) and 0.50 (minimum acceptable threshold), using one-sample proportion tests, with 95% confidence intervals reported for model accuracies. Cochran’s Q test was employed to compare relative performance across models. Results indicated that all models performed significantly above both reference levels. Although Gemini achieved the highest and Qwen the lowest correct-response rates, Cochran’s Q revealed no statistically significant differences in overall performance. Importantly, results are interpreted within an open-book scenario, and given the potential for data leakage and the multiple-choice nature of the questions, findings should not be construed as evidence of deep conceptual understanding or independent reasoning. Overall, the findings suggest that LLMs, even without advanced tuning or specialized training, possess substantial capacity for producing correct responses in standard accounting examinations and may serve as complementary tools in accounting education and assessment design.&lt;/span&gt;</Abstract>
			<OtherAbstract Language="FA">&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;پیشرفت شتابان مدل‌های زبانی بزرگ، توجه پژوهشگران را به عملکرد این ابزارها در پاسخ‌گویی به پرسش‌های تخصصی و پیامدهای بالقوه آن‌ها برای یادگیری و ارزشیابی معطوف کرده است. هدف پژوهش حاضر، ارزیابی و مقایسه عملکرد شش مدل زبانی بزرگ شامل &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;ChatGPT&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;، &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Gemini&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;، &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Perplexity&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;، &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Grok&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;، &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;DeepSeek&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; و &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Qwen&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; در پاسخ‌گویی به سؤالات آزمون دکتری حسابداری ایران است. داده‌های پژوهش شامل &lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۳۰۰&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; سؤال چهارگزینه‌ای رسمی آزمون دکتری حسابداری طی سال‌های &lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۱۴۰۰&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; تا &lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۱۴۰۴&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; در سه درس حسابرسی، حسابداری مدیریت و تئوری حسابداری است. پاسخ‌های هر مدل به‌صورت دودویی (صحیح/غلط) کدگذاری شد و با استفاده از آزمون نسبت تک‌نمونه‌ای، عملکرد آن‌ها نسبت به دو سطح مرجع &lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۰&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;Arial&#039;,sans-serif; mso-fareast-font-family: Calibri;&quot;&gt;٫&lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۲۵ (&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;عملکرد تصادفی) و &lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۰&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;Arial&#039;,sans-serif; mso-fareast-font-family: Calibri;&quot;&gt;٫&lt;/span&gt;&lt;span lang=&quot;FA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;; mso-bidi-language: FA;&quot;&gt;۵۰ (&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt;سطح پایه قابل قبول) ارزیابی گردید. همچنین، برای مقایسه عملکرد نسبی مدل‌ها از آزمون &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Cochran’s Q&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; استفاده شد. نتایج نشان داد که عملکرد تمامی مدل‌ها به‌طور معناداری فراتر از هر دو سطح مرجع است. اگرچه مدل &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Gemini&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; بالاترین و مدل &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Qwen&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; پایین‌ترین درصد پاسخ صحیح را ثبت کردند، آزمون &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;Cochran’s Q&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; تفاوت معناداری میان عملکرد کلی مدل‌ها نشان نداد. با این حال، نتایج در چارچوب یک سناریوی عملیاتی &lt;/span&gt;&lt;span dir=&quot;LTR&quot; style=&quot;font-size: 10.0pt; mso-bidi-font-size: 11.0pt; font-family: &#039;Times New Roman&#039;,serif; mso-fareast-font-family: Calibri; mso-bidi-font-family: &#039;B Zar&#039;;&quot;&gt;open-book&lt;/span&gt;&lt;span lang=&quot;AR-SA&quot; style=&quot;mso-ansi-font-size: 10.0pt; font-family: &#039;B Zar&#039;; mso-ascii-font-family: &#039;Times New Roman&#039;; mso-fareast-font-family: Calibri; mso-hansi-font-family: &#039;Times New Roman&#039;;&quot;&gt; تفسیر می‌شوند و با توجه به احتمال نشت داده و ماهیت چندگزینه‌ای سؤالات، نباید به‌عنوان شواهدی از درک مفهومی عمیق یا استدلال مستقل مدل‌ها تلقی شوند. به‌طور کلی، یافته‌ها نشان می‌دهد که مدل‌های زبانی بزرگ، حتی بدون تنظیمات پیشرفته یا آموزش اختصاصی، از توان قابل توجهی در عملکرد صحیح در آزمون‌های استاندارد حسابداری برخوردارند و می‌توانند به‌عنوان ابزارهای مکمل در آموزش و طراحی فعالیت‌های ارزشیابی در آموزش عالی حسابداری مورد توجه قرار گیرند.&lt;/span&gt;</OtherAbstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">مدل‌های زبانی بزرگ</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">آموزش حسابداری</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">آزمون دکتری حسابداری</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">ارزیابی عملکرد</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">هوش مصنوعی در آموزش</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://www.iitasa.org.ir/article_240947_fd7c140b56af05f71644a608bb4f05f8.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
