<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//TaxonX//DTD Taxonomic Treatment Publishing DTD v0 20100105//EN" "../../nlm/tax-treatment-NS0.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:tp="http://www.plazi.org/taxpub" article-type="research-article" dtd-version="3.0" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">69</journal-id>
      <journal-id journal-id-type="index">urn:lsid:arphahub.com:pub:8D21F818-6EEF-540F-91C7-D50E3E5A13E0</journal-id>
      <journal-title-group>
        <journal-title xml:lang="en">Maandblad voor Accountancy en Bedrijfseconomie</journal-title>
        <abbrev-journal-title xml:lang="en">MAB</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="ppub">0924-6304</issn>
      <issn pub-type="epub">2543-1684</issn>
      <publisher>
        <publisher-name>Amsterdam University Press</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5117/mab.99.132881</article-id>
      <article-id pub-id-type="publisher-id">132881</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group subj-group-type="scientific_subject">
          <subject>Accountantscontrole (Auditing)</subject>
          <subject>Externe verslaggeving (External reporting)</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>﻿Artificial Intelligence in fraud detection: textual analysis of 10-K filings</article-title>
      </title-group>
      <contrib-group content-type="authors">
        <contrib contrib-type="author" corresp="yes">
          <name name-style="western">
            <surname>Ketelaar</surname>
            <given-names>Florian</given-names>
          </name>
          <email xlink:type="simple">florian@ketelaar.tv</email>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
        <contrib contrib-type="author" corresp="no">
          <name name-style="western">
            <surname>Mićković</surname>
            <given-names>Ana</given-names>
          </name>
          <xref ref-type="aff" rid="A1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="A1">
        <label>1</label>
        <addr-line content-type="verbatim">University of Amsterdam, Amsterdam, Netherlands</addr-line>
        <institution>University of Amsterdam</institution>
        <addr-line content-type="city">Amsterdam</addr-line>
        <country>Netherlands</country>
      </aff>
      <author-notes>
        <fn fn-type="corresp">
          <p>Corresponding author: Florian Ketelaar (<email xlink:type="simple">florian@ketelaar.tv</email>).</p>
        </fn>
        <fn fn-type="edited-by">
          <p>Academic editor: Oscar van Leeuwen</p>
        </fn>
      </author-notes>
      <pub-date pub-type="collection">
        <year>2025</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>25</day>
        <month>04</month>
        <year>2025</year>
      </pub-date>
      <volume>99</volume>
      <issue>2</issue>
      <fpage>61</fpage>
      <lpage>71</lpage>
      <uri content-type="arpha" xlink:href="http://openbiodiv.net/94241BFB-E105-5F20-A62D-7E4972FA4F83">94241BFB-E105-5F20-A62D-7E4972FA4F83</uri>
      <history>
        <date date-type="received">
          <day>23</day>
          <month>07</month>
          <year>2024</year>
        </date>
        <date date-type="accepted">
          <day>22</day>
          <month>01</month>
          <year>2025</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>Florian Ketelaar, Ana Mićković</copyright-statement>
        <license license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/" xlink:type="simple">
          <license-p>This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY-NC-ND 4.0), which permits to copy and distribute the article for non-commercial purposes, provided that the article is not altered or modified and the original author and source are credited.</license-p>
        </license>
      </permissions>
      <abstract>
        <label>﻿Abstract</label>
        <p>In this paper, we investigate the potential of Artificial Intelligence (<abbrev xlink:title="Artificial Intelligence" id="ABBRID0ENC">AI</abbrev>) in detecting fraud by analyzing linguistic indicators in 10-K filings. We analyze the word frequencies (positive, negative, uncertainty, litigious), consistency, and readability in the MD&amp;A sections. The <abbrev xlink:title="Artificial Intelligence" id="ABBRID0ERC">AI</abbrev> model, <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EVC">BERT</abbrev>, was trained on these factors to predict fraud, showing significant promise compared to traditional models. The findings suggest that fraudulent filings tend to have more positive words, inconsistent language, and higher readability. This highlights <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EZC">AI</abbrev>’s practical role in improving fraud detection in financial reports.</p>
      </abstract>
      <kwd-group>
        <label>Keywords</label>
        <kwd>Fraud detection</kwd>
        <kwd>Artificial Intelligence</kwd>
        <kwd>financial reports</kwd>
        <kwd>textual analysis</kwd>
        <kwd>BERT model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec sec-type="﻿Relevance to practice" id="SECID0EBD">
      <title>﻿Relevance to practice</title>
      <p>This research demonstrates the practical application of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EHD">AI</abbrev>, specifically <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0ELD">BERT</abbrev>, in enhancing fraud detection in financial reports. By identifying key linguistic indicators of fraud, it provides a tool for auditors and regulators to improve accuracy and efficiency in monitoring and investigating potential financial misconduct.</p>
    </sec>
    <sec sec-type="﻿1. Introduction" id="SECID0EPD">
      <title>﻿1. Introduction</title>
      <p>Technological advancements have continually reshaped industries. One of the most significant transformations currently underway is the rise of Artificial Intelligence (<abbrev xlink:title="Artificial Intelligence" id="ABBRID0EVD">AI</abbrev>), which is becoming increasingly important across various industries and research fields. <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EZD">AI</abbrev> is already impacting many data-related jobs, and therefore it is no surprise that the audit industry is expected to be affected as well. Studies, such as <xref ref-type="bibr" rid="B19">Hasan (2021)</xref>, suggest that <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EBE">AI</abbrev> will play a significant role in transforming audit processes.</p>
      <p>Consulting companies have started adopting <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EHE">AI</abbrev> tools to improve their audit processes, particularly in areas like fraud detection, which is a major risk for firms due to its potential reputational and legal consequences. As fraud evolves, auditors must innovate. <abbrev xlink:title="Artificial Intelligence" id="ABBRID0ELE">AI</abbrev> tools like HeadStart or Argus are designed to help auditors navigate through complex regulations and analyze entire datasets, rather than just samples, to identify risks, anomalies, and trends (<xref ref-type="bibr" rid="B11">Davenport 2016</xref>).</p>
      <p>We aim to answer the following research question: <italic>What factors does <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EXE">AI</abbrev> detect as potentially fraudulent from specific linguistic patterns within 10-K filings</italic>?</p>
      <p>To be able to correctly and efficiently detect financial fraud, particularly in corporate financial statements like 10-K SEC (Securities and Exchange Commission) filings, is very important and cannot be ignored. Traditional methods of financial fraud detection, that are mostly based on human analysis and standard statistical techniques, have shown to be insufficient at identifying fraud. In contrast to traditional fraud detection methods, <abbrev xlink:title="Artificial Intelligence" id="ABBRID0E5E">AI</abbrev> has advanced computational and learning capabilities making it a good tool for an auditor. The study of <xref ref-type="bibr" rid="B25">Kureljusic and Karger (2023)</xref> finds that <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EGF">AI</abbrev> is highly relevant in financial accounting and their results show that <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EKF">AI</abbrev> has several practical uses and benefits for practitioners.</p>
      <p>Current literature has focused on the potential of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EQF">AI</abbrev> in multiple domains of financial analysis, but its use in fraud detection within 10-K filings remains under-explored (<xref ref-type="bibr" rid="B10">Craja et al. 2020</xref>). The research question shows relevance because it tests the potential of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EYF">AI</abbrev> to improve the traditional methods or even become better than these. From a social perspective, to be able to effectively detect fraudulent financial statements and to have automated mechanisms for it is key to maintain confidence in financial reports. To better answer the research question, we use the Management Discussion and Analysis (<abbrev xlink:title="Management Discussion and Analysis" id="ABBRID0E3F">MD&amp;A</abbrev>) section of the 10-K SEC filings. These MD&amp;A sections are applicable as <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EAG">AI</abbrev> is capable of analyzing text and finding structures in it that might indicate fraudulent behaviour.</p>
    </sec>
    <sec sec-type="﻿2. Literature review" id="SECID0EEG">
      <title>﻿2. Literature review</title>
      <sec sec-type="﻿2.1. Fraud" id="SECID0EIG">
        <title>﻿2.1. Fraud</title>
        <p>Financial fraud, particularly accounting fraud, is a major concern in manipulating financial statements (<xref ref-type="bibr" rid="B8">Campa et al. 2023</xref>). <xref ref-type="bibr" rid="B41">Wang et al. (2006)</xref> define fraud as “a deliberate act contrary to law, rule, or policy with intent to obtain unauthorized financial benefit.” <xref ref-type="bibr" rid="B22">Joyce and Biddle (1981)</xref> noted that fraud committed by higher-level management is especially hard to detect. <xref ref-type="bibr" rid="B24">Kieso et al. (2020)</xref> emphasized that financial fraud involves intentional misstatements or omissions of material information. <xref ref-type="bibr" rid="B1">Beasley et al. (2010)</xref> identified improper revenue recognition and asset overstatement as the most common methods of fraud. The Fraud Triangle, outlined by SAS No. 99, describes the three conditions under which fraud occurs: pressure, opportunity, and rationalization.</p>
      </sec>
      <sec sec-type="﻿2.2. Readability of text and sentiment" id="SECID0ECH">
        <title>﻿2.2. Readability of text and sentiment</title>
        <p>The readability of financial reports, particularly in the MD&amp;A sections, can provide insights into fraudulent behaviour. The Fog Index, developed by <xref ref-type="bibr" rid="B17">Gunning (1952)</xref>, measures readability by scoring the complexity of a text, where a higher Fog Index indicates a more difficult to read text. <xref ref-type="bibr" rid="B28">Li (2008)</xref> found that firms with lower earnings tend to have higher Fog Index scores, suggesting a link between complex, harder to read reports and poorer financial performance. <xref ref-type="bibr" rid="B32">Martinc et al. (2021)</xref> demonstrated that <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EUH">AI</abbrev> can effectively assess the readability of documents, making it a useful tool for identifying anomalies in financial reports.</p>
        <p>In addition to readability, the sentiment expressed through positive and negative words in a text can signal underlying financial conditions. Positive words convey positive sentiment, while negative words reflect a negative tone. However, as <xref ref-type="bibr" rid="B16">Ghosh et al. (2015)</xref> noted, context can alter meaning, for example, sarcastic use of positive words can imply negativity, and phrases like “not bad” can turn negative meanings positive. <xref ref-type="bibr" rid="B29">Loughran and McDonald (2011)</xref> developed word lists to capture sentiment in financial disclosures, categorizing words into positive, negative, litigious, and uncertain. Examples include “gains” and “profitability” for positive words, “loss” and “impairment” for negative, “contracts” and “regulatory” for litigious, and “may” and “risk” for uncertain. These sentiment patterns are critical in analyzing the MD&amp;A sections of 10-K filings.</p>
      </sec>
      <sec sec-type="﻿2.3. Artificial Intelligence and its role in accounting" id="SECID0EDAAC">
        <title>﻿2.3. Artificial Intelligence and its role in accounting</title>
        <p>Artificial Intelligence (<abbrev xlink:title="Artificial Intelligence" id="ABBRID0EJAAC">AI</abbrev>) encompasses advancements like Machine Learning (<abbrev xlink:title="Machine Learning" id="ABBRID0ENAAC">ML</abbrev>) and Natural Language Processing (<abbrev xlink:title="Natural Language Processing" id="ABBRID0ERAAC">NLP</abbrev>), along with traditional statistical methods such as classification and clustering (<xref ref-type="bibr" rid="B39">Sutton et al. 2016</xref>). <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EZAAC">AI</abbrev> models can be trained using various methods, including supervised, semi-supervised, unsupervised, and reinforcement learning (<xref ref-type="bibr" rid="B7">Burkov 2019</xref>). In supervised learning, the dataset contains labeled examples, where each label represents a specific feature, such as identifying fraud in the MD&amp;A sections of 10-K filings.</p>
        <p>In this study, <abbrev xlink:title="Natural Language Processing" id="ABBRID0EDBAC">NLP</abbrev> is employed to automatically classify companies as fraudulent or non-fraudulent. The model used is based on <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EHBAC">BERT</abbrev> (Bidirectional Encoder Representations from Transformers), introduced by <xref ref-type="bibr" rid="B13">Devlin (2018)</xref>. <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EPBAC">BERT</abbrev> is designed to pre-train deep bidirectional representations from unlabeled text, jointly conditioning on both left and right contexts in all layers. <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0ETBAC">BERT</abbrev> is one of the earliest Large Language Models (<abbrev xlink:title="Large Language Models" id="ABBRID0EXBAC">LLMs</abbrev>) and can be applied to various tasks such as question answering and language inference without significant modification.</p>
        <p><xref ref-type="bibr" rid="B21">Huang et al. (2023)</xref> noted that <abbrev xlink:title="Large Language Models" id="ABBRID0EBCAC">LLMs</abbrev>, due to their vast number of parameters, can learn semantic and syntactic relationships between words. However, these models are often expensive and difficult to train. Moreover, once trained, <abbrev xlink:title="Large Language Models" id="ABBRID0EFCAC">LLMs</abbrev> become “black boxes,” making their decision-making processes challenging to interpret (<xref ref-type="bibr" rid="B20">Hassija et al. 2024</xref>).</p>
        <p><abbrev xlink:title="Artificial Intelligence" id="ABBRID0EPCAC">AI</abbrev> has also been widely implemented in accounting, transforming traditional tasks. Studies such as <xref ref-type="bibr" rid="B42">Zhang et al. (2020)</xref> show that Big 4 firms have integrated <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EXCAC">AI</abbrev> technologies to enhance their services through automated data entry, voice analysis, intelligent search engines, and predictive analytics. These innovations allow accountants to focus on higher-value activities like strategic planning and advisory services. <abbrev xlink:title="Artificial Intelligence" id="ABBRID0E2CAC">AI</abbrev>-based fraud detection methods, such as those used by <xref ref-type="bibr" rid="B6">Brown et al. (2020)</xref> and <xref ref-type="bibr" rid="B23">Ikhsan et al. (2022)</xref>, have demonstrated high accuracy in detecting financial misreporting, thereby improving audit quality.</p>
        <p><xref ref-type="bibr" rid="B37">Ranta et al. (2023)</xref> identified several areas where <abbrev xlink:title="Artificial Intelligence" id="ABBRID0ENDAC">AI</abbrev> has impacted accounting, including changes in the profession, textual analysis of accounting data, and improved prediction methods. <xref ref-type="bibr" rid="B40">Tang et al. (2018)</xref> developed an ontology-based fraud detection model using decision trees, achieving 86.67% accuracy in identifying fraudulent financial activity. <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EVDAC">AI</abbrev>-driven accounting systems, as described by <xref ref-type="bibr" rid="B27">Lee and Tajudeen (2020)</xref>, have greatly improved productivity, efficiency, and risk management in financial reporting.</p>
        <p><xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref> applied <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EDEAC">BERT</abbrev> to detect fraudulent reporting in MD&amp;A sections of 10-K filings between 1994 and 2013. Their models outperformed others by around 15% in detecting fraud, providing significant economic benefits to regulators, investors, and analysts.</p>
      </sec>
    </sec>
    <sec sec-type="﻿3. Hypotheses development" id="SECID0EHEAC">
      <title>﻿3. Hypotheses development</title>
      <p>In the study by <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EREAC">AI</abbrev> algorithms demonstrated significant potential in identifying fraudulent companies from 1994 to 2013, outperforming other models by 15%. With the increasing reliance on <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EVEAC">AI</abbrev> models for processing Big Data and detecting fraud, it is essential to validate their findings and test whether they hold true for more recent company filings. A key goal of our study is to evaluate the results of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref> regarding the ability of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0E4EAC">AI</abbrev> to detect fraudulent activity based on the MD&amp;A sections of 10-K filings. The relevance of this is highlighted by <xref ref-type="bibr" rid="B35">Oyewole et al. (2024)</xref>, who found that companies increasingly use text-generating tools for financial reporting, making it even more crucial to test the applicability of the findings of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref> in recent years.</p>
      <p>
        <italic>Hypothesis 1: A <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0ENFAC">BERT</abbrev> model remains effective in detecting fraudulent activity in more recent years, with accuracy comparable to or exceeding that reported by <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>.</italic>
      </p>
      <p>To detect fraud through word categorization, a model is needed to analyze the language and make well-substantiated conclusions. The agency theory highlights linguistic characteristics that may signal fraudulent behavior. <xref ref-type="bibr" rid="B36">Purda and Skillicorn (2015)</xref> demonstrated that a support vector machine (<abbrev xlink:title="support vector machine" id="ABBRID0E3FAC">SVM</abbrev>) classifier, using the top 200 fraud-linked words, could effectively identify fraudulent reporting in MD&amp;A sections of 10-K filings. <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref> also provided evidence that fraudulent firms tend to use more positive words while minimizing negative language, suggesting that they do so deliberately to conceal fraudulent activities. Their study relied on <xref ref-type="bibr" rid="B29">Loughran and McDonald’s (2011)</xref> word classifications to analyze the frequency of positive and negative words. Additionally, <xref ref-type="bibr" rid="B26">Larcker and Zakolyukina (2012)</xref> found that deceptive CEOs often use more positive emotional language, which could help detect fraud, particularly in higher-level management, as noted by <xref ref-type="bibr" rid="B22">Joyce and Biddle (1981)</xref>.</p>
      <p>Building on the work of <xref ref-type="bibr" rid="B36">Purda and Skillicorn (2015)</xref>, <xref ref-type="bibr" rid="B4">Bochkay et al. (2023)</xref>, <xref ref-type="bibr" rid="B29">Loughran and McDonald (2011)</xref>, <xref ref-type="bibr" rid="B26">Larcker and Zakolyukina (2012)</xref>, <xref ref-type="bibr" rid="B22">Joyce and Biddle (1981)</xref>, and the suggestions of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, we propose the following hypothesis. In this hypothesis, we examine the frequency of positive, negative, litigious, and uncertainty words in fraudulent versus non-fraudulent firms’ MD&amp;A sections.</p>
      <p>
        <italic>Hypothesis 2: Firms engaged in fraudulent activities use a higher/lower frequency of positive, negative, litigious, and uncertainty words in their MD&amp;A sections compared to non-fraudulent firms.</italic>
      </p>
      <p>To further explore the linguistic factors <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EQHAC">AI</abbrev> detects, we analyze the frequency of specific word categories over time. <xref ref-type="bibr" rid="B21">Huang et al. (2023)</xref> suggest that <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EYHAC">AI</abbrev> can interpret both the semantics and syntactics of text, while <xref ref-type="bibr" rid="B26">Larcker and Zakolyukina (2012)</xref> note that deceptive CEOs tend to use more positive emotional language. This indicates that one factor <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EAIAC">AI</abbrev> may detect is the variance in word category frequency. Fraudulent firms may exhibit more fluctuation in their language to manipulate perceptions, while non-fraudulent firms are likely to use more consistent language over time. This variance in word usage could be a key indicator <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EEIAC">AI</abbrev> picks up when detecting fraudulent activities.</p>
      <p>
        <italic>Hypothesis 3: The frequency of word categories in non-fraudulent filings is more consistent than in fraudulent filings.</italic>
      </p>
      <p>In addition to analyzing the frequency and consistency of word categories, it is crucial to consider the readability of the MD&amp;A sections as a potential indicator of fraudulent activity. A common measure of readability is the in section 2.2 mentioned Fog Index (<xref ref-type="bibr" rid="B17">Gunning 1952</xref>). <xref ref-type="bibr" rid="B28">Li (2008)</xref> found that companies with lower earnings tend to have a higher Fog Index, which aligns with agency theory (<xref ref-type="bibr" rid="B14">Eisenhardt 1989</xref>) that suggests firms with lower earnings may have incentives to fraudulently inflate them. <xref ref-type="bibr" rid="B32">Martinc et al. (2021)</xref> demonstrated that <abbrev xlink:title="Artificial Intelligence" id="ABBRID0E5IAC">AI</abbrev> can interpret the readability of documents, while <xref ref-type="bibr" rid="B34">Nakashima et al. (2022)</xref> found significant differences in readability between fraudulent and non-fraudulent financial statements in Japan, with fraudulent filings being more difficult to read.</p>
      <p>
        <italic>Hypothesis 4: Firms engaged in fraudulent activities have a higher Fog Index in their MD&amp;A sections compared to non-fraudulent firms.</italic>
      </p>
    </sec>
    <sec sec-type="methods" id="SECID0EKJAC">
      <title>﻿4. Research method</title>
      <p>We employ <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EQJAC">AI</abbrev> models based on machine learning techniques like Natural Language Processing (<abbrev xlink:title="Natural Language Processing" id="ABBRID0EUJAC">NLP</abbrev>) and anomaly detection to identify patterns indicative of fraud. Specifically, we use <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EYJAC">BERT</abbrev> (Bidirectional Encoder Re­presentations from Transformers), designed by Google (<xref ref-type="bibr" rid="B13">Devlin 2018</xref>) to understand word context by examining surrounding text. To evaluate the performance of our <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EAKAC">AI</abbrev> model, we compare it with models from previous studies. The model is trained using the Management Discussion and Analysis (MD&amp;A) sections from 10-K filings. This research aims to explore <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EEKAC">AI</abbrev>’s potential in financial fraud detection, contributing both theoretical and practical insights into <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EIKAC">AI</abbrev>’s role in financial oversight.</p>
      <p>In our analysis, we use the EDGAR database from the US Securities and Exchange Commission, which contains 10-K annual reports of US firms from 1994 to 2024. These 10-K reports, publicly available, were downloaded via the University of Notre Dame (<xref ref-type="bibr" rid="B30">Loughran and McDonald 2016</xref>). Item 7, the MD&amp;A sections, are extracted from these reports for analysis, as prior research shows they provide valuable data for text analysis, fraud detection, and investor information (<xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>; <xref ref-type="bibr" rid="B15">Goel and Uzuner 2016</xref>; <xref ref-type="bibr" rid="B6">Brown et al. 2020</xref>; <xref ref-type="bibr" rid="B36">Purda and Skillicorn 2015</xref>). The extracted text is parsed for machine learning analysis, following the method of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref> as described below, who highlight the significance of MD&amp;A in detecting fraud through <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EGLAC">AI</abbrev>-based text analysis. According to the SEC, MD&amp;A sections provide critical insights that clarify and supplement financial statements (<xref ref-type="bibr" rid="B9">Cole and Jones 2004</xref>).</p>
      <p>Our dataset is constructed from the EDGAR database, using 10-K SEC filings from 1994 to 2021, pre-parsed by the University of Notre Dame (<xref ref-type="bibr" rid="B30">Loughran and McDonald 2016</xref>). The AAER dataset (<xref ref-type="bibr" rid="B12">Dechow et al. 2011</xref>), which tracks material fraud occurrences as reported by the SEC, is updated through 2021, serving as our cutoff period. This differs from <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, who used an earlier version of the AAER with a 2013 cutoff. The AAER database contains detailed information about fraud cases, including managers involved and the periods during which the fraud took place, though there is typically a three-year time lag due to processing. Studies, such as <xref ref-type="bibr" rid="B2">Bertomeu et al. (2021)</xref>, show that models trained with the AAER dataset are more effective at detecting misstatements.</p>
      <p>For our supervised learning model, we label examples as fraudulent or not by matching the Central Index Key (<abbrev xlink:title="Central Index Key" id="ABBRID0ECMAC">CIK</abbrev>), a unique identifier for U.S. firms, between the AAER database and the MD&amp;A sections of 10-K filings. We extract the MD&amp;A text from Item 7 of the 10-K annual reports. A search algorithm identifies the <abbrev xlink:title="Central Index Key" id="ABBRID0EGMAC">CIK</abbrev> so the MD&amp;A text can be linked to the correct company. If the <abbrev xlink:title="Central Index Key" id="ABBRID0EKMAC">CIK</abbrev> is found in the AAER database, we label the firm as fraudulent (Fraud = 1); otherwise, it is labeled as non-fraudulent (Fraud = 0). We label all filings by a firm identified as fraudulent in a given year as fraudulent, under the assumption that fraudulent activity may span multiple years and may not be isolated to the filing flagged by regulators. While this may overestimate the number of fraudulent filings, it aligns with prior work suggesting that fraud is often persistent and difficult to detect in a single filing.</p>
      <p>The dataset includes a total of 123,415 filings from 1994 to 2021, with 5,785 identified as fraudulent, representing 4.66% of the total. Figure <xref ref-type="fig" rid="F1">1</xref> visually represents this data, showing the non-fraudulent filings per year on the left vertical axis and fraudulent filings on the right vertical axis. For brevity, we have omitted the detailed table of numbers from the manuscript; however, it is available upon request.</p>
      <fig id="F1" position="float" orientation="portrait">
        <object-id content-type="arpha">15B93C1B-D977-5785-AD2A-866DE5AAD691</object-id>
        <label>Figure 1.</label>
        <caption>
          <p>Fraudulent vs. non-fraudulent 10-K filings per year. Note: The blue line represents the number of non-fraudulent filings per year, shown on the left vertical axis, while the red line represents the number of fraudulent filings per year, shown on the right vertical axis.</p>
        </caption>
        <graphic xlink:href="mab-99-061-g001.jpg" position="float" orientation="portrait" xlink:type="simple" id="oo_1315279.jpg">
          <uri content-type="original_file">https://binary.pensoft.net/fig/1315279</uri>
        </graphic>
      </fig>
      <sec sec-type="﻿4.1. Descriptive statistics" id="SECID0EANAC">
        <title>﻿4.1. Descriptive statistics</title>
        <p>To understand the sample, we calculated the total number of words, sentences, and words per sentence in the MD&amp;A sections. We found that the number of words per year in MD&amp;A sections increased significantly over time. Before the year 2000, MD&amp;A sections contained around 2,000 to 3,000 words per year, but after 2015, this range grew to 12,000 to 14,000 words; an increase of nearly 10,000 words over the last 15 years. This aligns with the findings of <xref ref-type="bibr" rid="B5">Brown and Tucker (2011)</xref>, who suggest that the increase in word count may result from managers using boilerplate disclosures, which are standardized and provide little firm-specific information, reducing the overall usefulness of the MD&amp;A.</p>
        <p>On average, the MD&amp;A sections in our dataset contain 8,854 words, 498 sentences, and 17 words per sentence. We observed that both the number of sentences and the average sentence length increased over time, with the rise in total sentences likely explained by the increase in total words. However, the increase in sentence length was an unexpected finding, potentially due to the inclusion of boilerplate information, as noted by <xref ref-type="bibr" rid="B5">Brown and Tucker (2011)</xref>.</p>
        <p>Further, Figure <xref ref-type="fig" rid="F2">2</xref> provides a visual comparison of the word count between fraudulent and non-fraudulent MD&amp;A sections. Interestingly, fraudulent filings have an average of 12,482 words, exceeding the 8,926 word average for non-fraudulent filings. This contradicts <xref ref-type="bibr" rid="B34">Nakashima et al. (2022)</xref>, who found that MD&amp;A disclosures were insignificantly shorter for Japanese fraudulent firms. One possible explanation is that fraudsters may attempt to bury fraudulent activities in longer texts, as readers are less likely to scrutinize lengthy reports. This is supported by <xref ref-type="bibr" rid="B18">Hancock (2007)</xref>, who suggests that liars tend to use more words, which could explain our findings.</p>
        <fig id="F2" position="float" orientation="portrait">
          <object-id content-type="arpha">6A5E1966-D8C2-569A-90C3-83C1F6860F25</object-id>
          <label>Figure 2.</label>
          <caption>
            <p>Comparison of word count in fraudulent vs. non-fraudulent MD&amp;A sections. Note: The figure compares the average word count in MD&amp;A sections of fraudulent (red line) and non-fraudulent (blue line) 10-K filings, with the overall average word count (green line) included for reference.</p>
          </caption>
          <graphic xlink:href="mab-99-061-g002.jpg" position="float" orientation="portrait" xlink:type="simple" id="oo_1315280.jpg">
            <uri content-type="original_file">https://binary.pensoft.net/fig/1315280</uri>
          </graphic>
        </fig>
      </sec>
      <sec sec-type="﻿4.2. Data analysis" id="SECID0EKOAC">
        <title>﻿4.2. Data analysis</title>
        <p>To validate the findings of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, we use the <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EUOAC">BERT</abbrev>-Base-uncased model from Hugging Face, a pre-trained language model. Similar to <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, we apply the WordPiece tokenizer, which splits text into tokens by removing punctuation and identifying word roots. Following <xref ref-type="bibr" rid="B13">Devlin (2018)</xref> and <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, we use a classification token (<abbrev xlink:title="classification token" id="ABBRID0EEPAC">CLS</abbrev>) at the start and a separation token (<abbrev xlink:title="separation token" id="ABBRID0EIPAC">SEP</abbrev>) at the end of each input, ensuring sequences do not exceed the 512-token limit. We fine-tune the <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EMPAC">BERT</abbrev> models using Google Colab Pro’s Nvidia Tesla A100 GPU. The fine-tuning process is done with a learning rate of 2e-5, the AdamW optimizer, over 3 epochs, and with a batch size of 16.</p>
        <p><xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref> approach fraud detection as a ranking task, using rank averaging to ensure that fraudulent samples are ranked higher than non-fraudulent ones, which reduces variability in predictions. They combine predictions from two models, <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EWPAC">BERT</abbrev><sub>first</sub> and <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0E2PAC">BERT</abbrev><sub>last</sub>, trained on the first and last 512 tokens of the MD&amp;A section, respectively, to capture both introductory summaries and future outlooks. To replicate their approach, we construct two models (<abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EBAAE">BERT</abbrev><sub>first</sub> and <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EGAAE">BERT</abbrev><sub>last)</sub>, train them on the first and last 512 tokens of the MD&amp;A sections, and then average their predictions. Our final prediction is the rank average of the outputs from <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0ELAAE">BERT</abbrev><sub>first</sub> and <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EQAAE">BERT</abbrev><sub>last,</sub> where pred<sub>final</sub> = 1/2 * rank(pred<sub>first</sub>) + 1/2 * rank(pred<sub>last</sub>).</p>
        <p>To validate the model and test Hypothesis 1 while staying consistent with <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, we use a rolling window of consecutive five years for training, followed by the subsequent year as the test set. The validation set includes the years 1994 to 1999, in line with their study, to optimize our model parameters. For example, we train the model on data from 2000 to 2004 and test it on 2005. This approach is applied from 1994 to 2021, giving us a 23-year test period, extending beyond their 2013 cutoff. For evaluation, we also use the area under the ROC curve (<abbrev xlink:title="area under the ROC curve" id="ABBRID0EBBAE">AUC</abbrev>) as our primary metric. Given the class imbalance common in fraud detection, <abbrev xlink:title="area under the ROC curve" id="ABBRID0EFBAE">AUC</abbrev> is suitable as it measures the likelihood that a randomly selected fraud sample is ranked higher than a non-fraud sample, ensuring comparability with <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>.</p>
        <p>To test Hypotheses 2 and 3, we use the word list from <xref ref-type="bibr" rid="B31">Loughran and McDonald (2023)</xref>. We first split the dataset into fraud and non-fraud instances and then filtered the text based on the words from the Loughran and McDonald list. This list categorizes words into four sections: Positive, Negative, Litigious, and Uncertain. The dataset is organized by year, and we extracted the corresponding MD&amp;A sections for each year. We then counted the occurrences of each word per year and categorized them according to Loughran and McDonald’s classification. Additionally, we calculated the total number of words per year to determine the percentage occurrence of each category annually.</p>
        <p>To test Hypothesis 4, we use the Fog Index developed by <xref ref-type="bibr" rid="B17">Gunning (1952)</xref>. We split the dataset into fraudulent and non-fraudulent filings per year and calculated the average Fog Index for each year, following Gunning’s method. We use an Ordinary Least Squares (<abbrev xlink:title="Ordinary Least Squares" id="ABBRID0EZBAE">OLS</abbrev>) regression analysis to test whether there are significant differences in word categories and readability between fraudulent and non-fraudulent MD&amp;A sections. We also examine if these differences have changed over time, particularly with the rise of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0E4BAE">AI</abbrev>-generated financial reports.</p>
        <p>Separate <abbrev xlink:title="Ordinary Least Squares" id="ABBRID0EDCAE">OLS</abbrev> regressions are conducted for the Fog Index and the occurrence of word categories (positive, negative, uncertainty, and litigious) to test for significant differences. We compare trends between two periods, 1994–2007 and 2008–2021, to see if word category usage has shifted. The regression includes variables for baseline measures, whether the MD&amp;A section is non-fraudulent, year, and an interaction term to assess how trends differ between fraudulent and non-fraudulent sections over time.</p>
      </sec>
    </sec>
    <sec sec-type="﻿5. Results" id="SECID0EHCAE">
      <title>﻿5. Results</title>
      <sec sec-type="﻿5.1. Hypothesis 1" id="SECID0ELCAE">
        <title>﻿5.1. Hypothesis 1</title>
        <p>We formulated Hypothesis 1 to assess whether a <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0ERCAE">BERT</abbrev> model remains effective in detecting fraudulent activity in recent years, with accuracy comparable to or exceeding that of BM (2024). To gauge performance, we use their results as a benchmark, directly comparing our final model to their <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EVCAE">BERT</abbrev> models.</p>
        <p>Table <xref ref-type="table" rid="T1">1</xref> displays the yearly and average <abbrev xlink:title="area under the ROC curve" id="ABBRID0E6CAE">AUC</abbrev> scores for our models from 2014 to 2021. Our <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EDDAE">BERT</abbrev> models, trained on MD&amp;A sections from 10-K filings, are evaluated against the benchmark models. Our <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EHDAE">BERT</abbrev><sub>first</sub>, <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EMDAE">BERT</abbrev><sub>last</sub>, and <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0ERDAE">BERT</abbrev><sub>final</sub> models achieved average <abbrev xlink:title="area under the ROC curve" id="ABBRID0EWDAE">AUC</abbrev> scores of 0.844, 0.661, and 0.797, respectively, compared to <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, which reported <abbrev xlink:title="area under the ROC curve" id="ABBRID0E5DAE">AUC</abbrev> scores of 0.804, 0.816, and 0.826. While the <abbrev xlink:title="area under the ROC curve" id="ABBRID0ECEAE">AUC</abbrev> results for 2000–2013 are omitted for brevity, they are available upon request and show comparable performance between our models and <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>.</p>
        <table-wrap id="T1" position="float" orientation="portrait">
          <label>Table 1.</label>
          <caption>
            <p><abbrev xlink:title="area under the ROC curve" id="ABBRID0ETEAE">AUC</abbrev> score per <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EXEAE">BERT</abbrev> model per year.</p>
          </caption>
          <table id="TID0EBJAG" rules="all">
            <tbody>
              <tr>
                <th rowspan="1" colspan="1">
                  <abbrev xlink:title="area under the ROC curve" id="ABBRID0EEFAE">AUC</abbrev>
                </th>
                <th rowspan="1" colspan="1">2014</th>
                <th rowspan="1" colspan="1">2015</th>
                <th rowspan="1" colspan="1">2016</th>
                <th rowspan="1" colspan="1">2017</th>
                <th rowspan="1" colspan="1">2018</th>
                <th rowspan="1" colspan="1">2019</th>
                <th rowspan="1" colspan="1">2020</th>
                <th rowspan="1" colspan="1">2021</th>
                <th rowspan="1" colspan="1">Average</th>
              </tr>
              <tr>
                <td rowspan="1" colspan="1">
                  <italic>BERTfirst</italic>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,738</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,892</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,846</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,878</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,824</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,878</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,823</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,869</bold>
                </td>
                <td rowspan="1" colspan="1">
                  <bold>0,844</bold>
                </td>
              </tr>
              <tr>
                <td rowspan="1" colspan="1">
                  <italic>BERTlast</italic>
                </td>
                <td rowspan="1" colspan="1">0,521</td>
                <td rowspan="1" colspan="1">0,665</td>
                <td rowspan="1" colspan="1">0,711</td>
                <td rowspan="1" colspan="1">0,636</td>
                <td rowspan="1" colspan="1">0,648</td>
                <td rowspan="1" colspan="1">0,658</td>
                <td rowspan="1" colspan="1">0,736</td>
                <td rowspan="1" colspan="1">0,711</td>
                <td rowspan="1" colspan="1">0,661</td>
              </tr>
              <tr>
                <td rowspan="1" colspan="1">
                  <italic>BERTfinal</italic>
                </td>
                <td rowspan="1" colspan="1">0,657</td>
                <td rowspan="1" colspan="1">0,825</td>
                <td rowspan="1" colspan="1">0,842</td>
                <td rowspan="1" colspan="1">0,814</td>
                <td rowspan="1" colspan="1">0,771</td>
                <td rowspan="1" colspan="1">0,823</td>
                <td rowspan="1" colspan="1">0,817</td>
                <td rowspan="1" colspan="1">0,827</td>
                <td rowspan="1" colspan="1">0,797</td>
              </tr>
            </tbody>
          </table>
          <table-wrap-foot>
            <fn>
              <p>Note: Yearly performance of <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EIKAE">BERT</abbrev> models on the textual data. <abbrev xlink:title="area under the ROC curve" id="ABBRID0EMKAE">AUC</abbrev> - area under the ROC curve, <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EQKAE">BERT</abbrev> - Bidirectional Encoder Representations from Transformers, Average - average across 8 test years between 2014 and 2021.</p>
            </fn>
          </table-wrap-foot>
        </table-wrap>
        <p>We tested Hypothesis 1 by comparing our <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EWKAE">BERT</abbrev> models to the benchmark models of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>. The results indicate no significant differences between the models for the period 2000–2013, confirming that our <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0E5KAE">BERT</abbrev><sub>first</sub> model performs comparably. Furthermore, our <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EDLAE">BERT</abbrev><sub>first</sub> model continues to be effective in detecting fraud from 2014–2021, with an average <abbrev xlink:title="area under the ROC curve" id="ABBRID0EILAE">AUC</abbrev> score of 0.844, exceeding the benchmark’s 0.804. This confirms Hypothesis 1, showing that <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EMLAE">BERT</abbrev> remains effective in recent years.</p>
      </sec>
      <sec sec-type="﻿5.2. Hypothesis 2" id="SECID0EQLAE">
        <title>﻿5.2. Hypothesis 2</title>
        <p>Hypothesis 2 postulates that firms engaged in fraudulent activities use a different frequency of positive, negative, litigious, and uncertainty words in their MD&amp;A sections compared to non-fraudulent firms. To test this, we examine the occurrences of these specific word categories in both fraudulent and non-fraudulent MD&amp;A sections, aiming to determine whether linguistic characteristics differ between the two groups. Figure <xref ref-type="fig" rid="F3">3</xref> presents the average percentage of each word category—positive, negative, litigious, and uncertainty—used per year from 1994 to 2021 for both fraudulent and non-fraudulent filings.</p>
        <fig id="F3" position="float" orientation="portrait">
          <object-id content-type="arpha">02CA4BD8-1492-514B-8ABD-22B5F41EA67C</object-id>
          <label>Figure 3.</label>
          <caption>
            <p>Percentage of word occurrences in MD&amp;A sections for fraudulent vs. non-fraudulent companies. Note: The y-axis represents the average yearly percentage of each word category occurring in the MD&amp;A sections. The word categories (Positive, Negative, Uncertainty, and Litigious) are based on the classifications by <xref ref-type="bibr" rid="B31">Loughran and McDonald (2023)</xref>.</p>
          </caption>
          <graphic xlink:href="mab-99-061-g003.jpg" position="float" orientation="portrait" xlink:type="simple" id="oo_1315281.jpg">
            <uri content-type="original_file">https://binary.pensoft.net/fig/1315281</uri>
          </graphic>
        </fig>
        <p>We also conduct Ordinary Least Squares (<abbrev xlink:title="Ordinary Least Squares" id="ABBRID0ENMAE">OLS</abbrev>) regression analyses to test for significant differences in word category usage between fraudulent and non-fraudulent MD&amp;A sections over the period 1994–2021. The results show that fraudulent MD&amp;A sections contain significantly more positive words, with the difference being statistically significant at the 99% level. This finding aligns with the suggestion of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, confirming that fraudulent filings use more positive language. A separate test for the 1994–2013 period also showed a significant difference at the 5% level (not included in the manuscript for brevity).</p>
        <p>However, the analysis reveals no significant differences in the usage of negative, uncertainty, and litigious words between fraudulent and non-fraudulent firms at the 95% significance level, leading us to partially reject H2 for these categories. There is some indication, at the 90% significance level, that non-fraudulent MD&amp;A sections contain fewer negative words.</p>
      </sec>
      <sec sec-type="﻿5.3. Hypothesis 3" id="SECID0EWMAE">
        <title>﻿5.3. Hypothesis 3</title>
        <p>In Hypothesis 3, we propose that the frequency of word categories in non-fraudulent filings is more consistent over time compared to fraudulent filings. To test this, we examine whether the trend in the frequency of word categories in MD&amp;A sections differs between non-fraudulent and fraudulent firms.</p>
        <p>Interestingly, Figure <xref ref-type="fig" rid="F3">3</xref> reveals a stabilizing trend in the occurrence of word categories in recent years. This aligns with the study by <xref ref-type="bibr" rid="B35">Oyewole et al. (2024)</xref>, which suggests that the use of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EFNAE">AI</abbrev>-generated financial reports may contribute to this stabilization. To investigate this, we tested whether the slope of the regression for word usage trends was less steep during the last 14 years (2008–2021) compared to the first 14 years (1994–2007). The difference in slopes is shown in Figure <xref ref-type="fig" rid="F4">4</xref>, which illustrates the trends in positive, negative, uncertainty, and litigious word usage across both periods. The figure presents scatter plots and trend lines, with the first time period T1 (1994–2007) indicated in red and second time period T2 (2008–2021) in blue, allowing for a clear visual comparison of word usage over time for each category.</p>
        <fig id="F4" position="float" orientation="portrait">
          <object-id content-type="arpha">F77B4DAC-90CD-52DA-9582-ACDDC6DD6B8F</object-id>
          <label>Figure 4.</label>
          <caption>
            <p>Trends in word usage for positive, negative, uncertainty, and litigious words across two time periods. Note: The figure presents scatter plots and trend lines for the usage of positive, negative, uncertainty, and litigious words in MD&amp;A sections across two periods: T1 (1994–2007) shown in red, and T2 (2008–2021) shown in blue. It provides a visual comparison of word usage trends over time for each category, illustrating the difference in slopes between the two periods.</p>
          </caption>
          <graphic xlink:href="mab-99-061-g004.jpg" position="float" orientation="portrait" xlink:type="simple" id="oo_1315282.jpg">
            <uri content-type="original_file">https://binary.pensoft.net/fig/1315282</uri>
          </graphic>
        </fig>
        <p>Our regression analysis (not tabulated for brevity) examined word category trends (positive, negative, uncertainty, and litigious) between two periods (T1 and T2), as shown in Figure <xref ref-type="fig" rid="F4">4</xref>.</p>
        <p>The results show a significant decline in the use of positive words in both fraudulent and non-fraudulent MD&amp;A during T1. However, in T2, the decline in positive word usage was less steep for fraudulent MD&amp;A, indicating some stabilization, while the trend remained insignificant for non-fraudulent filings.</p>
        <p>Negative word usage was significantly higher in the first period for both groups. Interestingly, there was a significant reduction in the use of negative words in the last period, suggesting a shift in language over time.</p>
        <p>Both fraudulent and non-fraudulent MD&amp;A showed an increase in uncertainty words during T1, but this trend reversed to a decline in T2. Litigious word usage showed an insignificant positive trend in fraudulent MD&amp;A during T1, which shifted to a negative trend in T2. This suggests that fraudulent filings may be using less legalistic language over time, while the results for non-fraudulent filings were not significant.</p>
      </sec>
      <sec sec-type="﻿5.4. Hypothesis 4" id="SECID0ECOAE">
        <title>﻿5.4. Hypothesis 4</title>
        <p>In Hypothesis 4, we propose that firms engaged in fraudulent activities have a higher Fog Index, indicating more complex and difficult to read MD&amp;A sections compared to non-fraudulent firms. To test this, we calculated the average Gunning Fog Index for the MD&amp;A sections, with the results visualized in Figure <xref ref-type="fig" rid="F5">5</xref>. Contrary to our hypothesis, the results show that the average Fog Index of non-fraudulent firms is actually higher than that of fraudulent firms. This contradicts our expectation and the findings of <xref ref-type="bibr" rid="B34">Nakashima et al. (2022)</xref>, who suggested that fraudulent filings are harder to read, and <xref ref-type="bibr" rid="B33">Moffitt and Burns (2009)</xref>, who found that fraudulent 10-Ks tend to contain more complex language.</p>
        <fig id="F5" position="float" orientation="portrait">
          <object-id content-type="arpha">6DF27EC5-E708-5BE4-9343-B71EBB8CA474</object-id>
          <label>Figure 5.</label>
          <caption>
            <p>Average Fog Index per year for fraudulent vs. non-fraudulent MD&amp;A sections. Note: The y-axis displays the Fog Index, which measures the readability of the text. The blue solid line represents the average Fog Index for non-fraudulent MD&amp;A sections per year, while the red solid line represents the average Fog Index for fraudulent MD&amp;A sections per year.</p>
          </caption>
          <graphic xlink:href="mab-99-061-g005.jpg" position="float" orientation="portrait" xlink:type="simple" id="oo_1315283.jpg">
            <uri content-type="original_file">https://binary.pensoft.net/fig/1315283</uri>
          </graphic>
        </fig>
        <p>We also conducted a regression analysis to further explore the relationship between the Fog Index and fraudulent vs. non-fraudulent MD&amp;A sections over time. While Figure <xref ref-type="fig" rid="F5">5</xref> suggests that the average Fog Index for fraudulent firms is lower in many years, the difference is not statistically significant. Therefore, we cannot conclusively state that firms engaging in fraudulent activities have a significantly lower Fog Index than non-fraudulent firms. However, the results do indicate a statistically significant increase in the Fog Index over time, suggesting that MD&amp;A sections are becoming more difficult to read overall.</p>
      </sec>
    </sec>
    <sec sec-type="﻿6. Discussion and conclusion" id="SECID0EGPAE">
      <title>﻿6. Discussion and conclusion</title>
      <p>Our research explores the use of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EMPAE">AI</abbrev>, specifically <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EQPAE">BERT</abbrev> models, to detect fraud through textual analysis of MD&amp;A sections in 10-K filings. The study tests four hypotheses on the effectiveness of <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EUPAE">BERT</abbrev> models and linguistic indicators of fraudulent activity.</p>
      <p><italic>Hypothesis 1</italic> confirmed that <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0E3PAE">BERT</abbrev> models remain effective in detecting fraud, replicating the findings of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>. This highlights the continued relevance of <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EEQAE">BERT</abbrev> models for fraud detection, aligning with <xref ref-type="bibr" rid="B25">Kureljusic and Karger (2023)</xref>, who emphasize <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EMQAE">AI</abbrev>’s usefulness in financial accounting. Given the findings of <xref ref-type="bibr" rid="B35">Oyewole et al. (2024)</xref> on <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EUQAE">AI</abbrev>-generated reports, it is possible that the <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EYQAE">AI</abbrev> models themselves are influencing the linguistic patterns detected by <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0E3QAE">BERT</abbrev>, suggesting an area for further research.</p>
      <p><italic>Hypothesis 2</italic> tested whether fraudulent MD&amp;A sections use a higher frequency of positive, litigious, and uncertainty words, and fewer negative words. The results supported the hypothesis for positive words, consistent with <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, suggesting that fraudulent firms use positive language to mask fraud. However, contrary to expectations, fraudulent MD&amp;A sections also contained more negative words, indicating a more complex linguistic strategy. No significant differences were found for uncertainty and litigious words.</p>
      <p><italic>Hypothesis 3</italic> explored whether non-fraudulent MD&amp;A sections show more consistent trends in word categories over time compared to fraudulent sections. While we found some evidence of a decreasing trend in litigious word usage in non-fraudulent filings, the trend lines for both fraudulent and non-fraudulent sections became more stable in recent years. This stabilization could reflect the increasing use of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EMRAE">AI</abbrev>-generated reports, as suggested by <xref ref-type="bibr" rid="B35">Oyewole et al. (2024)</xref>.</p>
      <p><italic>Hypothesis 4</italic> examined whether fraudulent MD&amp;A sections are less readable, as measured by the Fog Index. Contrary to the hypothesis, non-fraudulent MD&amp;As had a slightly higher Fog Index, although the difference was not statistically significant. This finding contrasts with <xref ref-type="bibr" rid="B34">Nakashima et al. (2022)</xref>, who found that fraudulent reports were harder to read. The overall increase in the Fog Index over time suggests that corporate reports are becoming more complex, possibly due to regulatory or industry changes.</p>
      <p>A limitation of this study, noted by <xref ref-type="bibr" rid="B26">Larcker and Zakolyukina (2012)</xref> and <xref ref-type="bibr" rid="B16">Ghosh et al. (2015)</xref>, is the reliance on word-count methods based on psychosocial dictionaries, which may not capture contextual meaning. Future research could improve word categorization by combining sentiment analysis with word structures, as suggested by <xref ref-type="bibr" rid="B16">Ghosh et al. (2015)</xref>. Another limitation is our reliance on the AAER database to identify fraudulent firms, which could benefit from broader validation.</p>
      <p>In summary, this study validates and extends the findings of <xref ref-type="bibr" rid="B3">Bhattacharya and Mićković (2024)</xref>, confirming that <abbrev xlink:title="Bidirectional Encoder Representations from Transformers" id="ABBRID0EQSAE">BERT</abbrev> models are highly effective in detecting fraudulent activity. The results indicate that fraudulent MD&amp;A sections contain more positive and negative words, though no significant differences were found for uncertainty and litigious words. The analysis of readability found no significant differences between fraudulent and non-fraudulent MD&amp;As, though readability has decreased over time. This study offers valuable insights into the role of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EUSAE">AI</abbrev> in fraud detection and provides practical implications for auditors and regulators. Future research should further investigate linguistic features associated with fraud and explore the impact of <abbrev xlink:title="Artificial Intelligence" id="ABBRID0EYSAE">AI</abbrev>-generated financial reports on fraud detection.</p>
      <boxed-text id="box1" position="float" orientation="portrait">
        <p><bold>F.H.E. Ketelaar MSc – Florian</bold><sup><xref ref-type="fn" rid="en1">1</xref></sup> is an <abbrev xlink:title="Artificial Intelligence" id="ABBRID0ELTAE">AI</abbrev> enthusiast and auditor at a Big 4 audit firm.</p>
      </boxed-text>
      <boxed-text id="box2" position="float" orientation="portrait">
        <p><bold>Dr. A. Mićković – Ana</bold> is an Assistant Professor of Accounting at the University of Amsterdam.</p>
      </boxed-text>
    </sec>
  </body>
  <back>
    <fn-group>
      <title>Note</title>
      <fn id="en1">
        <p>Florian Ketelaar is one of the winners of the MAB Thesis Award 2024. This article is based on his master thesis.</p>
      </fn>
    </fn-group>
    <ref-list>
      <title>﻿References</title>
      <ref id="B1">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Beasley</surname><given-names>MS</given-names></name><name name-style="western"><surname>Carcello</surname><given-names>JV</given-names></name><name name-style="western"><surname>Hermanson</surname><given-names>DR</given-names></name><name name-style="western"><surname>Neal</surname><given-names>TL</given-names></name></person-group> (<year>2010</year>) Fraudulent Financial Reporting 1998–2007: An Analysis of US Public Companies. In: COSO Report.</mixed-citation>
      </ref>
      <ref id="B2">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Bertomeu</surname><given-names>J</given-names></name><name name-style="western"><surname>Cheynel</surname><given-names>E</given-names></name><name name-style="western"><surname>Floyd</surname><given-names>E</given-names></name><name name-style="western"><surname>Pan</surname><given-names>W</given-names></name></person-group> (<year>2021</year>) <article-title>Using machine learning to detect misstatements.</article-title><source>Review of Accounting Studies</source><volume>26</volume>: <fpage>468</fpage>–<lpage>519</lpage>. <ext-link xlink:href="10.1007/s11142-020-09563-8" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1007/s11142-020-09563-8</ext-link></mixed-citation>
      </ref>
      <ref id="B3">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Bhattacharya</surname><given-names>I</given-names></name><name name-style="western"><surname>Mićković</surname><given-names>A</given-names></name></person-group> (<year>2024</year>) Accounting fraud detection using contextual language learning. International Journal of Accounting Information Systems 53: 100682. <ext-link xlink:href="10.1016/j.accinf.2024.100682" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1016/j.accinf.2024.100682</ext-link></mixed-citation>
      </ref>
      <ref id="B4">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Bochkay</surname><given-names>K</given-names></name><name name-style="western"><surname>Brown</surname><given-names>SV</given-names></name><name name-style="western"><surname>Leone</surname><given-names>AJ</given-names></name><name name-style="western"><surname>Tucker</surname><given-names>JW</given-names></name></person-group> (<year>2023</year>) <article-title>Textual Analysis in Accounting: What’s Next?*.</article-title><source>Contemporary Accounting Research</source><volume>40</volume>(<issue>2</issue>): <fpage>765</fpage>–<lpage>805</lpage>. <ext-link xlink:href="10.1111/1911-3846.12825" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/1911-3846.12825</ext-link></mixed-citation>
      </ref>
      <ref id="B5">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Brown</surname><given-names>SV</given-names></name><name name-style="western"><surname>Tucker</surname><given-names>JW</given-names></name></person-group> (<year>2011</year>) <article-title>Large‐sample evidence on firms’ year‐over‐year MD&amp;A modifications.</article-title><source>Journal of Accounting Research,</source><volume>49</volume>(<issue>2</issue>): <fpage>309</fpage>–<lpage>346</lpage>. <ext-link xlink:href="10.1111/j.1475-679X.2010.00396.x" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/j.1475-679X.2010.00396.x</ext-link></mixed-citation>
      </ref>
      <ref id="B6">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Brown</surname><given-names>NC</given-names></name><name name-style="western"><surname>Crowley</surname><given-names>RM</given-names></name><name name-style="western"><surname>Elliott</surname><given-names>WB</given-names></name></person-group> (<year>2020</year>) <article-title>What Are You Saying? Using topic to Detect Financial Misreporting.</article-title><source>Journal of Accounting Research</source><volume>58</volume>(<issue>1</issue>): <fpage>237</fpage>–<lpage>291</lpage>. <ext-link xlink:href="10.1111/1475-679X.12294" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/1475-679X.12294</ext-link></mixed-citation>
      </ref>
      <ref id="B7">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Burkov</surname><given-names>A</given-names></name></person-group> (<year>2019</year>) The hundred-page machine learning book the hundred-page machine learning book. Andriy Burkov.</mixed-citation>
      </ref>
      <ref id="B8">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Campa</surname><given-names>D</given-names></name><name name-style="western"><surname>Quagli</surname><given-names>A</given-names></name><name name-style="western"><surname>Ramassa</surname><given-names>P</given-names></name></person-group> (<year>2023</year>) <article-title>The roles and interplay of enforcers and auditors in the context of accounting fraud: a review of the accounting literature.</article-title><source>Journal of Accounting Literature</source><volume>47</volume>(<issue>5</issue>): <fpage>151</fpage>–<lpage>183</lpage>. <ext-link xlink:href="10.1108/JAL-07-2023-0134" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1108/JAL-07-2023-0134</ext-link></mixed-citation>
      </ref>
      <ref id="B9">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Cole</surname><given-names>CJ</given-names></name><name name-style="western"><surname>Jones</surname><given-names>CL</given-names></name></person-group> (<year>2004</year>) <article-title>The usefulness of MD&amp;A disclosures in the retail industry.</article-title><source>Journal of Accounting, Auditing &amp; Finance</source><volume>19</volume>(<issue>4</issue>): <fpage>361</fpage>–<lpage>388</lpage>. <ext-link xlink:href="10.1177/0148558X0401900401" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1177/0148558X0401900401</ext-link></mixed-citation>
      </ref>
      <ref id="B10">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Craja</surname><given-names>P</given-names></name><name name-style="western"><surname>Kim</surname><given-names>A</given-names></name><name name-style="western"><surname>Lessmann</surname><given-names>S</given-names></name></person-group> (<year>2020</year>) Deep learning for detecting financial statement fraud. Decision Support Systems 139: 113421. <ext-link xlink:href="10.1016/j.dss.2020.113421" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1016/j.dss.2020.113421</ext-link></mixed-citation>
      </ref>
      <ref id="B11">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Davenport</surname><given-names>TH</given-names></name></person-group> (<year>2016</year>) Deloitte - The power of advanced audit analytics. <ext-link xlink:href="https://www2.deloitte.com/content/dam/Deloitte/us/Documents/deloitte-analytics/us-daadvanced-audit-analytics.pdf" ext-link-type="uri" xlink:type="simple">https://www2.deloitte.com/content/dam/Deloitte/us/Documents/deloitte-analytics/us-daadvanced-audit-analytics.pdf</ext-link> [Accessed: 1 June. 2024]</mixed-citation>
      </ref>
      <ref id="B12">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Dechow</surname><given-names>PM</given-names></name><name name-style="western"><surname>Ge</surname><given-names>W</given-names></name><name name-style="western"><surname>Larson</surname><given-names>CR</given-names></name><name name-style="western"><surname>Sloan</surname><given-names>RG</given-names></name></person-group> (<year>2011</year>) <article-title>Predicting material accounting misstatements.</article-title><source>Contemporary Accounting Research</source><volume>28</volume>(<issue>1</issue>): <fpage>17</fpage>–<lpage>82</lpage>. <ext-link xlink:href="10.1111/j.1911-3846.2010.01041.x" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/j.1911-3846.2010.01041.x</ext-link></mixed-citation>
      </ref>
      <ref id="B13">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Devlin</surname><given-names>J</given-names></name></person-group> (<year>2018</year>) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.</mixed-citation>
      </ref>
      <ref id="B14">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Eisenhardt</surname><given-names>KM</given-names></name></person-group> (<year>1989</year>) <article-title>Agency Theory: An Assessment and Review.</article-title><source>The Academy of Management Review</source><volume>14</volume>(<issue>1</issue>): <fpage>57</fpage>–<lpage>74</lpage>. <ext-link xlink:href="10.2307/258191" ext-link-type="doi" xlink:type="simple">https://doi.org/10.2307/258191</ext-link></mixed-citation>
      </ref>
      <ref id="B15">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Goel</surname><given-names>S</given-names></name><name name-style="western"><surname>Uzuner</surname><given-names>O</given-names></name></person-group> (<year>2016</year>) <article-title>Do sentiments matter in fraud detection? Estimating semantic orientation of annual reports.</article-title><source>Intelligent Systems in Accounting, Finance and Management</source><volume>23</volume>(<issue>3</issue>): <fpage>215</fpage>–<lpage>239</lpage>. <ext-link xlink:href="10.1002/isaf.1392" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1002/isaf.1392</ext-link></mixed-citation>
      </ref>
      <ref id="B16">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Ghosh</surname><given-names>D</given-names></name><name name-style="western"><surname>Guo</surname><given-names>W</given-names></name><name name-style="western"><surname>Muresan</surname><given-names>S</given-names></name></person-group> (<year>2015</year>) [September] Sarcastic or not: Word embeddings to predict the literal or sarcastic meaning of words. In proceedings of the 2015 conference on empirical methods in natural language processing,1003–1012. <ext-link xlink:href="10.18653/v1/D15-1116" ext-link-type="doi" xlink:type="simple">https://doi.org/10.18653/v1/D15-1116</ext-link></mixed-citation>
      </ref>
      <ref id="B17">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Gunning</surname><given-names>R</given-names></name></person-group> (<year>1952</year>) The technique of clear writing.</mixed-citation>
      </ref>
      <ref id="B18">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Hancock</surname><given-names>JT</given-names></name></person-group> (<year>2007</year>) <article-title>Digital deception.</article-title><source>Oxford Handbook of Internet Psychology</source><volume>61</volume>(<issue>5</issue>): <fpage>289301</fpage>.</mixed-citation>
      </ref>
      <ref id="B19">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Hasan</surname><given-names>AR</given-names></name></person-group> (<year>2021</year>) <article-title>Artificial Intelligence (AI) in accounting &amp; auditing: A Literature review.</article-title><source>Open Journal of Business and Management</source><volume>10</volume>(<issue>1</issue>): <fpage>440</fpage>–<lpage>465</lpage>. <ext-link xlink:href="10.4236/ojbm.2022.101026" ext-link-type="doi" xlink:type="simple">https://doi.org/10.4236/ojbm.2022.101026</ext-link></mixed-citation>
      </ref>
      <ref id="B20">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Hassija</surname><given-names>V</given-names></name><name name-style="western"><surname>Chamola</surname><given-names>V</given-names></name><name name-style="western"><surname>Mahapatra</surname><given-names>A</given-names></name><name name-style="western"><surname>Singal</surname><given-names>A</given-names></name><name name-style="western"><surname>Goel</surname><given-names>D</given-names></name><name name-style="western"><surname>Huang</surname><given-names>K</given-names></name><name name-style="western"><surname>Scardapane</surname><given-names>S</given-names></name><name name-style="western"><surname>Spinelli</surname><given-names>I</given-names></name><name name-style="western"><surname>Mahmud</surname><given-names>M</given-names></name><name name-style="western"><surname>Hussain</surname><given-names>A</given-names></name></person-group> (<year>2024</year>) <article-title>Interpreting black-box models: a review on explainable artificial intelligence.</article-title><source>Cognitive Computation</source><volume>16</volume>(<issue>1</issue>): <fpage>45</fpage>–<lpage>74</lpage>. <ext-link xlink:href="10.1007/s12559-023-10179-8" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1007/s12559-023-10179-8</ext-link></mixed-citation>
      </ref>
      <ref id="B21">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Huang</surname><given-names>AH</given-names></name><name name-style="western"><surname>Wang</surname><given-names>H</given-names></name><name name-style="western"><surname>Yang</surname><given-names>Y</given-names></name></person-group> (<year>2023</year>) <article-title>FinBERT: A large language model for extracting information from financial text.</article-title><source>Contemporary Accounting Research</source><volume>40</volume>(<issue>2</issue>): <fpage>806</fpage>–<lpage>841</lpage>. <ext-link xlink:href="10.1111/1911-3846.12832" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/1911-3846.12832</ext-link></mixed-citation>
      </ref>
      <ref id="B22">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Joyce</surname><given-names>EJ</given-names></name><name name-style="western"><surname>Biddle</surname><given-names>GC</given-names></name></person-group> (<year>1981</year>) <article-title>Anchoring and adjustment in probabilistic inference in auditing.</article-title><source>Journal of Accounting Research</source><volume>19</volume>(<issue>1</issue>): <fpage>120</fpage>–<lpage>145</lpage>. <ext-link xlink:href="10.2307/2490965" ext-link-type="doi" xlink:type="simple">https://doi.org/10.2307/2490965</ext-link></mixed-citation>
      </ref>
      <ref id="B23">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Ikhsan</surname><given-names>WM</given-names></name><name name-style="western"><surname>Ednoer</surname><given-names>EH</given-names></name><name name-style="western"><surname>Kridantika</surname><given-names>WS</given-names></name><name name-style="western"><surname>Firmansyah</surname><given-names>A</given-names></name></person-group> (<year>2022</year>) <article-title>Fraud detection automation through data analytics and artificial intelligence.</article-title><source>Riset</source><volume>4</volume>(<issue>2</issue>): <fpage>103</fpage>–<lpage>119</lpage>. <ext-link xlink:href="10.37641/riset.v4i2.166" ext-link-type="doi" xlink:type="simple">https://doi.org/10.37641/riset.v4i2.166</ext-link></mixed-citation>
      </ref>
      <ref id="B24">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Kieso</surname><given-names>DE</given-names></name><name name-style="western"><surname>Warfield</surname><given-names>TD</given-names></name><name name-style="western"><surname>Weygandt</surname><given-names>JJ</given-names></name></person-group> (<year>2020</year>) Intermediate accounting: IFRS edition.</mixed-citation>
      </ref>
      <ref id="B25">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Kureljusic</surname><given-names>M</given-names></name><name name-style="western"><surname>Karger</surname><given-names>E</given-names></name></person-group> (<year>2023</year>) <article-title>Forecasting in financial accounting with artificial intelligence – A systematic literature review and future research agenda.</article-title><source>Journal of Applied Accounting Research</source><volume>25</volume>(<issue>1</issue>): <fpage>81</fpage>–<lpage>104</lpage>. <ext-link xlink:href="10.1108/JAAR-06-2022-0146" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1108/JAAR-06-2022-0146</ext-link></mixed-citation>
      </ref>
      <ref id="B26">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Larcker</surname><given-names>DF</given-names></name><name name-style="western"><surname>Zakolyukina</surname><given-names>AA</given-names></name></person-group> (<year>2012</year>) <article-title>Detecting deceptive discussions in conference calls.</article-title><source>Journal of Accounting Research</source><volume>50</volume>(<issue>2</issue>): <fpage>495</fpage>–<lpage>540</lpage>. <ext-link xlink:href="10.1111/j.1475-679X.2012.00450.x" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/j.1475-679X.2012.00450.x</ext-link></mixed-citation>
      </ref>
      <ref id="B27">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Lee</surname><given-names>CS</given-names></name><name name-style="western"><surname>Tajudeen</surname><given-names>FP</given-names></name></person-group> (<year>2020</year>) <article-title>Usage and impact of artificial intelligence on accounting: Evidence from Malaysian organisations.</article-title><source>Asian Journal of Business and Accounting</source><volume>13</volume>(<issue>1</issue>): <fpage>213</fpage>–<lpage>239</lpage>. <ext-link xlink:href="10.22452/ajba.vol13no1.8" ext-link-type="doi" xlink:type="simple">https://doi.org/10.22452/ajba.vol13no1.8</ext-link></mixed-citation>
      </ref>
      <ref id="B28">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Li</surname><given-names>F</given-names></name></person-group> (<year>2008</year>) <article-title>Annual report readability, current earnings, and earnings persistence.</article-title><source>Journal of Accounting and Economics</source><volume>45</volume>(<issue>2–3</issue>): <fpage>221</fpage>–<lpage>247</lpage>. <ext-link xlink:href="10.1016/j.jacceco.2008.02.003" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1016/j.jacceco.2008.02.003</ext-link></mixed-citation>
      </ref>
      <ref id="B29">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Loughran</surname><given-names>T</given-names></name><name name-style="western"><surname>McDonald</surname><given-names>B</given-names></name></person-group> (<year>2011</year>) <article-title>When is a liability not a liability? textual analysis, dictionaries, and 10-Ks.</article-title><source>Journal of Finance</source><volume>66</volume>: <fpage>35</fpage>–<lpage>65</lpage>. <ext-link xlink:href="10.1111/j.1540-6261.2010.01625.x" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/j.1540-6261.2010.01625.x</ext-link></mixed-citation>
      </ref>
      <ref id="B30">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Loughran</surname><given-names>T</given-names></name><name name-style="western"><surname>McDonald</surname><given-names>B</given-names></name></person-group> (<year>2016</year>) <article-title>Textual analysis in accounting and finance: A survey.</article-title><source>Journal of Accounting Research</source><volume>54</volume>(<issue>4</issue>): <fpage>1187</fpage>–<lpage>1230</lpage>. <ext-link xlink:href="10.1111/1475-679X.12123" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/1475-679X.12123</ext-link></mixed-citation>
      </ref>
      <ref id="B31">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Loughran</surname><given-names>T</given-names></name><name name-style="western"><surname>McDonald</surname><given-names>B</given-names></name></person-group> (<year>2023</year>) Measuring firm complexity. Journal of Financial and Quantitative Analysis. <ext-link xlink:href="10.2139/ssrn.3645372" ext-link-type="doi" xlink:type="simple">https://doi.org/10.2139/ssrn.3645372</ext-link></mixed-citation>
      </ref>
      <ref id="B32">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Martinc</surname><given-names>M</given-names></name><name name-style="western"><surname>Pollak</surname><given-names>S</given-names></name><name name-style="western"><surname>Robnik-Šikonja</surname><given-names>M</given-names></name></person-group> (<year>2021</year>) <article-title>Supervised and Unsupervised Neural Approaches to Text Readability.</article-title><source>Computational Linguistics</source><volume>47</volume>(<issue>1</issue>): <fpage>141</fpage>–<lpage>179</lpage>. <ext-link xlink:href="10.1162/coli_a_00398" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1162/coli_a_00398</ext-link></mixed-citation>
      </ref>
      <ref id="B33">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Moffitt</surname><given-names>K</given-names></name><name name-style="western"><surname>Burns</surname><given-names>MB</given-names></name></person-group> (<year>2009</year>) What does that mean? Investigating obfuscation and readability cues as indicators of deception in fraudulent financial reports. AMCIS 2009 Proceedings, 399.</mixed-citation>
      </ref>
      <ref id="B34">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Nakashima</surname><given-names>M</given-names></name><name name-style="western"><surname>Hirose</surname><given-names>Y</given-names></name><name name-style="western"><surname>Hirai</surname><given-names>H</given-names></name></person-group> (<year>2022</year>) Fraud Detection by Focusing on Readability of MD&amp;A Disclosure: Evidence from Japan. Journal of Forensic and Investigative Accounting 14(2).</mixed-citation>
      </ref>
      <ref id="B35">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Oyewole</surname><given-names>AT</given-names></name><name name-style="western"><surname>Adeoye</surname><given-names>OB</given-names></name><name name-style="western"><surname>Addy</surname><given-names>WA</given-names></name><name name-style="western"><surname>Okoye</surname><given-names>CC</given-names></name><name name-style="western"><surname>Ofodile</surname><given-names>OC</given-names></name><name name-style="western"><surname>Ugochukwu</surname><given-names>CE</given-names></name></person-group> (<year>2024</year>) <article-title>Automating financial reporting with natural language processing: A review and case analysis.</article-title><source>World Journal of Advanced Research and Reviews</source><volume>21</volume>(<issue>3</issue>): <fpage>575</fpage>–<lpage>589</lpage>. <ext-link xlink:href="10.30574/wjarr.2024.21.3.0688" ext-link-type="doi" xlink:type="simple">https://doi.org/10.30574/wjarr.2024.21.3.0688</ext-link></mixed-citation>
      </ref>
      <ref id="B36">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Purda</surname><given-names>L</given-names></name><name name-style="western"><surname>Skillicorn</surname><given-names>D</given-names></name></person-group> (<year>2015</year>) <article-title>Accounting variables, deception, and a bag of words: Assessing the tools of fraud detection.</article-title><source>Contemporary Accounting Research</source><volume>32</volume>(<issue>3</issue>): <fpage>1193</fpage>–<lpage>1223</lpage>. <ext-link xlink:href="10.1111/1911-3846.12089" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1111/1911-3846.12089</ext-link></mixed-citation>
      </ref>
      <ref id="B37">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Ranta</surname><given-names>M</given-names></name><name name-style="western"><surname>Ylinen</surname><given-names>M</given-names></name><name name-style="western"><surname>Järvenpää</surname><given-names>M</given-names></name></person-group> (<year>2023</year>) <article-title>Machine learning in management accounting research: Literature review and pathways for the future.</article-title><source>European Accounting Review</source><volume>32</volume>(<issue>3</issue>): <fpage>607</fpage>–<lpage>636</lpage>. <ext-link xlink:href="10.1080/09638180.2022.2137221" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1080/09638180.2022.2137221</ext-link></mixed-citation>
      </ref>
      <ref id="B38">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Rijsenbilt</surname><given-names>A</given-names></name><name name-style="western"><surname>Commandeur</surname><given-names>H</given-names></name></person-group> (<year>2013</year>) <article-title>Narcissus enters the courtroom: CEO narcissism and fraud.</article-title><source>Journal of Business Ethics</source><volume>117</volume>: <fpage>413</fpage>–<lpage>429</lpage>. <ext-link xlink:href="10.1007/s10551-012-1528-7" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1007/s10551-012-1528-7</ext-link></mixed-citation>
      </ref>
      <ref id="B39">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Sutton</surname><given-names>SG</given-names></name><name name-style="western"><surname>Holt</surname><given-names>M</given-names></name><name name-style="western"><surname>Arnold</surname><given-names>V</given-names></name></person-group> (<year>2016</year>) <article-title>“The reports of my death are greatly exaggerated”—Artificial intelligence research in accounting.</article-title><source>International Journal of Accounting Information Systems</source><volume>22</volume>: <fpage>60</fpage>–<lpage>73</lpage>. <ext-link xlink:href="10.1016/j.accinf.2016.07.005" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1016/j.accinf.2016.07.005</ext-link></mixed-citation>
      </ref>
      <ref id="B40">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Tang</surname><given-names>XB</given-names></name><name name-style="western"><surname>Liu</surname><given-names>GC</given-names></name><name name-style="western"><surname>Yang</surname><given-names>J</given-names></name><name name-style="western"><surname>Wei</surname><given-names>W</given-names></name></person-group> (<year>2018</year>) <article-title>Knowledge-based financial statement fraud detection system: Based on an ontology and a decision tree.</article-title><source>Knowledge Organization</source><volume>45</volume>(<issue>3</issue>): <fpage>205</fpage>–<lpage>219</lpage>. <ext-link xlink:href="10.5771/0943-7444-2018-3-205" ext-link-type="doi" xlink:type="simple">https://doi.org/10.5771/0943-7444-2018-3-205</ext-link></mixed-citation>
      </ref>
      <ref id="B41">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Wang</surname><given-names>JH</given-names></name><name name-style="western"><surname>Liao</surname><given-names>YL</given-names></name><name name-style="western"><surname>Tsai</surname><given-names>TM</given-names></name><name name-style="western"><surname>Hung</surname><given-names>G</given-names></name></person-group> (<year>2006</year>) Technology-based financial frauds in Taiwan: issues and approaches. IEEE International Conference on Systems, Man and Cybernetics, Vol. 2, 1120–1124. IEEE. <ext-link xlink:href="10.1109/ICSMC.2006.384550" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1109/ICSMC.2006.384550</ext-link></mixed-citation>
      </ref>
      <ref id="B42">
        <mixed-citation xlink:type="simple"><person-group><name name-style="western"><surname>Zhang</surname><given-names>Y</given-names></name><name name-style="western"><surname>Xiong</surname><given-names>F</given-names></name><name name-style="western"><surname>Xie</surname><given-names>Y</given-names></name><name name-style="western"><surname>Fan</surname><given-names>X</given-names></name><name name-style="western"><surname>Gu</surname><given-names>H</given-names></name></person-group> (<year>2020</year>) <article-title>The impact of artificial intelligence and blockchain on the accounting profession.</article-title><source>IEEE Accessvol</source><volume>8</volume>: <fpage>110461</fpage>–<lpage>110477</lpage>. <ext-link xlink:href="10.1109/ACCESS.2020.3000505" ext-link-type="doi" xlink:type="simple">https://doi.org/10.1109/ACCESS.2020.3000505</ext-link></mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>
