MAB-scriptieprijs
Print
MAB-scriptieprijs
Signal or noise? The role of firm narratives in earnings prediction
expand article infoZixian Yin, Alexandre Madelaine§
‡ Erasmus University of Rotterdam, Cappelle aan den IJssel, Netherlands
§ Erasmus University of Rotterdam, Rotterdam, Netherlands
Open Access

Abstract

This study examines whether narrative disclosures in corporate financial reports enhance the prediction of future firm performance. While quantitative financial data such as current earnings and cash flows are established predictors, the added value of narratives remains uncertain. Using machine learning and topic modeling on over 5,000 10-K filings, we test whether narrative features improve predictive accuracy. Results indicate that narratives add limited and inconsistent value, though sections like Management Discussion and Risk Factors have higher predictive power. Our findings underline both the promise and current limitations of integrating textual analysis into financial prediction models.

Keywords

Textual analysis, narrative disclosures, financial prediction, annual reports, machine learning, topic modeling

Relevance to practice

This research informs analysts, investors, and regulators about the practical limitations of using automated textual analysis in financial prediction. It shows that while certain report sections provide valuable insights, textual data must be carefully selected and processed to meaningfully complement traditional quantitative financial data in predicting firm performance.

1. Introduction

Quantitative information is regarded as easier for investors to process and compare, offering advantages in decision-making (Viswanathan and Childers 1996; Engelberg 2008; Huang et al. 2014; Liberti and Petersen 2019; Campbell et al. 2025; Madelaine et al. 2026). Yet, numbers rarely speak for themselves. They often require the interpretive context provided by narrative disclosures (Ahn et al. 2022; Allee et al. 2023). In some settings, such ‘soft’ information can even convey insights that are more informative than traditional quantitative indicators alone (Brockman and Cicon 2013). Despite this potential, the virtually unlimited amount of narrative content in annual reports raises a practical concern for practitioners and researchers: when does more text stop adding signal and start adding noise? In this paper, we offer a cautionary examination of that question in the context of short-horizon earnings prediction. Specifically, we test whether narrative topic features extracted from thousands of 10-K filings provide incremental predictive value, beyond quantitative financial data, when forecasting next-year earnings changes.

We investigate three research questions:

  • whether 10-K annual reports provide incremental predictive power for earnings changes beyond traditional financial variables;
  • which topics within these reports are most strongly associated with earnings changes; and
  • which sections of these reports are most predictive of earnings changes.

The predictive value of structured financial data has been established by extensive research on earnings prediction. For instance, Ou and Penman (1989) establish that financial ratios predict changes in earnings and produce profitable investment strategies. More recently, Chen et al. (2022) found that ensemble methods – techniques that combine multiple models – can significantly improve earnings prediction accuracy. Yet, despite the potential relevance of narrative disclosures, most research in this area has largely overlooked them.

Our research questions are not without tension. Unlike quantitative data, narratives are high-dimensional, noisy, and not easily comparable across firms, in particular because relevant information may be diffuse or buried in regulatory language. Dyer et al. (2017) conduct descriptive analysis of the evolution of 10-K filings, finding that regulatory requirements have led to increasingly lengthy documents with declining readability and increasing boilerplate language, particularly in topics related to risks, fair value, and internal controls. Therefore, it is unclear whether narratives in those filings contain any predictive signals or constitute mere noise.

To capture narrative content, we employ Latent Dirichlet Allocation (LDA) to extract topic distributions from 10-K filings of S&P 500 companies between 2013 and 2023. Then, we develop prediction models for earnings changes using Lasso, Random Forest, and XGBoost, both with and without LDA-derived topic features. Our main findings indicate that, although topic modeling offers valuable insights into narrative patterns in 10-K filings, we do not find robust evidence that topic-based features meaningfully complement traditional financial variables in predicting earnings changes. These results call for a cautious interpretation of the predictive value of textual analysis and underscore the importance of critically evaluating emerging methodologies in financial predictions.

2. Literature review

2.1. Prediction of earnings changes

Reported earnings are a key metric for market participants when considering investment decisions (e.g., Beyer et al. (2010)), reflecting operational fundamentals more than transitory market noise (Penman 2013). The prediction of earnings changes has long been a cornerstone of fundamental analysis. Prior research shows that financial statement variables can predict future earnings (Ou and Penman 1989; Penman and Zhang 2006). Building on this foundation, Chen et al. (2022) set a new academic benchmark by applying ensemble learning, such as Random Forest and Gradient Boosting, to a granular set of over 4,000 distinct financial items. Their models achieved an Area Under the Curve (AUC) value of 68.66%, significantly outperforming traditional linear regressions. While their work identifies which financial variables carry the most information, their models exclude unstructured textual content, leaving open the question of whether narrative disclosures can further enhance predictive performance.

2.2. 10-K narrative disclosures: information signal vs. noise

While quantitative data provides the ‘what,’ narrative disclosures in 10-K filings provide the ‘why,’ offering interpretive context and forward-looking insights into management’s expectations. However, the practical utility of these narratives is under scrutiny. Dyer et al. (2017) provide a comprehensive evolutionary analysis of 10-K filings from 1996 to 2013, documenting an increase in document length driven largely by regulatory requirements. Crucially for practitioners, they found that this expansion often results in declining readability and an increase in boilerplate language, that is, repetitive text that may obscure rather than reveal firm-specific risks. While narratives like the Management’s Discussion and Analysis (MD&A) can signal future performance (Brown and Tucker 2011), they can also be used strategically to obfuscate poor results through linguistic complexity (Li 2008). For investors and analysts, the primary challenge lies in distinguishing between genuine economic signals and regulatory noise. Our study addresses this by testing whether automated tools can effectively filter this noise to enhance prediction.

2.3. Textual analysis and LDA

Textual analysis has become central in accounting and finance as narrative disclosures complement traditional financial information. To process the thousands of pages in a typical 10-K corpus, accounting research has shifted from simple word-counting to sophisticated Natural Language Processing (NLP). Early studies focus on readability (Li 2008) and sentiment (Loughran and McDonald 2011), offering simple textual measures but limited semantic depth. More recent NLP tools, including transformer-based models like FinBERT (Huang et al. 2023), capture richer narrative patterns, though their complexity and high dimensionality make the underlying drivers of prediction difficult to interpret. In this context, topic modeling, particularly LDA, provides a useful middle ground. LDA represents each document as a mixture of latent topics, allowing users to identify the specific themes (e.g., ‘Renewable energy’ or ‘Corporate governance’) being discussed. Each topic is characterized by a distribution over words, producing interpretable topic proportions that can be readily incorporated into predictive models.

LDA has been widely applied to various financial texts. For example, Campbell et al. (2014) extract topics from risk disclosures in 10-K filings; Huang et al. (2018) compare analyst reports with earnings calls, and Feuerriegel et al. (2016) study market reactions to news announcements. Beyond those analyses, several studies have used topic distributions as predictive variables. Hanley and Hoberg (2010) incorporate topics into regressions predicting IPO underpricing, and Brown et al. (2020) use topic proportions from 10-K filings to predict financial misreporting. For practitioners, LDA offers an automated way to analyze thousands of filings simultaneously, identifying thematic shifts. By integrating these LDA-derived topics with the financial benchmark established by Chen et al. (2022), we provide a test of whether automated textual analysis truly adds value in the context of short-horizon earnings prediction.

3. Methodology

3.1. Sample and variables

The sample consists of S&P 500 firms between 2013 and 2023 for three reasons:

  1. the index captures roughly 80% of U.S. equity market capitalization (S&P Global 2025);
  2. these firms follow strict reporting standards, producing documents that are standardized yet retain firm-specific narrative content, making them suitable for topic modeling (Li 2010); and
  3. the sample size balances representativeness, computational feasibility, and statistical power.

To ensure a consistent definition of the S&P 500 universe throughout the sample period, we use constituents as of February 2024. Pre-processed 10-K filings are obtained from the Notre Dame Software Repository1 (accessed: May 18, 2025). After matching filings to the S&P 500 constituents using Central Index Keys (CIK), the sample includes 5,250 documents.

Building upon the one-stage parsed text files from the repository, we apply additional preprocessing steps to optimize the corpus for topic modeling. First, paragraphs shorter than 80 characters or containing over 50% non-alphabetic characters are removed. URLs and emails are eliminated using regular expressions, and the text is tokenized into individual words with English stopwords filtered out. Words shorter than four characters and domain-specific terms such as ‘fiscal’ or ‘quarterly’ are also removed. A document-term matrix (DTM) is then constructed and dimensionality reduced by excluding terms appearing in fewer than 0.1% or more than 80% of documents (Dyer et al. 2017). This approach preserves terms with potential topical significance while minimizing noise, forming the foundation for subsequent LDA topic modeling. After text preprocessing, we obtain a document-term matrix of 5,241 10-K filings and 87,106 unique terms.

Then, we extract specific sections from 10-K filings including: Management’s Discussion and Analysis (MD&A, typically Item 7), Risk Factors (Item 1A), Controls and Procedures (Item 9A), Changes in Accounting and Disagreements (Item 9), Unresolved Staff Comments (Item 1B), and Quantitative and Qualitative Disclosures (Item 7A). The section extraction algorithm employs a multi-step approach. It first scans the document for standard SEC item headers using pattern recognition to identify section boundaries. Next, it extracts the text content between identified headers, creating separate text segments for each section. Finally, the same preprocessing procedures used for the main corpus (see previous paragraph) are applied to each section, ensuring consistency in tokenization, stop-word removal, and term filtering.

We obtain annual financial data from Compustat for 2009–2024, a timeframe necessary to calculate earnings trend components while remaining consistent with the 10-K sample. To construct the dependent variable for earnings prediction, we follow Chen et al. (2022) and focus on earnings per share excluding extraordinary items (EPSPX). We then define a binary indicator equal to 1 if a firm’s EPS increases and 0 if it decreases from year t to year t + 1. The EPS change is detrended by subtracting a drift term, calculated as the average EPS change over the previous four years to control for firm-specific growth trends and enable meaningful cross-sectional comparisons.

We calculate a set of 21 financial variables identified in prior literature as significant predictors of earnings changes. These variables span five key dimensions of financial performance: profitability, liquidity, leverage, efficiency, and market valuation. Ou and Penman (1989) evaluate 68 financial statement items and identify 16 indicators that collectively predict earnings direction more accurately than price-based measures alone. Sloan (1996) highlights the differential persistence of earnings components, showing that firms with high accrual components typically experience subsequent earnings reversals, while cash flow components persist more reliably. This explains why we include both accrual-based and operating cash flow measures to capture the varying information content of earnings components. Piotroski (2000) develops the F-Score using nine fundamental signals across profitability (ROA, CFO, change in ROA, accruals), leverage/liquidity (change in debt ratio, change in current ratio, equity offerings), and operating efficiency (change in gross margin, change in asset turnover). Instead of calculating the composite F-Score, we include its component ratios individually to allow differential weighting in the predictive model.

For a subset of variables where a missing value corresponds to the absence of the underlying activity, we replace missing entries with zero. For instance, this applies to inventory (INVT), capital expenditures (CAPX), or depreciation and amortization (DP). For other variables, missing values are imputed using industry averages based on the Fama-French 48 industry classification, constructed from each firm’s SIC code (see Chen and McCoy (2024)). Of all firm-year observations, 1,759 require imputation using the industry mean, with 1,322 cases due to missing inventory turnover. While this accounts for about 33% of observations, it affects only 4% of the underlying data entries in the dataset. Table 1 presents the definitions and summary statistics of our financial variables.

Table 1.

Financial variables.

Variable Definition Mean Std. Dev. Min. Max.
Profitability variables
ROA Net income / Total assets 0.066 0.069 −0.161 0.282
ROE Net income / Common equity 0.169 0.579 −2.880 3.324
EBIT_Margin EBIT / Sales 0.192 0.136 −0.261 0.569
Gross_Margin (Sales − COGS) / Sales 0.447 0.223 0.033 0.949
CFO_to_Assets Operating cash flow / Total assets 0.106 0.070 −0.056 0.321
Accruals (Net income – Operating cash flow) / Total assets −0.040 0.048 −0.221 0.117
Liquidity variables
Current_Ratio Current assets / Current liabilities 1.631 1.007 0.379 6.013
Quick_Ratio (Current assets − Inventory) / Current liabilities 1.333 0.891 0.180 5.413
Cash_Ratio Cash & equivalents / Current liabilities 0.675 0.783 0.009 4.558
Leverage variables
Debt_to_Equity Total liabilities / Common Equity 2.484 6.028 −28.976 28.903
Debt_to_Assets Total liabilities / Total assets 0.644 0.214 0.144 1.272
Long_Term_Debt_to_Assets Long-term debt / Total assets 0.276 0.185 0.000 0.951
Efficiency variables
Asset_Turnover Sales / Total assets 0.697 0.608 0.040 3.342
Inventory_Turnover COGS / Inventory 411.560 953.578 0.253 2732.488
Receivables_Turnover Sales / Accounts receivable 12.075 20.744 0.073 133.127
CAPEX_to_Assets Capital expenditures / Total assets 0.034 0.034 0.000 0.167
Market valuation variables
Book_Value 14,759.054 27,454.617 −4822.789 182,828.230
Market_Cap 49,179.189 75,294.089 2647.402 480,021.371
Market_to_Book Market capitalization / Book Value 4.881 14.202 −73.669 74.239
Price_to_Earnings Stock price / Earnings per share 25.439 46.903 −176.452 286.697
Dividend_Yield Common dividends / Market capitalization 0.018 0.016 0.000 0.073

3.2. Topic modeling on 10-K filings

To identify the main themes discussed in the 10-K filings, we apply LDA, a widely used unsupervised machine-learning technique (Blei et al. 2003). LDA treats each filing as a mixture of underlying topics and each topic as a collection of words that tend to appear together. The key decision in LDA is choosing the number of topics (K). Too few topics merge distinct ideas into overly broad categories; too many create specific or hard-to-interpret themes.

We evaluate five candidate values K ∈ {20, 60, 100, 140, 180} covering a range from coarse to highly granular topic structures, using two complementary criteria. Perplexity provides a measure of model fit, with lower values indicating better out-of-sample performance, but it does not guarantee that topics are meaningful. As shown by Mimno et al. (2011), models with strong perplexity can still produce semantically incoherent topics. To address this limitation, we complement perplexity with topic coherence (Röder et al. 2015), which evaluates how consistently a topic’s top words co-occur in the corpus and better reflects human interpretability. As shown in Figure 1, perplexity decreases sharply at lower values of K and then levels off between 140 and 180. Coherence scores, however, continue to rise over this interval, indicating greater interpretability without clear signs of overfitting. We therefore select K = 140 as the optimal number of topics. This choice aligns with prior work, including Dyer et al. (2017), who use 150 topics.

Figure 1.

Selection of optimal K. Notes: This figure presents the perplexity and coherence scores for different numbers of topics. Perplexity measures the model’s predictive performance as follows:Perplexity(D)=exp{d=1Mlogp(wd)d=1MNd} where p(wd) denotes the document likelihood under the fitted model and Nd is the number of words in each document. We train the model on 75% of the data and then calculate the perplexity using a random hold-out sample of the remaining 25% of the observations. The coherence score is computed as follows:Cv(T)=1Ni=2Nj=1i1logP(wi,wj)+ϵP(wi)P(wj) where P(wi, wj) denotes the probability of terms co-occurring within a sliding window (default = 110 words) in the 10-K corpus, and 𝜖 is a smoothing constant. In a second stage, we narrow the search to K ∈ {140, 150, 160, 170, 180} to capture potential marginal gains. Untabulated results highlight that coherence increases further but only modestly. After fixing the number of topics to 140, we estimate the final LDA model using Gibbs sampling with 1,000 iterations and a burn-in of 250 to ensure convergence. Hyperparameters α and β follow standard symmetric values of 50/K and 0.1, respectively.

The estimation yields two outputs: (1) the document-topic distribution matrix, which provides each 10-K’s probability distribution over the 140 topics and serves as textual features in the earnings prediction models, and (2) the top words for each topic, enabling qualitative interpretation. At this stage, we exclude filings with fewer than 100 characters after preprocessing. The final dataset comprises 5,155 firm-year observations including 495 companies over the period 2013–2023.

3.3. Machine learning predictive models

3.3.1. Dataset structure

Our dependent variable is a binary indicator constructed following Chen et al. (2022). We examine the direction of earnings changes after adjusting for firm-specific trends. Specifically, we subtract the average EPS change over the prior four years (the drift term) from the current change. An earnings increase is then coded as 1 and a decrease as 0 based on this de-trended measure. This adjustment serves three purposes (Chen et al. 2022): (1) it reduces class imbalance (as raw earnings increases are more common), (2) makes the prediction task more relevant for investment decisions by removing anticipated changes, and (3) allows direct comparison with prior literature. After this procedure, the sample is nearly balanced with 2,719 observations (52.7%) indicating an earnings increase and 2,436 observations (47.3%) indicating a decrease. This distribution mitigates concerns about class imbalance, which could otherwise bias predictive model performance, and supports the use of standard machine learning algorithms and evaluation metrics.

The predictive modeling framework uses a time-based data split consistent with the panel structure of the dataset, where firms appear across multiple years. The training set covers 2013–2021 (4,212 observations), and the test set covers 2022–2023 (943 observations). This temporal split serves three methodological purposes:

  1. it prevents data leakage ensuring that information from future periods does not influence predictions;
  2. it reflects real-world investment settings where forecasts rely on historical data; and
  3. it provides a meaningful gap to evaluate the model’s ability to generalize beyond the training period.

To evaluate the incremental value of textual information, we estimate two models for each machine learning algorithm. The baseline model uses only the 21 financial variables presented in Table 1, while the enhanced model combines these variables with the 140 topic probability distributions generated by the LDA analysis of 10-K filings.

3.3.2. Algorithms and parameter optimization

We employ three machine learning prediction models. First, we employ a Lasso regression (Tibshirani 1996). It is particularly useful in high-dimensional settings because it automatically performs variable selection and shrinks less important coefficients toward zero, which helps prevent overfitting. To select the optimal regularization parameter λ, we conduct time series cross validation that respects chronological order of the data to ensure that future observations do not predict the past – a crucial requirement in financial predictions (Bergmeir et al. 2016). We use five folds and perform tuning separately for the baseline and enhanced models. We select λ using the ‘1 standard error rule’ which chooses the most regularized models within one standard error of the minimum cross-validated AUC, yielding more parsimonious and robust models (Hastie et al. 2009). The selected λ values are 0.013 for the baseline model and 0.031 for the enhanced model.

Second, we employ a Random Forest algorithm (Breiman 2001), an ensemble learning method that builds multiple decision trees and combines their predictions through majority voting. It is particularly suitable for earnings prediction because it captures complex nonlinear relationships and interactions among financial variables without requiring explicit specification of these relationships (Chen et al. 2022).

Hyperparameter optimization focuses on two parameters: the number of trees (ntree) and the number of features considered at each split (mtry). Larger ntree improves performance at the cost of computational time, while mtry balances individual tree strength against inter-tree correlation. Following standard practice (Hastie et al. 2009), we conduct a grid search to identify optimal values for both parameters. Consistent with Lasso, we employ time series cross validation rather than random fold assignment. For the baseline model, the grid includes ntree ∈ {100, 200, 500, 1000} and mtry ∈ {2, 4, 6, 8} allowing exploration from highly correlated trees (large mtry) to diverse but weaker trees (small mtry). The enhanced model uses an expanded mtry ∈ {6, 8, 10, 12, 14} to accommodate the larger set of 161 features. Hyperparameter tuning selects ntree = 1000 and mtry = 6 for the baseline model, indicating that deeper trees and larger ensemble improve predictions. For the enhanced model, the optimal parameters are ntree = 500 and mtry = 14.

Third, we employ XGBoost (Extreme Gradient Boosting) (Chen and Guestrin 2016). Like Random Forest, XGBoost is tree-based, but it builds trees sequentially, with each tree learning to correct the residual errors of its predecessors. This boosting approach iteratively improves predictions and often achieves superior performance on structured data. We tune three key parameters: (1) the number of boosting rounds (nrounds) controls ensemble size, with more rounds improving accuracy but increasing overfitting risk, (2) the maximum tree depth (max_depth) sets tree complexity, balancing the capture of nonlinear patterns and generalization and (3) the learning rate (eta) scales each tree’s contribution, with smaller values requiring more trees but improving performance. The hyperparameter search employs a grid search on those three parameters, setting subsample and colsample_bytree at 0.8 (i.e., two other parameters representing proportion of features considered for each tree) for computational efficiency. The grids are specified as maximum nrounds ∈ {100, 200, 300, 500}, max_depth ∈ {2, 4, 6, 8, 10, 12}, and eta ∈ {0.01, 0.03, 0.05, 0.10, 0.15, 0.20}. An early stopping mechanism with a patience of 10 rounds is applied to prevent overfitting, which dynamically determines the optimal number of trees rather than strictly sticking to the predefined discrete grid values. Consequently, the baseline (enhanced) model achieved optimal performance with an actual nrounds = 83 (122), max_depth = 4 (2), and eta = 0.05 (0.03). The reduced tree depth and learning rate for the enhanced model suggest that the textual features allow effective learning with simpler trees, likely due to richer feature representation from the LDA topic distributions.

3.3.3. Evaluation metrics

To evaluate the performance of the models in predicting the direction of earnings change, we use three widely adopted metrics: Area Under the Curve (AUC), Accuracy, and the F1 Score. Each metric highlights a different aspect of prediction quality and provides a balanced view of model performance.

AUC measures how effectively the prediction model separates firms with earnings increases from those with decreases. It is computed as the area under the Receiver Operating Characteristic (ROC) curve, which plots the true positive rate against the false positive rate across classification thresholds. It has the advantage of being independent of any specific probability cutoff and is widely used in prediction settings. Chen et al. (2022), for example, report AUC values between 67.52% and 68.66%.

Accuracy is the simplest and most intuitive metric. It shows the percentage of correct predictions (i.e., how often the model correctly predicted the direction of the earnings change):

Accuracy =TP+TNTP+TN+FP+FN

where TP and TN denote true positives and true negatives, respectively, and FP and FN denote false positives and false negatives.

The F1 Score combines precision (how many of the firms predicted to increase earnings actually did) and recall (how many of the firms that actually increased earnings were correctly identified by the model):

F1 Score =2× Precision × Recall Precision + Recall

It offers a balanced evaluation of performance when false positives and false negatives carry similar importance. It is also particularly useful when class distributions are imbalanced, although in our nearly balanced sample it serves as a complementary measure.

4. Results

4.1. Predictive power of narratives in 10-K filings

The empirical results provide mixed evidence on the incremental predictive power of narratives from 10-K filings for predicting earnings changes. Table 2 presents the performance metrics for all six model configurations. XGBoost outperforms both Lasso regression and Random Forest, with the baseline model achieving the highest performance across all metrics: AUC = 0.750, Accuracy = 0.695, and F1 Score = 0.699. This finding is consistent with recent machine learning literature emphasizing the superior performance of gradient boosting methods for structured prediction tasks (Chen et al. 2022).

Table 2.

Predictive performance of models on the test data.

Lasso Random Forest XGBoost
Baseline Enhanced Baseline Enhanced Baseline Enhanced
AUC 0.676 0.654 0.733 0.714 0.750 0.737
Accuracy 0.623 0.592 0.667 0.664 0.695 0.672
F1 Score 0.651 0.652 0.667 0.674 0.699 0.695

Contrary to expectations, the enhanced models, incorporating 140 LDA-derived topic features alongside financial variables, generally exhibit slightly lower performance than the baseline models using only financial variables. This pattern is consistent across AUC and Accuracy metrics, where enhanced models underperform baseline models. Specifically, the AUC (Accuracy) declines from 0.676 to 0.654 (0.623 to 0.592) for Lasso, from 0.733 to 0.714 (0.667 to 0.664) for Random Forest and from 0.750 to 0.737 (0.695 to 0.672) for XGBoost. The F1 Score metric presents a more nuanced picture, showing minor improvements for Lasso and Random Forest enhanced models (0.651 to 0.652 and 0.667 to 0.674, respectively). Overall, these findings suggest that the relationship between narratives and earnings changes is complex. While incorporating LDA-derived topic features can yield modest gains on a specific metric, using 140 topics does not provide systematic incremental predictive power beyond traditional financial variables.

We then conduct two untabulated robustness tests. First, we examine whether including COVID-19 years influences model performance. We re-estimate the models using 2013–2019 as the training period and 2023 as the test period, thereby excluding 2020–2022. The results are consistent with the main findings, with baseline models outperforming their enhanced counterparts across all three algorithms. Second, we investigate whether the high dimensionality of LDA explains the lack of incremental predictive value. Reducing the number of topics from 140 to 100 yields mixed results. Random Forest exhibits improved prediction with 100 topics, whereas XGBoost continues to underperform, and Lasso remains largely unchanged.

4.2. Topics in 10-K filings predictive of earnings changes

We examine feature importance when using the enhanced XGBoost model. Figure 2 reports the top 20 features. Traditional financial ratios – particularly price-to-earnings ratio, ROA, accruals, and gross margin – dominate the ranking, while the most important topic features exhibit low importance. This result is consistent with prior earnings-prediction research (Ou and Penman 1989; Chen et al. 2022) and highlights the continued relevance of structured financial information relative to narrative disclosures.

Figure 2.

XGBoost – Top 20 most important features. Notes: The different scales on the horizontal axis reflect distinct feature importance. XGBoost uses relative gain (0–1 scale).

To evaluate the contribution of narratives, we further inspect the top 40 topic features from the XGBoost model with 140 topics in Figure 3. Since LDA provides word distributions rather than semantic labels, topic interpretation relies on qualitative assessment of each topic’s most important words. The examination of topic composition reveals several meaningful patterns that contribute to the prediction of earnings changes.

Figure 3.

XGBoost – Top 40 most important topic features. Notes: The different scales on the horizontal axis reflects distinct feature importance. XGBoost uses relative gain (0–1 scale).

A first pattern is that many influential topics capture industry-specific disclosure language. Examples include topics related to healthcare and pharmaceuticals (Topics 118, 36, 70, 61, 108), energy and utilities (Topics 62, 87, 93, 40, 111), and technology-related themes (Topics 81, 7, 127). For instance, words in Topic 118 include ‘fertility’, ‘kidney’, ‘capitation’, and ‘practitioner’ and words in Topic 62 include ‘liquefaction’, ‘compressor’, and ‘energy’. This suggests that industry-specific linguistic patterns may embed information relevant for predicting earnings changes. Within these industry-oriented patterns, some topics further capture detailed elements of firms’ business models and operations. For instance, Topic 78 (Retail & Product Assortment) includes keywords such as ‘assortment’, ‘markdown’, and ‘ecommerce’, pointing to retail merchandising and store-format adjustments. Topic 44 (Real Estate & Land Development) features ‘tract’, ‘propco’, and ‘plat’, reflecting land development and property-holding structures. Topic 87 (Renewable Energy Technology) is defined by technical terms such as ‘microinverter’, ‘inverter’, and ‘connector’, capturing solar-energy hardware disclosures. These operational themes suggest that LDA extracts variation in firms’ activities that is not fully captured by financial variables. Second, beyond industry-specific themes, several topics reflect functional aspects of complex financial and governance practices. For instance, Topic 106 (Financial Instruments) encompasses specialized banking and regulatory terminology, including words such as ‘noncumulative’ (non-cumulative preferred shares), ‘tlac’ (Total Loss-Absorbing Capacity requirements), and ‘hqla’ (High-Quality Liquid Assets). Topic 49 (Corporate Governance) contains procedural language related to shareholder meetings and securities administration, such as ‘adjourn’, ‘securityholder’ and ‘CUSIP’.

4.3. Sections in 10-K filings predictive of earnings changes

We conduct a section-level analysis to examine how the most important topics are distributed across different sections of the 10-K filings. After identifying each section, we compute the average term frequency for every section type across all filings in the sample. To assess the predictive relevance of each section, we compute cosine similarity scores between the 40 most important topics identified through XGBoost and the term distributions of each 10-K section. This approach allows us to map topics to their most likely source sections without re-estimating separate topic models for individual sections, thereby maintaining consistency with the main analysis.

Figure 4 presents the distribution of important topics across 10-K sections. The results reveal that Management’s Discussion and Analysis contains the highest concentration of predictive topics, capturing more than 25 of the 40 most important topics. The next most influential sections are Risk Factors (7 topics) and Controls and Procedures (4 topics).

Figure 4.

Number of important topics associated with each 10-K section. Notes: The horizontal axis reflects the number of important topics, obtained with XGBoost, associated with each 10-K section.

5. Discussion and conclusion

We extend recent earnings prediction research (Chen et al. 2022) by incorporating full- document topic distributions extracted using LDA as explanatory variables in machine learning models. Overall, we find that incorporating topic features derived from LDA analysis of 10-K filings offers limited incremental predictive value beyond traditional financial variables. Across several machine learning models, including Lasso, Random Forest, and XGBoost, adding the topic features does not substantially improve performance and frequently even reduces it. Regarding our three research questions, we find that: (1) narrative disclosures in 10-K filings offer only modest and inconsistent incremental power for predicting earnings changes; (2) certain topics, particularly those capturing industry-specific operations and business model details, provide stronger predictive power than others; and (3) predictive content is concentrated in specific sections, MD&A and Risk Factors contributing the most, albeit modestly.

5.1. Implications for practitioners

For practitioners, the implications are threefold:

  1. The targeted use of narrative disclosures may still provide qualitative value. The higher relevance of certain sections suggests that analysts should focus selectively on these portions rather than ingesting entire filings. Such prioritization can enhance efficiency while preserving the most decision-relevant insights.
  2. Our findings highlight important caveats regarding automation. Automated textual features can introduce noise and methodological sensitivity – for example, through choices related to topic modeling parameters or preprocessing techniques. Practitioners should rigorously validate any text-augmented model using out-of-sample testing and benchmark its performance against traditional approaches using economically meaningful metrics. This caution extends to the broader adoption of modern AI systems, which often function as black boxes and require careful scrutiny before integration into decision-making processes.
  3. Corporate narratives frequently echo quantitative disclosures or contain low-information content. As a result, for most investment and valuation workflows, the marginal benefit of systematically incorporating narrative data appears limited. In many cases, these benefits may not justify the additional complexity, implementation challenges, and data-processing costs involved.

Conceptually, the results suggest an information redundancy interpretation: narrative disclosures in 10-K filings may largely echo information already incorporated in traditional financial metrics. This interpretation is consistent with prior evidence that narrative disclosures often contain boilerplate and low-information content (Dyer et al. 2017), and with the idea that management may strategically obfuscate poor performance by increasing textual complexity (Li 2008). Predictive information appears to be localized in specific sections, and whole-document analysis may dilute these concentrated signals, consistent with the evidence in Campbell et al. (2014). From a practical perspective, our findings provide regulators, such as the SEC, with empirical guidance on which sections of the 10-K might convey information most relevant for investors, potentially motivating refinements to disclosure standards that discourage excessive boilerplate while reinforcing economically meaningful content.

At the same time, the mixed predictive results highlight the instability and limited robustness of textual features in earnings prediction in this context. Whether these results generalize to non-US settings remains an open question. On one hand, the structured nature of US disclosures makes narrative information easier to compare and analyze. On the other hand, less structured reporting in other settings may provide firms with greater opportunities to disclose unique, non-boilerplate information. Model performance is sensitive to methodological choices, particularly the number of topics, reflecting well-established concerns over topic-modeling instability and the challenges of extracting reliable signals from high-dimensional text (Greene et al. 2014; Lewis and Young 2019). While dimension reduction techniques can improve performance for specific algorithms, the overall pattern suggests that the incremental value of LDA-derived features remains modest. Alternatively, the limited incremental value of textual features may reflect a temporal mismatch: narratives could be more informative for multi-year horizons rather than immediate annual earnings changes. Further research could disentangle between the two explanations.

5.2. Limitations

Our study contains several limitations. First, the limited predictive value of LDA-derived topics partly reflects methodological constraints inherent in standard unsupervised topic modeling. The bag-of-words representation disregards word order and syntax, leading to information loss when converting qualitative text into quantitative features (Loughran and McDonald 2016). LDA assumes that each document is a mixture of latent topics optimized for textual coherence rather than predictive accuracy (Taddy 2013), which may cause misalignment with financially relevant information. Future work could explore transformer-based models (e.g., BERT or FinBERT) to better capture the semantic and rhetorical structure of narrative disclosures. Second, many extracted topics resemble industry labels dominated by sector-specific jargon, suggesting that LDA predominantly captures broad industry traits rather than firm-level signals. This dilution is especially pronounced in cross-industry samples, highlighting the potential value of industry-specific topic modeling. Third, topic modeling remains highly sensitive to the number of topics. Despite using grid search and perplexity-coherence metrics, topic interpretability and predictive performance vary substantially across specifications, illustrating trade-offs between interpretability, granularity, and model fit. These limitations suggest that more nuanced approaches, such as supervised LDA (Gentzkow et al. 2019), embedding-based representations (Huang et al. 2023), or stability-enhancing methods (Greene et al. 2014), may better capture the economic substance of narrative disclosures.

Zixian Yin is an alumnus of the MSc Business Analytics & Management, Rotterdam School of Management, Erasmus University Rotterdam.Zixian Yin is one of the winners of the MAB Thesis Award 2025. This article is based on her master thesis.

Dr. A. Madelaine – Alexandre is an Assistant Professor of Accounting, Rotterdam School of Management, Erasmus University Rotterdam.

Note

References

  • Blei DM, Ng AY, Jordan MI (2003) Latent dirichlet allocation. Journal of Machine Learning Research 3: 993–1022. https://dl.acm.org/doi/10.5555/944919.944937
  • Brown NC, Crowley RM, Elliott WB (2020) What are you saying? Using topic to detect financial misreporting. Journal of Accounting Research 58(1): 237–291. https://doi.org/10.1111/1475-679X.12294
  • Campbell JL, Chen H, Dhaliwal DS, Lu H, Steele LB (2014) The information content of mandatory risk factor disclosures in corporate filings. Review of Accounting Studies 19(1): 396–455. https://doi.org/10.1007/s11142-013-9258-3
  • Chen T, Guestrin C (2016) XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785
  • Chen X, Cho YH, Dou Y, Lev B (2022) Predicting future earnings changes using machine learning and detailed financial data. Journal of Accounting Research 60(2): 467–515. https://doi.org/10.1111/1475-679X.12429
  • Dyer T, Lang M, Stice-Lawrence L (2017) The evolution of 10-K textual disclosure: Evidence from Latent Dirichlet Allocation. Journal of Accounting and Economics 64(2–3): 221–245. https://doi.org/10.1016/j.jacceco.2017.07.002
  • Feuerriegel S, Ratku A, Neumann D (2016) Analysis of how underlying topics in financial news affect stock prices using Latent Dirichlet Allocation. Proceedings of the 2016 49th Hawaii International Conference on System Sciences (HICSS), 1072–1081. https://doi.org/10.1109/HICSS.2016.137
  • Greene D, O’Callaghan D, Cunningham P (2014) How many topics? Stability analysis for topic models. In: Calders T, Esposito F, Hüllermeier E, Meo R (Eds) Machine Learning and Knowledge Discovery in Databases 8724: 498–513. https://doi.org/10.1007/978-3-662-44848-9_32
  • Huang AH, Zang AY, Zheng R (2014) Evidence on the information content of text in analyst reports. The Accounting Review 89(6): 2151–2180. https://doi.org/10.2308/accr-50833
  • Huang AH, Lehavy R, Zang AY, Zheng R (2018) Analyst information discovery and interpretation roles: A topic modeling approach. Management Science 64(6): 2833–2855. https://doi.org/10.1287/mnsc.2017.2751
  • Huang AH, Wang H, Yang Y (2023) FinBERT: A large language model for extracting information from financial text. Contemporary Accounting Research 40(2): 806–841. https://doi.org/10.1111/1911-3846.12832
  • Li F (2010) Textual analysis of corporate disclosures: A survey of the literature. Journal of Accounting Literature 29(1): 143–165.
  • Mimno D, Wallach H, Talley E, Leenders M, McCallum A (2011) Optimizing semantic coherence in topic models. Proceedings of the Conference on Empirical Methods in Natural Language Processing, 262–272. https://dl.acm.org/doi/10.5555/2145432.2145462
  • Penman SH (2013) Financial statement analysis and security valuation. McGraw-Hill Education, 5th edition.
  • Piotroski JD (2000) Value investing: The use of historical financial statement information to separate winners from losers. Journal of Accounting Research 38: 1–41. https://doi.org/10.2307/2672906
  • Röder M, Both A, Hinneburg A (2015) Exploring the space of topic coherence measures. Proceedings of the Eighth ACM International Conference on Web Search and Data Mining, 399–408. https://doi.org/10.1145/2684822.2685324
login to comment