When AI Summaries Change the Investment Call

When AI Summaries Change the Investment Call

When AI Summaries Change the Investment Call

Understanding

by

Yongjae Lee

Jacob Chanyeol Choi

An LLM-generated summary can keep every sentence true and still reverse the call. Our EMNLP 2026 Industry Track paper measures that failure and tests a source-grounded way to reduce it.

When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis

Authors

Hoyoung Lee¹・², Suhwan Park¹, Seunghan Lee³, Jun Seo³, Jaehoon Lee³, Sungdong Yoo³, Minjae Kim³, CheolWon Na⁴, Zhangyang Wang⁵, V. Zach Golkhou⁶, Minkyu Kim⁷, Sotirios Sabanis⁸・⁹・¹⁰, Alejandro Lopez-Lira¹¹, Dhagash Mehta¹², Soonyoung Lee³, Chanyeol Choi², Wonbin Ahn³・*, and Yongjae Lee¹・²・*

Affiliations

¹ UNIST · ² LinqAlpha · ³ LG AI Research · ⁴ Sungkyunkwan University · ⁵ University of Texas at Austin · ⁶ J.P. Morgan Chase · ⁷ State Street Investment Management · ⁸ University of Edinburgh · ⁹ National Technical University of Athens · ¹⁰ Archimedes/Athena Research Centre · ¹¹ University of Florida · ¹² BlackRock · * Corresponding authors

Venue

2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026) · Industry Track

Presented

October 24–29, 2026 · Budapest, Hungary

arXiv ↗

What we found

  • Compression changed the call. When we asked four different LLMs to summarize the same documents, decision-flip rates reached 23.7–31.5% for MD&A and 24.9–29.5% for earnings calls.

  • Context was lost during summarization. When LLMs summarized the source documents, they omitted caveats, comparisons, timing, and offsets more often than headline facts.

  • Source checks helped. Comparing summaries, checking their differences against the source, and selecting the better-supported summary reduced decision flips by 38.1% for MD&A and 20.1% for earnings calls versus naive summarization, while retaining roughly 95% token savings.


Financial filings and earnings-call transcripts are long, and reading every document in full takes time. Asking AI for a summary is a natural shortcut, but that shortcut can change the investment judgment. Imagine a filing that reports strong growth, then explains that the increase reflects an unusually weak comparison period. If the summary keeps the growth figure but drops that qualifier, every retained sentence can be true while the company looks more attractive than the full report suggests.

Conceptual illustration; the 7/4/3 evidence counts are illustrative, not measured results. Compression preserves headline facts but thins out context, caveats, offsets, and comparisons. The resulting 20-bullet context can remain factual while shifting a source-supported bearish decision to bullish.

An LLM summary can be factually correct at the sentence level and still be decisionally unfaithful at the document level.

That is the gap between factuality and information fidelity. Factuality asks whether the generated summary says something true. Information fidelity asks whether it still supports the same decision as the source.

How often the decision moves

We tested 300 10-Q MD&A sections and 297 earnings-call transcripts from S&P 100 firms over fiscal 2025 Q1–Q3, comparing investment judgments based on each full source with those based on its summary.

Summaries from all four LLMs change investment decisions much more often than rereading the original documents. The red dashed lines mark the reread floor: how often the decision changes when the model reads the same full source again, without any summarization (7.0% for MD&A and 5.7% for earnings calls).

Across the four summary models, 23.7% to 31.5% of MD&A decisions and 24.9% to 29.5% of earnings-call decisions flipped. The control is the same decision model reading the same full source a second time, with nothing summarized in between. That flipped only 7.0% of MD&A decisions and 5.7% of earnings-call decisions. Summarization moves the call roughly three to five times as often as the model's own run-to-run instability. The underlying probabilities move on the same scale, which the figure above reports.

* Decision flip means the model’s top label (bear, neutral, or bull) changes when it reads the summary instead of the full source. Total variation distance (TVD) measures how much the probabilities across all three labels change.

Close to 30% of the time, reading a summary instead of the source changed the call. Rereading the source itself changed it in under one case in ten.

The distortion has a direction

A flip rate says how often the top label changes. It does not say whether the LLM summary pushes the decision bearish or bullish, or whether the same stocks move under every compressor.

The scatter plots below show each disclosure by signed decision shift and movement magnitude. If compression were model-invariant, the four panels would look alike. They do not.

Results pool MD&A sections and earnings-call transcripts, with each panel showing one compressor. The horizontal axis shows direction from bearish to bullish; the vertical axis shows movement magnitude as TVD.

The labeled examples for major companies reveal both shared patterns and model differences. MSFT appears on the bearish side in all four panels, while AAPL and AMZN appear on the bullish side, although the magnitude varies. GOOG appears bearish under GPT and Qwen but bullish under Gemini and DeepSeek. NVDA appears bearish under GPT, Qwen, and DeepSeek but slightly bullish under Gemini. These are shifts relative to the full source, not absolute bearish or bullish recommendations. Even for widely followed companies, the summary model can affect both the direction and size of the distortion.

What is consistent is that the decision moves. Which way it moves is not.

What gets lost

The missing information is usually not the headline fact. It is the caveat, comparison, timing, causal driver, or offset that tells an investor what the fact means.

Context made up 25% of the MD&A source inventory but only 9% of a naive LLM summary. In flipped cases, adding the omitted context back recovered the source decision 37% of the time, more than adding boilerplate or random facts.

Naive MD&A compression disproportionately removes contextual qualifiers. Adding that context back recovers more source-supported decisions than adding the same amount of boilerplate or random information.

Source-grounded summarization reduces decision distortion

Different compressors agreed on the top decision only about 75% of the time. ACC turns that disagreement into a source-checking signal.

Not every compressor is naive about this. Contextualization works under the same fixed budget but changes what counts as worth keeping. Instead of picking the facts that look most important on their own, it asks the compressor to preserve whatever makes each material point interpretable: the relevant comparison, caveat, offsetting signal, or qualifier. The mechanism is unchanged. Only the selection criterion moves, from isolated salience to interpretable evidence. That change alone brings flips down to 21.2% for MD&A and 22.1% for earnings calls.

Agentic Context Compression (ACC) generates two contextualized candidates with GPT-5.4-Mini and DeepSeek-V4-Flash, finds their material disagreements, checks them against the source, and selects the better-supported candidate without rewriting it.

ACC separates generation from adjudication, investigates disagreement against the source, and returns one candidate intact. The lower panel shows decision-flip performance under the same 20-bullet budget.

ACC brought flips down to 19.5% for MD&A and 19.9% for earnings calls. Against naive compression the reduction holds for both document types under one judge model, and for MD&A under the other. For earnings calls under the second judge it did not clear the bar. Neither did the gap between ACC and contextualization on its own, under either judge. ACC is a safeguard, not a complete solution.

The effect survives in production

We applied ACC inside a deployed equity-forecasting product used by institutional investors, on S&P 100 disclosures from fiscal 2025 Q1 to Q3. Absolute product metrics stay confidential, but the direction is unambiguous. Naive compression degraded forecasting IC: the decision changes it introduced took forecast-relevant signal out with them. ACC went the other way. It improved IC by 23.8% against naive compression and by 8.3% against the original-source baseline, while cutting source-relative decision flips by 59.5%.

Information coefficient measures how well a signal's ranking lines up with what actually happened. It is the one result in this study judged against something other than agreement with a document.

Both methods saved more than 90% of tokens. The benchmark's 20-bullet output runs 4% to 5% of the source. Different budgets, same order of magnitude.

The benchmark and the case study are different settings. The benchmark is a fixed 20-bullet budget over S&P 100 filings and calls, scored by judge models. The case study is a production system. The numbers are not directly comparable. They point the same way.

What to do

The research points at a workflow, not a warning.

Ask two different models to summarize the same document. Read the summaries against each other and mark where they disagree: a figure one keeps and the other drops, a caveat that survives in one and not the other, a different read on the same quarter. Go back to the source only on those points. Then pick the better-supported summary and use it as it stands.

The last step is the one that is easy to get wrong. The instinct after checking the source is to write a better third summary combining both. The paper tested that. Merging two summaries into one does help. Auditing where they disagree and then taking one candidate intact helps more, and costs less. A fresh synthesis is another compression, and it can introduce a new distortion after the checking is already done.

This is slower than asking once, and it is not the same thing as ACC, which runs across every document rather than the ones an analyst happens to find suspicious. What carries over is the principle. Disagreement between two summaries marks where the source needs rereading, and two compressors disagree about a quarter of the time.

Claim AI summaries are neutral.
Finding Even a factually accurate summary can distort the decision supported by the source.

Claim A good summary preserves the headline facts.
Finding Headline facts can be misinterpreted when the context that gives them meaning is lost.

Claim Summarization pushes investment judgments in a consistent direction.
Finding Both the direction and strength of distortion vary by stock and summary model. A summary can push one judgment more bullish and another more bearish.

Claim AI summaries are too risky to use for investment decisions.
Finding Comparing multiple summaries, checking their disagreements against the full source, and selecting the better-supported summary can substantially reduce distortion while keeping most of the token savings.

What this does not prove

Information fidelity asks whether an LLM summary preserves the source-implied decision. It does not say that the decision is economically correct. The benchmark covers S&P 100 disclosures and model-based judges; human and broader-market validation still matters.

REFERENCES
  • Lee, H., Park, S., Lee, S., et al. (2026). “When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis.” Accepted to the EMNLP 2026 Industry Track. arXiv:2606.29251 ↗.