Insights

Can AI Read Your Factsheet?

Written by Marc Quintavalle | Aug 19, 2026, 7:02:50 PM

Can AI Read Your Factsheet?

A hedge fund factsheet has traditionally been designed for one audience: the investor looking at it. That audience is changing. Investors can now use artificial intelligence to extract performance, compare strategies, identify inconsistencies and summarize fund materials before deciding what deserves closer attention.

This creates a new technical requirement for managers. Your factsheet must still look polished, but it must also work as a data source.

A document can appear perfectly clear on screen while becoming fragmented or misleading when processed by a machine. Text may be extracted in the wrong order. Table headings may become separated from their values. A chart may be recognized as an image but not as a reliable source of performance data. Footnotes may be assigned to the wrong metric—or missed entirely.

The question is no longer simply whether an investor can read your factsheet. It is whether the systems assisting that investor can read it accurately.

Marketing Alpha

A visually polished factsheet may earn attention. A machine-readable factsheet has a better chance of surviving analysis.


The PDF Is No Longer Just a Page

Most factsheets are distributed as PDFs because the format preserves fonts, colors, spacing and layout. But a PDF can be created in several technically different ways.

One may contain selectable text, structured tables and a defined reading order. Another may be little more than an image of a page. They can look nearly identical to a person while producing very different results when information is extracted.

The most common mistake is flattening the entire factsheet into an image. This often happens when a document is exported through a design platform, scanned, or processed through certain print-to-PDF workflows. Once flattened, the document must rely on optical character recognition to recover its content. OCR can be effective, but it introduces another opportunity for errors—particularly with decimal points, negative signs, superscripts and small footnotes.

Managers can perform an elementary check by opening the final PDF and attempting to select individual sentences and values. If the page behaves like one large image, the document is not truly text-based. If the text can be selected but appears in an illogical sequence when copied into a plain-text document, the factsheet has a reading-order problem.

Structure Matters More Than It Appears

AI document systems do not necessarily process a page in the same way a person does. They may extract text blocks, identify tables, analyze images and then attempt to reconstruct the relationship between those elements.

Complex layouts make that task harder. Multiple columns, floating text boxes, decorative rules and disconnected footnotes can cause content to be extracted out of sequence. A statistic shown beside a chart may be associated with the wrong chart, while a disclaimer running across the bottom of the page may be inserted into the middle of a paragraph.

A tagged PDF helps establish the document’s logical structure. Tags can identify titles, headings, paragraphs, tables, figures and the intended reading order. Adobe notes that automatic tagging can struggle with complex layouts, closely spaced columns and tables without borders, which means tagging should be checked rather than simply switched on during export. Adobe’s guidance on creating accessible PDFs provides a useful technical framework for reviewing these issues.

Accessibility and AI readability are not identical, but they share an important principle: information should remain understandable when its visual presentation is removed.

Performance Tables Need Real Relationships

A performance table is not simply a collection of numbers. Each number depends on its relationship to a row heading, column heading, date and unit of measurement.

Consider a table containing monthly returns. A person can usually infer that the value at the intersection of “March” and “2025” represents the fund’s return for March 2025. A document processor must correctly detect the table, preserve its rows and columns, recognize the headers and maintain that relationship during extraction.

Tables built from individually positioned text boxes may look correct but contain no underlying table structure. The values can be extracted as an unorganized stream of numbers. Merged cells, repeated header rows and extensive use of blank spacing can create additional ambiguity.

Where possible, use actual tables with clearly identified header cells. The W3C’s guidance for PDF tables recommends preserving row, column, header and data-cell relationships through proper tags such as Table, TR, TH and TD. Its PDF table technique explains how those relationships are represented programmatically.

Charts should also be supported by extractable data. A growth-of-$1,000 chart may communicate the trajectory of the fund effectively, but it should not be the only place where performance appears. AI systems may be able to interpret the image, but managers should not assume that every system can recover exact values reliably from lines, bars or axes.

The safest approach is to use charts for communication and tables for verification.

Marketing Alpha

Never make a machine estimate a value that you can state explicitly.


Every Metric Needs Context

Labels such as “Return,” “Volatility” and “Sharpe” may feel self-explanatory, but they leave important questions unanswered.

Is the return gross or net? Which share class does it represent? What currency is being used? Is volatility annualized? What risk-free rate was used to calculate the Sharpe ratio? Does “Since Inception” refer to the fund, the strategy, a predecessor vehicle or a simulated track record?

A person may find some of those answers in nearby footnotes. An automated system may not connect the footnote to the relevant number.

Factsheets should therefore use precise, consistent labels. Instead of presenting “Return: 8.4%,” consider “2025 Net Return — Class A, USD: 8.4%.” Instead of “Volatility,” use “Annualized Volatility — Monthly Net Returns, Since Fund Inception.”

The same principle applies to dates. “As of July” is less precise than “As of July 31, 2026.” “Since inception” should be accompanied by the actual inception date. Benchmarks should be named in full at first reference, and the same name should be used throughout the document.

These details do more than help AI. They reduce the amount of interpretation required from every investor reviewing the factsheet.

Version Control Is Part of Readability

Machine readability extends beyond the content of the page. Investors may receive several versions of the same document through email, a CRM, a consultant database or a virtual data room. If those versions have generic names such as Factsheet_Final.pdf or Factsheet_New2.pdf, both people and retrieval systems have difficulty determining which one is current.

Use a predictable naming convention that includes the fund and reporting date:

FundName_Factsheet_2026-07-31.pdf

The file’s internal title and metadata should match the document displayed on the page. The publication date should be clearly visible, and revised documents should be distinguishable from the originals. If a correction is material, changing the filename or adding a revision date is preferable to silently replacing the file.

This creates a simple hierarchy: one fund, one reporting period and one clearly identifiable version.

Run the Extraction Test

Before distributing a factsheet, test the document in an AI environment approved by your compliance and information-security teams. Do not upload confidential or nonpublic information to an unapproved consumer tool.

Ask the system to return:

  1. The fund name, strategy, share class, currency and reporting date.
  2. Monthly and annual returns in a structured table.
  3. Each risk statistic together with its measurement period.
  4. The benchmark and every comparison made against it.
  5. All qualifications, footnotes and disclosures attached to performance.

Then compare the response with the original factsheet.

The objective is not to see whether the AI produces an elegant summary. It is to identify omissions, incorrect associations and ambiguous labels. If the system assigns a footnote to the wrong statistic, reverses two table columns or cannot determine whether performance is gross or net, the factsheet needs to be revised.

Testing should also include copying the full document into a plain-text file and reviewing the extraction order. For a more formal review, Adobe Acrobat’s accessibility tools can identify missing tags, reading-order problems and improperly structured tables.

Design for Two Readers

None of this means managers should abandon good design. Investors still respond to clarity, hierarchy and visual polish. The goal is to create a document that works in two forms: as a designed page and as structured information.

Use real text rather than images of text. Build genuine tables instead of manually arranging numbers. Give metrics complete labels. Support charts with underlying values. Connect footnotes clearly to the information they qualify. Apply a logical reading order and test the exported PDF—not just the source document.

The strongest factsheet is no longer merely easy to look at. It is difficult to misinterpret, regardless of whether the first reader is an allocator, an analyst or an algorithm.

Marketing Alpha

Design your factsheet for investors—and structure it for the systems helping them decide.


Put Your Factsheet to the Test

At Thalēs, we help hedge fund managers turn complex investment strategies into clear, institutional-quality marketing systems. That work extends beyond the appearance of an individual document. It connects the substance of a manager’s investment program with how that program is presented across factsheets, pitch books, websites, virtual data rooms and the broader investor due diligence experience.

Machine readability is the latest layer of that work. Your investment story should remain accurate and persuasive regardless of whether it is first reviewed by an allocator, an analyst or an AI platform working on their behalf.

If you are unsure how your materials perform under that scrutiny, Thalēs can review them from both a human and machine perspective—identifying extraction problems, ambiguous metrics, structural weaknesses and opportunities to make your story easier to evaluate.

Contact us to put your current marketing materials through an investor-readiness test



The views expressed above are not necessarily the views of Thalēs Trading Solutions or any of its affiliates (collectively, “Thalēs”). The information presented above is only for informational and educational purposes and is not an offer to sell or the solicitation of an offer to buy any securities or other instruments. Additionally, the above information is not intended to provide, and should not be relied upon for investment, accounting, legal or tax advice. Thalēs makes no representations, express or implied, regarding the accuracy or completeness of this information, and the reader accepts all risks in relying on the above information for any purpose whatsoever.