AI in Research

A very fast analyst who is confidently wrong about numbers.

Used carelessly, these tools produce consensus faster than you can think. Used well, they buy back the single scarcest input in this job — attention.

01

There is a structural problem with using language models for investment research that I think gets discussed far too little.

These systems are trained on what has already been written. Asked what they think about a company, they return a well-organised distillation of the existing published view — which is to say, they are consensus machines. Consensus is precisely the thing an active manager is being paid to differ from. A tool that makes it faster and more fluent to arrive at the standard opinion is not obviously an advantage. In a market as thinly covered as Canada's, where the edge comes from doing work nobody else has done, it can be an active liability.

So I don't use them to form views. I use them to remove the mechanical work that stands between me and forming my own — and to attack the views once I have them.

The gain isn't better answers. It's reading everything instead of sampling it.

02
  1. USE 01

    Reading the whole corpus, not a sample

    The honest constraint in fundamental research has always been that there are more filings, transcripts, and disclosures than any person can read. So you sample, and you hope the sample was representative.

    Machine assistance changes the arithmetic. Every transcript for a whole sub-sector across forty quarters becomes tractable — not to summarise, but to locate the handful of passages worth reading closely myself. The model's job is retrieval and triage. The reading is still mine.

  2. USE 02

    Tracking how language changes over time

    Management commentary drifts before the numbers do. A segment that was "a priority" becomes "an area we continue to evaluate." Hedging language appears around a contract that used to be described flatly. Across one company over twelve quarters, that drift is hard to see; across a sector, it is invisible without help.

    This is the application where I think the technology is genuinely differentiating rather than merely convenient — it surfaces questions, which I then answer with primary work.

  3. USE 03

    Structured extraction from unstructured disclosure

    Canadian disclosure is inconsistent in ways that make comparison genuinely laborious: segment definitions that shift, non-standard adjusted metrics, note structures that vary by issuer. Pulling a consistent series out of that has historically been a manual exercise, which is why it often doesn't get done.

    Extraction is the one task where these models are both reliable and checkable — every extracted figure carries a citation back to the source document, and spot-checking against the filing is fast. Anything that can't be traced is discarded rather than trusted.

  4. USE 04

    Arguing against me, on demand

    The highest-value use I've found, and the one that plays directly to what these systems actually are. Because they encode consensus, they are excellent at reconstructing the view I am betting against.

    I hand over a completed thesis and ask for the strongest possible case that it's wrong, the base rates that argue against it, and what a sceptical analyst would attack first. It is a tireless, unoffendable counterparty — and unlike a colleague, I can ask it to do this eleven times without social cost. Several positions have been resized, and a few abandoned, off the back of this step.

  5. USE 05

    Compressing the build time on quantitative work

    The distance between "I wonder whether X holds in this market" and a defensible test of X used to be measured in days, which meant most such questions never got asked. That distance is now short enough that the marginal question gets tested.

    The risk is obvious and worth naming: when testing becomes cheap, the multiple-comparisons problem gets much worse. More shots on goal means more spurious results that look real. The discipline described under Models — hypothesis stated first, specification count tracked — matters more now than it did when the friction did that policing for me.

03

Where I don't let it near the work.

Knowing the failure modes precisely is the whole of the skill. These are the ones I treat as hard boundaries.

04

Where I think this goes.

My working assumption is that the advantage does not accrue to whoever adopts the tools — that will be everyone, quickly — but to whoever is clearest about which parts of the job were never the bottleneck.

Reading was a bottleneck. Extraction was a bottleneck. Building a test was a bottleneck. Judgment was not, and removing friction around judgment mostly produces faster bad decisions. The managers who do well with this are likely to be the ones who use it to widen the funnel at the top and leave the narrow end alone.

I'm genuinely interested in how other practitioners are drawing these lines, and I'd welcome the argument if you draw them somewhere else. iam@brandontu.com.

Next Personal →