AI in Research
A very fast analyst who is confidently wrong about numbers.
Used carelessly, these tools produce consensus faster than you can think. Used well, they buy back the single scarcest input in this job — attention.
There is a structural problem with using language models for investment research that I think gets discussed far too little.
These systems are trained on what has already been written. Asked what they think about a company, they return a well-organised distillation of the existing published view — which is to say, they are consensus machines. Consensus is precisely the thing an active manager is being paid to differ from. A tool that makes it faster and more fluent to arrive at the standard opinion is not obviously an advantage. In a market as thinly covered as Canada's, where the edge comes from doing work nobody else has done, it can be an active liability.
So I don't use them to form views. I use them to remove the mechanical work that stands between me and forming my own — and to attack the views once I have them.
The gain isn't better answers. It's reading everything instead of sampling it.
-
USE 01
Reading the whole corpus, not a sample
The honest constraint in fundamental research has always been that there are more filings, transcripts, and disclosures than any person can read. So you sample, and you hope the sample was representative.
Machine assistance changes the arithmetic. Every transcript for a whole sub-sector across forty quarters becomes tractable — not to summarise, but to locate the handful of passages worth reading closely myself. The model's job is retrieval and triage. The reading is still mine.
-
USE 02
Tracking how language changes over time
Management commentary drifts before the numbers do. A segment that was "a priority" becomes "an area we continue to evaluate." Hedging language appears around a contract that used to be described flatly. Across one company over twelve quarters, that drift is hard to see; across a sector, it is invisible without help.
This is the application where I think the technology is genuinely differentiating rather than merely convenient — it surfaces questions, which I then answer with primary work.
-
USE 03
Structured extraction from unstructured disclosure
Canadian disclosure is inconsistent in ways that make comparison genuinely laborious: segment definitions that shift, non-standard adjusted metrics, note structures that vary by issuer. Pulling a consistent series out of that has historically been a manual exercise, which is why it often doesn't get done.
Extraction is the one task where these models are both reliable and checkable — every extracted figure carries a citation back to the source document, and spot-checking against the filing is fast. Anything that can't be traced is discarded rather than trusted.
-
USE 04
Arguing against me, on demand
The highest-value use I've found, and the one that plays directly to what these systems actually are. Because they encode consensus, they are excellent at reconstructing the view I am betting against.
I hand over a completed thesis and ask for the strongest possible case that it's wrong, the base rates that argue against it, and what a sceptical analyst would attack first. It is a tireless, unoffendable counterparty — and unlike a colleague, I can ask it to do this eleven times without social cost. Several positions have been resized, and a few abandoned, off the back of this step.
-
USE 05
Compressing the build time on quantitative work
The distance between "I wonder whether X holds in this market" and a defensible test of X used to be measured in days, which meant most such questions never got asked. That distance is now short enough that the marginal question gets tested.
The risk is obvious and worth naming: when testing becomes cheap, the multiple-comparisons problem gets much worse. More shots on goal means more spurious results that look real. The discipline described under Models — hypothesis stated first, specification count tracked — matters more now than it did when the friction did that policing for me.
Where I don't let it near the work.
Knowing the failure modes precisely is the whole of the skill. These are the ones I treat as hard boundaries.
- Any unsourced number These systems produce plausible figures with complete confidence and no signal that they were invented. A number without a traceable citation to a primary document does not enter a model, a note, or a conversation. This rule has no exceptions.
- Valuation judgment Deciding what a business is worth requires weighing things that aren't in the training data — the quality of a management team you've met, the durability of an advantage that hasn't been tested yet. Fluent output about valuation is the most dangerous kind, because it reads exactly like insight.
- Genuinely novel situations The cases where the analysis is most valuable are the ones with the least precedent, which is precisely where a model trained on precedent is weakest — and where it will be most confidently generic.
- Anything touching material non-public information Confidential and non-public material does not go into third-party systems. This is a compliance obligation before it is a preference, and it constrains the tooling choices rather than the other way around.
- Backtests where the model knows the future A system trained through a given date has, in effect, read the outcome. Asking it to reason about a historical period is a lookahead-bias generator wearing a helpful interface.
- The final step before a decision Machine output is never the last thing I look at. If it would be, the work isn't finished.
Where I think this goes.
My working assumption is that the advantage does not accrue to whoever adopts the tools — that will be everyone, quickly — but to whoever is clearest about which parts of the job were never the bottleneck.
Reading was a bottleneck. Extraction was a bottleneck. Building a test was a bottleneck. Judgment was not, and removing friction around judgment mostly produces faster bad decisions. The managers who do well with this are likely to be the ones who use it to widen the funnel at the top and leave the narrow end alone.
I'm genuinely interested in how other practitioners are drawing these lines, and I'd welcome the argument if you draw them somewhere else. iam@brandontu.com.