Research

Evaluating multi-agent materiality ranking.

If we're going to claim that agent-based reading ranks material events better than the alternatives, the claim needs a benchmark, a held-out set and numbers anyone can check. This is where we publish that work.

Working paper

Evaluating Multi-Agent Materiality Ranking for Institutional Investment Research

The question

Given a day's worth of filings, transcripts, news and price action across a portfolio, how well does each approach identify the events an experienced analyst would have escalated?

Method

Events are labelled by analyst review against a fixed materiality rubric. Each system produces a ranked list for the same day and portfolio. We report precision, recall and F1 against the labelled set, with the labelling protocol and rubric published alongside.

Systems compared

  • Keyword baseline — term matching over the same sources
  • Sentiment model — off-the-shelf financial sentiment classification
  • Single LLM — one model reading all sources in one pass
  • Portfolio Intelligence — specialised agents with thesis-conditioned scoring
SystemPrecisionRecallF1
Keyword baseline
Sentiment model
Single LLM
Portfolio Intelligence

Benchmarking in progress.

This table stays empty until the evaluation is complete. We publish measured results only — no projected, illustrative or selectively reported figures.

If you'd like to be notified when the results are published, or to review the labelling rubric before then, write to support@pullelainnovation.com.

Standards

How we'll report.

Held-out evaluation

Labels are fixed before systems are run. No tuning against the evaluation set.

Published rubric

The materiality definition used for labelling is published in full, so results can be contested.

Failure cases included

Where the system misses or over-ranks, those examples are published with the aggregates.