Held-out evaluation
Labels are fixed before systems are run. No tuning against the evaluation set.
Research
If we're going to claim that agent-based reading ranks material events better than the alternatives, the claim needs a benchmark, a held-out set and numbers anyone can check. This is where we publish that work.
Working paper
Given a day's worth of filings, transcripts, news and price action across a portfolio, how well does each approach identify the events an experienced analyst would have escalated?
Events are labelled by analyst review against a fixed materiality rubric. Each system produces a ranked list for the same day and portfolio. We report precision, recall and F1 against the labelled set, with the labelling protocol and rubric published alongside.
| System | Precision | Recall | F1 |
|---|---|---|---|
| Keyword baseline | — | — | — |
| Sentiment model | — | — | — |
| Single LLM | — | — | — |
| Portfolio Intelligence | — | — | — |
Benchmarking in progress.
This table stays empty until the evaluation is complete. We publish measured results only — no projected, illustrative or selectively reported figures.
If you'd like to be notified when the results are published, or to review the labelling rubric before then, write to support@pullelainnovation.com.
Standards
Labels are fixed before systems are run. No tuning against the evaluation set.
The materiality definition used for labelling is published in full, so results can be contested.
Where the system misses or over-ranks, those examples are published with the aggregates.