Interactive
What biology does each method go after?
Two methods can recover the same number of hits from completely different parts of the cell. Effective Pathways counts how many distinct Reactome pathways a method's picks span. This is done at the scale of a single 100-gene batch (EP-B), a whole screen (EP-S), and the full test dataset (EP-D). The vocabulary is Reactome's 186 level-2 groups (Innate Immune System, Cell Cycle Mitotic, Signaling by Receptor Tyrosine Kinases, …), so the numbers read as counts of named biological programs. For example, you could say, “The average batch predicted by AssayFormer reflects EP-B = 19.3 Reactome pathways, and predictions for the whole dataset cover EP-D = 65 pathways.” For reference, the expectation across uniform draws is 21.7 / 56.4 / 81.6 at the three scopes.
Effective pathways by scope
How the metric works
- Map genes to broad Reactome pathways. Each Reactome-annotated picked gene is mapped to one of 186 level-2 pathway groups. We use these broad groups because Reactome's 2,012 narrow leaf pathways are too specific: almost every gene in a 30-gene sample would fall into a different group, making nearly every method look equally diverse.
- Give every gene one vote. If a gene belongs to several pathways, one of them is sampled at random and the calculation is repeated many times. This prevents well-studied genes with many pathway annotations from counting more heavily.
- Turn the spread into an intuitive count. The pathway distribution is summarized as an effective number. A score of 19.3 means the picks are as broadly spread as they would be across 19.3 equally represented pathways. Concentrating most picks in one pathway lowers the score.
- Compare equal sample sizes. Every method is evaluated using the same
number of Reactome-annotated genes:
—per batch,—per screen, and—across the test dataset. This resampling step, called rarefaction, prevents methods with more annotated picks from appearing more diverse simply because they contributed more data. Compare methods within an EP column; the three scopes use different sample sizes and should not be compared directly with one another. - Leave unreliable comparisons blank. A score is reported only when at least — of the units at that scope contain enough annotated genes. Otherwise, keeping only the better-annotated units would make the method look artificially diverse. Incomplete batches are excluded first because their missing picks are already measured by Shortfall.
What that looks like
A visual illustration of the biology used by selected methods. Each sunburst lays one method's proposed genes out over Reactome. The inner ring shows the eight largest top-level categories plus everything else; the outer ring shows the more specific level-2 groups used by the EP metric.
The methods are pulling from visibly different parts of the cell. The analysis page has a heatmap showing the over- and under-representation of different types of biology →
Does breadth cost enrichment?
Does recovering more hits require a method to concentrate on fewer biological pathways? Each point compares one method's enrichment factor with its pathway diversity. LLMs show clearly lower diversity within individual batches (EP-B) and screens (EP-S), meaning their picks are more concentrated in particular kinds of biology. That difference largely disappears at the dataset level (EP-D): across all 20 test screens, LLMs collectively cover a similarly broad range of biology. Switch the diversity axis to see how the relationship changes by scope, and hover over a point to identify the method.
A dash is not a zero. It means fewer than — of the relevant units contained enough Reactome-annotated genes for a fair comparison. Reporting only the better-annotated units would bias the method's score upward, so the value is omitted.