Interactive

What biology does each method go after?

Two methods can recover the same number of hits from completely different parts of the cell. Effective Pathways counts how many distinct Reactome pathways a method's picks span. This is done at the scale of a single 100-gene batch (EP-B), a whole screen (EP-S), and the full test dataset (EP-D). The vocabulary is Reactome's 186 level-2 groups (Innate Immune System, Cell Cycle Mitotic, Signaling by Receptor Tyrosine Kinases, …), so the numbers read as counts of named biological programs. For example, you could say, “The average batch predicted by AssayFormer reflects EP-B = 19.3 Reactome pathways, and predictions for the whole dataset cover EP-D = 65 pathways.” For reference, the expectation across uniform draws is 21.7 / 56.4 / 81.6 at the three scopes.

Effective pathways by scope

How the metric works

What that looks like

A visual illustration of the biology used by selected methods. Each sunburst lays one method's proposed genes out over Reactome. The inner ring shows the eight largest top-level categories plus everything else; the outer ring shows the more specific level-2 groups used by the EP metric.

The methods are pulling from visibly different parts of the cell. The analysis page has a heatmap showing the over- and under-representation of different types of biology →

Does breadth cost enrichment?

Does recovering more hits require a method to concentrate on fewer biological pathways? Each point compares one method's enrichment factor with its pathway diversity. LLMs show clearly lower diversity within individual batches (EP-B) and screens (EP-S), meaning their picks are more concentrated in particular kinds of biology. That difference largely disappears at the dataset level (EP-D): across all 20 test screens, LLMs collectively cover a similarly broad range of biology. Switch the diversity axis to see how the relationship changes by scope, and hover over a point to identify the method.

A dash is not a zero. It means fewer than of the relevant units contained enough Reactome-annotated genes for a fair comparison. Reporting only the better-annotated units would bias the method's score upward, so the value is omitted.