AssayBench · Video explainer
Can a model predict a CRISPR screen before you run it?
A three-minute animated walkthrough of the paper: what a CRISPR screen is, what counts as a hit, how the 1,920-screen benchmark is built and split, and how AnDCG scores a ranking so that screens of wildly different hit rates can be compared.
Chapters
- 0:00What a screen is Knock out one gene, infect with SARS-CoV-2, and see what survives.
- 0:23What the cell did Endpoints, not transcriptomes: survival, infection, immune evasion.
- 0:36The dataset 1,920 screens, processed and curated, split by publication date.
- 0:51Five phenotypes Viability, drug response, infection, reporter activity, trafficking.
- 1:00Why it is hard Held-out screens are genuinely new, and the hits have a long tail.
- 1:23The task A screen in plain text in, a ranked list of 100 genes out.
- 1:30AnDCG, built up Cumulative gain to DCG to nDCG, then rescaled per screen.
- 2:23Where models stand Generalist LLMs lead, 16% of the way from guessing to perfect.