1
00:00:00,000 --> 00:00:04,114
A CRISPR screen switches off one gene, and asks what changes.

2
00:00:04,374 --> 00:00:10,318
Take lung cells, for example, and infect them with SARS-CoV-2. Almost all of them die.

3
00:00:10,578 --> 00:00:16,685
But switch off the gene the virus needs to get in, and they survive. That gene is a hit.

4
00:00:16,945 --> 00:00:20,615
Do that for every gene in the genome, and you have a screen.

5
00:00:20,875 --> 00:00:22,616
The goal is to find the hits.

6
00:00:23,116 --> 00:00:29,266
Most perturbation benchmarks stop at the molecular state of the cell. A screen asks what the cell actually did.

7
00:00:29,526 --> 00:00:35,467
Did it survive the chemotherapy? Did it fight off the virus? Did it hide from the immune system?

8
00:00:35,967 --> 00:00:40,367
AssayBench gathers one thousand, nine hundred and twenty of these screens.

9
00:00:40,627 --> 00:00:44,640
All of them processed and curated into one comparable format.

10
00:00:44,900 --> 00:00:50,413
Then split by publication date: train on the past, test on the future.

11
00:00:50,913 --> 00:00:59,610
They span five phenotypes: viability, drug response, infection, reporter activity, and trafficking.

12
00:01:00,110 --> 00:01:11,142
And that split has teeth. A held-out screen shares almost none of its hits with the most similar screen in training. This is a test of biological understanding, not memorisation.

13
00:01:11,402 --> 00:01:19,505
There is also a long tail. A thousand genes account for half of every hit on record. The other twenty thousand share the rest.

14
00:01:19,765 --> 00:01:22,515
And that tail is where screen-specific biology lives.

15
00:01:23,515 --> 00:01:29,177
The task is: give a model the screen in plain text, and ask for the hundred most likely genes to be hits.

16
00:01:30,077 --> 00:01:36,648
So how do you score a ranking? Here are the genes a model predicted, and what each one turned out to be worth.

17
00:01:36,908 --> 00:01:39,063
Add them up, and you have cumulative gain.

18
00:01:39,663 --> 00:01:47,141
But position matters. The top of the list is what actually gets tested. So discount each gene by how far down it sits.

19
00:01:47,401 --> 00:01:52,966
Then divide by the best ranking possible. Normalized D C G, between zero and one.

20
00:01:53,226 --> 00:02:06,997
But n D C G still is not comparable across screens. For example, in a screen where a quarter of the genes are hits, guessing already scores well. For a screen with only one in a hundred hits, guessing scores almost nothing.

21
00:02:07,257 --> 00:02:14,236
So we rescale every screen, so that random performance lands on zero, and a perfect ranking lands on one.

22
00:02:14,496 --> 00:02:22,512
We call it adjusted n D C G, or just A, n D C G. Zero is guessing. One is perfect.

23
00:02:23,012 --> 00:02:27,762
So how do models do? The best method scores zero point one six.

24
00:02:28,022 --> 00:02:31,376
That is sixteen percent of the way from guessing to a perfect ranking.

25
00:02:32,226 --> 00:02:41,715
General-purpose language models come out on top, ahead of biology-specific models and trained baselines. Real signal — but a long way from solved.

26
00:02:42,315 --> 00:02:47,053
AssayBench. Building a virtual cell takes more than predicting gene expression.

