All cases are synthetic, drawn by a script so that we know the truth for every one. 80 form the training set; 20 are held out as a test set the network never sees while training. Click any case to put it under the inspector.
The parameters the network can see
Where the cases come from
every nucleus was scanned at two labs · the other lab’s stain is weaker: paler nuclei, less contrastTraining set ·
Test set · held out
How separable are the measurements?
filled = training set · hollow = test set (when revealed) · click a point to inspectThe network, live
forward pass for the case in the inspector · hover any node, map or connectionThe single-layer network as a weighted checklist
every weight, largest first · orange pushes toward , blue awayWhat each hidden unit responds to
first hidden layer · updated at the end of each epochTraining set · what the network currently calls
Learning curves
The network, frozen
held-out cases go through one at a time · N classifies the next · the weights never change hereTest set ·
What the calls would mean in a real population
sensitivity and specificity from above · prevalence from the sliderThe game: spot the same nucleus
one view of the last batch (or of the nucleus you pick below) against every other view of the batch · no label anywhereThe encoder turns every view into 8 numbers and compares them by cosine; the vote is a softmax over the cosines at temperature 0.2. The answer is the nucleus’s other view, framed in orange. At the start the vote is spread thin over every candidate; pretraining is nothing but making the partner win this game, for every view of every batch. Watch the orange bar grow as it runs.
The encoder, live · one pair of views
the nucleus of the last batch (or the one you click below) goes through the encoder · hover any filter, map or connectionWhat the code learns to keep, and to ignore
the same nucleus in all 8 orientations from both labs, and the code of each view · at the start and now · then another nucleusEach strip is one view’s code, the 8 numbers as bars on one scale per row. At the start a nucleus’s 16 codes are no closer to each other than to another nucleus’s: a random encoder gives every view much the same code, so nothing is told apart. Pretraining makes a nucleus’s own codes agree (orientation and stain are ignored) while a different nucleus’s codes move away (its identity is kept). Untick the other lab’s scans above, reset, and the right-hand half no longer lines up.
This batch: 8 of its 20 nuclei, two views each
every view against every other · ringed cells are the pairsOrange = alike, blue = unlike. Pretraining pulls the ringed cells towards orange and everything else towards blue: each nucleus becomes recognisable across its views and distinct from the rest. Views are flips, rotations and (when ticked) the other lab’s scans.
Pretraining curve
The embedding map
Two of the eight code numbers, for every nucleus and its second view (joined by a line). As pretraining runs, the two views of a nucleus pull together and nuclei arrange themselves by what makes them recognisable. Colour by the answers above to see how far the classes separate without any label having been used.
Foundation: a code from 100 unlabelled nuclei
the specimens of the foundation question · no labels anywhere100 nuclei from the three image questions’ training sets (34 from atypia, 33 from enlargement, 33 from irregularity), our lab’s scans, and never a label. The foundation model will learn to turn each of them into a code of eight numbers without being told what any question asks; afterwards a single layer on that code, trained with a handful of labelled cases, answers each question. Colour the tray by each nucleus’s own answer to see what the code will have to capture on its own, then go to 2 · Train to pretrain it and to 3 · Test to see what the code is worth.
The 100 unlabelled nuclei
What the code is worth
This chart is measured on the side, never fed back: after each epoch a single layer is trained on the frozen code with 20 labelled training cases of each question and scored on 40 nuclei it never saw (the question’s 20 test nuclei and the 20 held-out ones). Pretraining itself never sees a label, yet the code becomes worth more for every question. Size is readable even from a random code; contour irregularity is not, and that is what the pretraining earns.
The slide
The nuclei ranked by attention; each bar is that nucleus’s share of the slide’s attention, and the shares add up to 100%. Frames on the slide grow with the share. Hover a nucleus to see how its score is computed below.
Who looks at whom
one layer of self-attention: every nucleus asks the others, and adds what it hears to its own code before the scorer sees itHover a nucleus on the slide, a row of the map or a row of the card below, or click one on the slide or the map to pin it: the lines on the slide show whom it listens to, and the card below takes its decision apart. With the distance cost the model learns how far to look; with a 2 × 2 focus every member has two atypical neighbours, a scattered one has none, and that difference is what reaches the scorer.
How a nucleus decides where to look
The whole model for this slide
How a nucleus is scored
the attention network, the same for every nucleus: its code → tanh units → one score · a softmax over the slide turns the 20 scores into weightsThe slide’s summary, and the call
the nuclei’s codes averaged with the attention weights, then a single layer on the summaryLearning curves
The right-hand chart is the point: the share of a positive slide’s attention that lands on its atypical nuclei. Nobody told the model which nuclei those are; with uniform weights they get their share of the slide, about 15%. A plain average stays there by construction.
The slides
Training slides
Test slides · never trained on, scored as it goes
Slides
The slide
The slides
Training slides
Test slides · held out
The slide under test
held-out slides go through one at a time · N classifies the next · the weights never change hereTest slides ·
Fields of bladder: is it invasive?
The field
Invasion is the conjunction of three cues, and every pattern that lacks one is a real mimic:
The fields
Training fields
Test fields · held out
The field
What the CIS head weighs
What the invasion head weighs
Who looks at whom
two layers of self-attention: every nucleus asks the others, with a learned cost per nucleus diameter of distance, and adds what it hears to its own code before the heads see itHover a nucleus on the field, a row of the map or a row of the card below, or click one on the field or the map to pin it: the lines on the field show whom it listens to, and the card below takes its decision apart. A nucleus in a solid nest has three neighbours a diameter away and one in the middle surrounded on every side; a nucleus in an angulated nest has two or three and an open side. That, and where the membrane runs, is what reaches the heads.
How a nucleus decides where to look
The whole model for this field
How a nucleus is scored
the attention network of this head, the same for every nucleus: its token → tanh units → one score · a softmax over the field turns the scores into weightsThe field’s summary, and the call
the nuclei’s tokens averaged with this head’s attention weights, then a single layer on the summaryThe mimic table
the test fields of each pattern: how often called CIS, how often called invasive · scored as the model trains, never trained onLearning curves
The fields
Training fields
Test fields · never trained on, scored as it goes
The field under test
held-out fields go through one at a time · N classifies the next · the weights never change hereWhat the CIS head weighed
What the invasion head weighed
The mimic table so far
the classified test fields of each patternTest fields ·
Reports
The report
The rule the diagnosis follows
what the generator did with the findings; the model has to find it in the wordsThe reports
Training reports
Test reports · held out
Never trained on · invasion under a normal surface
The report
How a word was chosen
How each head decides where to look
Who reads whom, across the report
Learning curves
The words as the model sees them
The reports
Training reports
Test reports · never trained on
The findings and the requisition go in, the description and the diagnosis come out
The findings
Change a finding and write again: the report follows the findings, right or wrong. Flip the surface to normal with invasive nests and see what it does with a combination it never saw.