Cancer diagnostic models perform well in the lab where they were built and degrade when moved to another hospital. The cause is usually not the cells — it is how the data was collected. We generate training data from optical first principles, sampling the device, the stain and the disease independently of one another.
Every statistic of our synthetic output is placed against the interval formed by two independent real datasets. The goal is not to imitate one of them — it is to occupy a generation space that contains both.
On this axis the two real datasets do not intersect at all — two entirely different imaging conditions. Both still fall inside the range we generate.
A fit with lower error was found and discarded: it assigned the red cell a physiologically impossible value. We chose the biology over the number.
Edge sharpness sits below both real datasets. We traced it to how the model handles axial depth; the correction is under test. We do not call it closed until it is.
A network learns the cheapest signal that separates the classes; that signal does not have to live in the cell. And in real data, the disease and the way the data was collected always arrive together.
Patient samples come from specialist centres, controls from routine labs — different scanners, different stain batches, often different years. This is not a collection error; it is how healthcare systems work.
The test set is drawn from the same two sources, so the shortcut is rewarded there too. The failure only surfaces at a third hospital, and nobody can explain it.
Diversity weakens the bias but does not remove it — what is needed is balance, and balance cannot be achieved for rare cases. And you cannot correct a variable you never recorded.
We build the scene first, then compute how light travels through the microscope to turn it into an image. The scene plan is also the answer key — nothing has to be annotated afterwards.
Cell types, proportions and arrangement set by disease stage
Nucleus, cytoplasm, membrane; refractive index assignment
The pixel-level label is produced here, never inferred
Point spread function, focus, numerical aperture, illumination
Measured stain vectors, colour profile, noise, artefacts
Stain batch, scanner profile and smear thickness are drawn without reference to the label. The probability that a diseased slide carries a given colour cast is identical to that of a healthy one. The model cannot use colour as a shortcut — it has to look at the cell. This can never be guaranteed by collecting real patient data.
The same scene can be regenerated with only the focus changed, only the stain changed, or only the numerical aperture changed. You see exactly where your model breaks. This experiment cannot be run on real data, because the same smear from the same patient cannot be captured through five different optics.
Each result comes from an experiment whose acceptance criterion was written before the run and repeated across eight random seeds.
A model trained purely on synthetic data, having never seen a real image, classifies expert-labelled real cells at this score. Majority baseline on the same test: 0.342.
A principle we wrote into the model — a ruptured cell carries the same chromatin — was measured across 988 real cells: real 0.920, model 0.948. Changing the measurement threshold sixfold left the real value unmoved.
Where real data is scarce (≤50 examples per class), synthetic pre-training reaches the same performance with at least 1.6× fewer real examples, rising to 3.5× depending on the task. In the data-rich regime there is no effect — we measured that too.
Synthetic data does not replace real data, and we make no such claim. Regulators accept synthetic samples for enriching a training set; the evidence for a device's clinical performance has to come from real patients. Our value sits in training, pre-adaptation and robustness testing.
| Layer | Status |
|---|---|
| Optical chain | verified — two independent acquisition paths agreed to within 0.6% at every stain density |
| Cell-level morphology | verified — against a 41,621-cell expert-labelled reference set, smudge cells included |
| Transfer on pathological cells | measured — 0.647 macro F1 against a 0.342 baseline |
| Scene level: stage and proportions | unverified — the relationship between disease stage and lymphocyte area fraction, smudge-cell percentage and rouleaux degree has not yet been checked against real data. No public image dataset carries stage information; a clinical reference study is under way. |
If you run on third-party scanners, every new device is a new calibration problem and there is no hardware you control. We shorten deployment time.
Without a decade of labelled archive, synthetic data is not an alternative — it is the only route. We work with early partners on reference terms.
If you want models that work on your instrument, we build a simulation environment tuned to your optical profile. You already know that profile; sharing it is not sensitive.
The optical behaviour of an instrument can be recovered from a few hundred anonymised images taken with it: point spread function, stain vectors, noise characteristics and illumination profile. The request is not “send us your PSF” — it is “send anonymised images and we will measure the profile ourselves.”
Send us roughly 200 anonymised images and the basic specification of your instrument. We extract your optical and staining profile and generate labelled synthetic data matched to those conditions. You run the evaluation — your model, your test set. We agree the success criterion together, in advance.