Synthetic microscopy data · Haematology

The same blood, as if imaged in two different labs.

Cancer diagnostic models perform well in the lab where they were built and degrade when moved to another hospital. The cause is usually not the cells — it is how the data was collected. We generate training data from optical first principles, sampling the device, the stain and the disease independently of one another.

Simulated field of view
Transfer A model that has never seen a real image scores 0.647 macro F1 on expert-labelled real cells — majority baseline 0.342
Prediction confirmed A physical principle derived from the model held to within 3% across 988 real cells
Status First paid delivery completed; clinical validation in progress
Measurement · Calibration status

We do not claim realism.
We measure where we stand.

Every statistic of our synthetic output is placed against the interval formed by two independent real datasets. The goal is not to imitate one of them — it is to occupy a generation space that contains both.

Interval spanned by two real datasets Synthetic output

No overlap in illumination

On this axis the two real datasets do not intersect at all — two entirely different imaging conditions. Both still fall inside the range we generate.

We rejected a better number

A fit with lower error was found and discarded: it assigned the red cell a physiologically impossible value. We chose the biology over the number.

Still open

Edge sharpness sits below both real datasets. We traced it to how the model handles axial depth; the correction is under test. We do not call it closed until it is.

Problem

The model learns the lab, not the disease

A network learns the cheapest signal that separates the classes; that signal does not have to live in the cell. And in real data, the disease and the way the data was collected always arrive together.

How it happens

Cases and controls take different routes

Patient samples come from specialist centres, controls from routine labs — different scanners, different stain batches, often different years. This is not a collection error; it is how healthcare systems work.

Why it goes unnoticed

Validation scores look excellent

The test set is drawn from the same two sources, so the shortcut is rewarded there too. The failure only surfaces at a third hospital, and nobody can explain it.

Why scale does not fix it

More data does not break the correlation

Diversity weakens the bias but does not remove it — what is needed is balance, and balance cannot be achieved for rare cases. And you cannot correct a variable you never recorded.

How it works · Generation chain

The label is an input, not an output

We build the scene first, then compute how light travels through the microscope to turn it into an image. The scene plan is also the answer key — nothing has to be annotated afterwards.

01

Scene

Cell types, proportions and arrangement set by disease stage

02

Geometry

Nucleus, cytoplasm, membrane; refractive index assignment

03

Mask is born

The pixel-level label is produced here, never inferred

04

Optics

Point spread function, focus, numerical aperture, illumination

05

Stain and sensor

Measured stain vectors, colour profile, noise, artefacts

What sets it apart

Conditions are sampled independently of the diagnosis

Stain batch, scanner profile and smear thickness are drawn without reference to the label. The probability that a diseased slide carries a given colour cast is identical to that of a healthy one. The model cannot use colour as a shortcut — it has to look at the cell. This can never be guaranteed by collecting real patient data.

What sets it apart

One variable can be isolated

The same scene can be regenerated with only the focus changed, only the stain changed, or only the numerical aperture changed. You see exactly where your model breaks. This experiment cannot be run on real data, because the same smear from the same patient cannot be captured through five different optics.

Measurement · Product utility

What we measured

Each result comes from an experiment whose acceptance criterion was written before the run and repeated across eight random seeds.

Transfer on pathological cells
0.647macro F1

A model trained purely on synthetic data, having never seen a real image, classifies expert-labelled real cells at this score. Majority baseline on the same test: 0.342.

Prediction confirmed
3%deviation

A principle we wrote into the model — a ruptured cell carries the same chromatin — was measured across 988 real cells: real 0.920, model 0.948. Changing the measurement threshold sixfold left the real value unmoved.

Annotation saving
1.6–3.5×

Where real data is scarce (≤50 examples per class), synthetic pre-training reaches the same performance with at least 1.6× fewer real examples, rising to 3.5× depending on the task. In the data-rich regime there is no effect — we measured that too.

Scope

What we do, and what we do not

Synthetic data does not replace real data, and we make no such claim. Regulators accept synthetic samples for enriching a training set; the evidence for a device's clinical performance has to come from real patients. Our value sits in training, pre-adaptation and robustness testing.

  • Labelled data for training and pre-training, generated at source
  • Controlled generation of rare morphologies
  • Model adaptation ahead of deployment at a new site
  • Single-variable robustness and edge-case testing
  • A technical annex documenting every generation parameter
  • Not a substitute for clinical validation data
  • No measurable benefit in the data-rich regime
  • Produces no diagnosis — this is not a medical device
Status · Verified and unverified

Knowing a gap beats assuming there is none

LayerStatus
Optical chain verified — two independent acquisition paths agreed to within 0.6% at every stain density
Cell-level morphology verified — against a 41,621-cell expert-labelled reference set, smudge cells included
Transfer on pathological cells measured — 0.647 macro F1 against a 0.342 baseline
Scene level: stage and proportions unverified — the relationship between disease stage and lymphocyte area fraction, smudge-cell percentage and rouleaux degree has not yet been checked against real data. No public image dataset carries stage information; a clinical reference study is under way.
Who this is for

Teams where the pain is sharpest

Primary

Software teams without hardware

If you run on third-party scanners, every new device is a new calibration problem and there is no hardware you control. We shorten deployment time.

Primary

New entrants

Without a decade of labelled archive, synthetic data is not an alternative — it is the only route. We work with early partners on reference terms.

Enterprise

Scanner and device manufacturers

If you want models that work on your instrument, we build a simulation environment tuned to your optical profile. You already know that profile; sharing it is not sensitive.

Data sharing

We do not ask for your design parameters

The optical behaviour of an instrument can be recovered from a few hundred anonymised images taken with it: point spread function, stain vectors, noise characteristics and illumination profile. The request is not “send us your PSF” — it is “send anonymised images and we will measure the profile ourselves.”

Pilot

Two weeks, free, on your own data

Send us roughly 200 anonymised images and the basic specification of your instrument. We extract your optical and staining profile and generate labelled synthetic data matched to those conditions. You run the evaluation — your model, your test set. We agree the success criterion together, in advance.