C&T Technology Co., Ltd. · Bắc Ninh, Vietnam
Clonal fitness in clonal haematopoiesis — a measurement pipeline for predicting which pre-leukaemic clones expand.
01 — The question
Clonal haematopoiesis is an expanded blood cell clone driven by a somatic mutation — common enough in later life that it turns up as a routine incidental finding in sequencing. Almost none of it will ever progress to disease. Today, nothing tells an individual carrier which group they are in.
carry a clonal haematopoiesis mutation, usually found incidentally on routine sequencing.
progress to overt myeloid disease. The rest never do — and current practice cannot tell them apart.
Clinical practice today stratifies by gene identity alone, scoring a known disease hotspot and a variant that has never been seen before in the same gene identically. LeukemiPrediag targets a different, measurable quantity instead: clonal fitness — the exponential growth rate of a clone, measured from repeat sampling of the same person over time. Fitness is measurable in every carrier, sick or not — a far richer training signal than a binary progression label that, given the 0.5–1%/year rate, would discard roughly 99% of carriers. LeukemiPrediag tests whether that fitness can be predicted at the level of the individual variant, rather than just the gene it sits in. Fitness is not a correlate of progression; it is the mechanism of progression itself.
02 — The instrument
Across harmonised cohorts, real signal was enriched 4.4× over the measured noise floor in the Lothian Birth Cohorts, and 4.1× in SardiNIA, against a chance expectation of 1.0×. The noise floor itself is measured directly: synonymous mutations have zero fitness by construction, so their spread defines the error bar.
The estimated growth rate per year, broken out by gene, was recovered without the model being told the answer in advance — spliceosome factor genes showed the fastest growth, DNMT3A the slowest, matching the exact ordering already reported in the published literature on this same data (Fabre et al.). The pipeline itself runs on NVIDIA BioNeMo's ESM-2 protein language model, cross-checked against a HuggingFace backend implementation, with results recorded per variant.
03 — Negative results
We tested four additional hypotheses for improving on the baseline fitness predictor. Every one of them scored worse than the training-fold mean, so no improvement is reported for any of them.
We report that plainly, on purpose: a team that only shows what worked is easy to trust and hard to verify; a team that also shows what it tried and ruled out is checkable.
04 — Roadmap
The workstation-scale result above — cohort harmonisation, signal enrichment, external validation against the published literature, and the NVIDIA BioNeMo integration — already produces a publishable number. The next stage is biobank-scale validation, gated on securing an institutional partnership; a clinical pilot follows, gated on that biobank-scale result.