App / Tech notes

MERCURIUS™ 1536 DRUG-seq Generates Reproducible Ultra-High-Throughput Transcriptomic Perturbation Datasets That Decipher Compound Mechanisms of Action and Toxicity Pathway Activation

A 33-compound study demonstrating reproducible transcriptomic profiling at 1536-well scale, with gene-level resolution of mechanisms of action, toxicity pathways, and dose-dependent responses.

How to Generate AI-Ready Perturbation Datasets for Reproducible Mechanism-of-Action and Toxicity Prediction with MERCURIUS™ 1536 DRUG-seq 

AI models trained to predict compound mechanism of action or toxicity are only as good as the perturbation datasets they learn from. These perturbation datasets must be standardized, reproducible, and large enough for AI-scale drug discovery. This is strategically crucial as proprietary perturbation data, not model architecture, is increasingly the competitive moat for AI-native biopharma, and standardized data generation is what makes rapid iteration possible within a lab-in-the-loop workflow. 

Yet, most screening-scale compound-perturbation profiling methods force a trade-off. Probe-based gene expression profiling captures only a fraction of the transcriptome, and morphological profiling methods like Cell Painting detect phenotypic changes without revealing which gene expression pathway or program drove it. 

MERCURIUS™ 1536 DRUG-seq closes that gap. In a 33-compound perturbation screen across two independently processed 1536-well plates, the assay generated transcriptome-wide data with the reproducibility and gene-level resolution to correctly cluster compounds by known mechanism, detect activation across 18 toxicity and 14 mechanism-of-action (MoA) pathways, and capture early apoptotic transcriptional signals before conventional viability assays could.  

 

What the MERCURIUS™ 1536 DRUG-seq Compound-Perturbation Data Shows 

  • 12,000+ genes detected per sample from just 800 HepG2 cells at 0.7M reads: enough depth per well to keep cost per data point low without sacrificing the gene detection resolution a model needs to learn from. 
  • Inter-plate DE gene correlation of R² = 0.94 across two independently processed plates: reproducibility strong enough to combine datasets across runs with confidence, without reliance on computational correction for batch effects that could otherwise bias training data. 
  • Unsupervised clustering correctly grouped compounds by known MoA, including PPARγ agonists, proteasome inhibitors, ER stress inducers, and topoisomerase inhibitors: gene expression pathway activation is strong enough for a model to learn MoA structure directly from transcriptomic data. 
  • 18 toxicity pathways and 14 MoA signatures scored per compound: a single MERCURIUS™ 1536 DRUG-seq run generates rich, multi-parameter biological outputs instead of one endpoint per assay. 
  • Gene-level resolution separated two compounds sharing the same MoA: mitoxantrone and etoposide both act as topoisomerase II inhibitors, but only the transcriptomic data revealed mitoxantrone’s distinct ferroptosis signature and its greater potency, differences invisible to imaging methods. 
  • Apoptosis signal detected before measurable cytotoxicity: for etoposide, transcriptomic stress response was found at doses below what a simultaneous CTBlue cell viability assay could detect, giving models an earlier, more sensitive toxicity signal to train on. 
  • Concordant tPODs and dose-response curves across replicate plates: dose-response structure was concordant on both independently processed plates, supporting its use for point-of-departure estimation. 

Why this matters for your model 

A dataset that clusters compounds correctly by MoA, distinguishes mechanistically related compounds at the gene level, and catches toxicity signal earlier than apical endpoints gives a foundation model richer, more informative data points per experiment. This makes the difference between a dataset that’s merely large and one that’s rich in broad biological information for training MoA classifiers and toxicity predictors or experimentally evaluating AI-generated predictions at scale. 

 

Authors: Alexandre Coudray¹, Elodie Koenig¹, Maya Wilson², Elizabeth Bourne², Mark Ofield², Yaoyao Xiong², Greg Slodkowicz², Adam Peall², Alix Buu Hoang¹, Vincent Hahaut¹, and Daniel Alpern¹

Ready to talk about your next RNA-seq study?

Tell us about your project and we will help you find the right approach.