Back to Project Directory

Explainable AI Methods for Interpretable Transcriptomic Cancer Data Analysis

Context & Background

Cancer is a leading cause of mortality globally, arising from complex genetic changes and mutations. High-throughput single-cell RNA sequencing (scRNA-seq) provides cellular resolution but introduces high noise, missing values (dropouts), and extreme dimensionality. To build clinical trust in machine learning tools, there is an urgent need for explainable AI (XAI) models that clarify how predictions are made for individual samples.

Problems to be Addressed

Genomic datasets suffer from severe batch effects, meaning measurements reflect laboratory conditions rather than actual biology. High-dimensional cancer data suffers from the 'curse of dimensionality,' rendering traditional distance metrics useless. Additionally, high rates of technical and biological zero values obscure genetic expression levels.

Aims and Objectives

1. Formulate joint clustering and batch effect removal models using Bayesian nonparametrics.
2. Devise local interpretable algorithms that explain gene expression forecasts for individual test cases.
3. Validate these tools using public databases (TCGA, METABRIC) and clinical leukemia samples.

Methodology

The project uses causal Bayesian inference and hierarchical models to analyze RNA profiles. In the first phase, models are trained on bulk-RNA datasets (METABRIC, SEER) to resolve high-dimensionality. In the second phase, these methods are extended to single-cell data, integrating imputation models to correct technical dropouts and establish concept-based explanations.

Expected Outcomes

Development of a software tool for interpretability in transcriptomic analysis, deployment at AIIMS cancer lab for clinician feedback, and high-impact publications in core bioinformatics journals.