A Pretrained Model-Based Approach to Characterisation of Functional Heterogeneity in Cancer
Context & Background
Human DNA contains roughly 3 billion base pairs. Predicting which genomic variations (particularly in non-coding regions) drive cancer is a major challenge. In genomics, generating patient sequence data is expensive, but massive amounts of data are centrally archived. Pretrained models can capture universal genomic features that can be repurposed to study functional heterogeneity in tumors, especially under-represented Indian patient cohorts.
Problems to be Addressed
Most genomic analysis models focus solely on coding mutations, which represent a tiny fraction of human cancers. In clinical settings, the unavailability of matched normal samples hinders the identification of tumor-specific somatic mutations. There is a lack of representation of gallbladder cancer in international databases like TCGA, despite India having some of the highest case rates globally.
Aims and Objectives
1. Learn universal embeddings of genomic alterations using large-scale exome databases.
2. Apply transfer learning to single-cell gene expression profiles.
3. Delineate cellular heterogeneity in gallbladder cancer tumors.
Methodology
The project uses word-embedding methodologies (like Skipgram) from NLP to represent mutations as vectors (ASSM - Aminoacid Switch Sequence Model) trained on ~60,000 exomes. Deep LSTM architectures classify variants into cancer-related categories. Single-cell RNA sequencing is then performed on clinical gallbladder cancer tissues to map tumor subtypes.
Expected Outcomes
A database of gallbladder cancer gene expression profiles, computational tools for predicting somatic mutations without normal matches, and commercial partnerships with genomics firms (MedGenome, Strand Life Sciences).