About Workshop
This 3-day workshop introduces participants to the emerging role of an AI Scientist in computational biology and medicine. The workshop focuses on how mathematical modeling, biological data engineering, machine learning, transformer-based sequence models, and interactive AI applications are used to analyze genomic, protein, and biomedical datasets. Through short lectures and guided hands-on activities, participants will understand how AI can support disease risk prediction, biological data analysis, rare anomaly detection, drug-target workflows, and real-time biomedical dashboards.
Aim
To equip students, researchers, academicians, and professionals with foundational and applied skills in AI-driven computational biology and medicine, including biological data handling, mathematical modeling, large-scale data workflows, sequence-based AI models, and biomedical AI application development.
What Participants Will Learn
- Understand how AI, statistics, and computational biology are connected in modern biomedical research.
- Learn how probability, likelihood estimation, conjugate priors, SVD, and eigenvalues support biological data interpretation.
- Access, clean, and prepare genomic and protein datasets from public biological databases such as NCBI and EBI.
- Explore large-scale biomedical data engineering using tools such as Apache Arrow, Polars, and Dask.
- Understand how Python and R can be combined for machine learning and statistical analysis workflows.
- Learn how transformer-based models are used for DNA, RNA, protein, and biomedical sequence analysis.
- Understand model fine-tuning, rare disease detection, imbalanced data handling, and Focal Loss.
- Build awareness of AI ethics, data privacy, bias, and future trends in autonomous biomedical research.
- Develop practical exposure to biomedical AI pipelines, dashboards, and real-time prediction systems.
Structure
```html
📅 Day 1: Foundational Data Setup
- Focus: Building the core math foundations and setting up biological datasets.
- Learning how to use conjugate priors and likelihood estimation to predict disease risks and variant changes.
- Using SVD and eigenvalues to reduce high-dimensional biological data into simple, visual groups.
- Writing custom, fast mathematical code using hardware acceleration to understand how AI models learn across a 3D loss surface.
- Accessing, downloading, and cleaning raw genomic and protein data from online databases such as NCBI and EBI.
🛠️ Hands-on:
- Write a Python script using PyMC to calculate disease risk probabilities with clear uncertainty boundaries.
- Build a basic SVD matrix system from scratch using NumPy to sort and group complex genetic data.
📅 Day 2: Data Engineering & Large-Scale Workflows
- Focus: Managing massive datasets and linking different coding tools without crashing your computer’s memory.
- Using Apache Arrow to read, filter, and clean huge data files without overloading your computer’s RAM.
- Setting up a workflow where Python handles the machine learning code and R handles the deep statistical analysis in the same environment.
- Handling heavy genomic datasets by splitting the work across distributed computing networks.
- Designing workflow paths to match candidate drugs with specific biological targets.
🛠️ Hands-on:
- Use Polars and Dask to run sorting and aggregation filters over a large-scale public health dataset.
- Run speed and memory benchmarks to find the fastest way to join two large biological data tables.
📅 Day 3: Sequence Models, Automation & Applications
- Focus: Fine-tuning transformer models, handling rare data, and deploying visual AI applications.
- Transformers for Biology: Understanding how deep learning architectures read DNA and protein sequences like sentences, including tools such as AlphaFold.
- Model Fine-Tuning: Customizing pre-trained language and sequence models on small, specific biological datasets using PyTorch Lightning.
- Managing Unbalanced Data: Using specialized loss functions such as Focal Loss to train AI models for rare diseases and anomaly detection.
- Live Dashboards: Creating interactive web apps that display model predictions, data trends, and feature importance in real time.
- Ethics & Future Trends: Discussing AI bias in health metrics, data privacy, and how autonomous laboratories generate new scientific hypotheses.
🛠️ Hands-on:
- Deploy an automated AI pipeline designed to flag rare anomalies in highly imbalanced medical datasets.
- Build a live-updating web app using Plotly Dash that displays automated predictions and structural metrics instantly.
Important Dates
Registration Ends
7:00 PM IST
Workshop Dates
2026-07-15
8:00 PM IST
8:00 PM IST
What You Will Gain

Outcomes
- Explain how AI is transforming computational biology, genomics, drug discovery, and medicine.
- Apply basic probabilistic modeling concepts to biomedical risk prediction problems.
- Use dimensionality reduction techniques to simplify and visualize high-dimensional biological data.
- Access and clean biological datasets from public repositories.
- Understand scalable data engineering workflows for large genomic and health datasets.
- Compare memory and speed performance across different data processing tools.
- Understand how transformer models are applied to DNA, RNA, protein, and biomedical sequence data.
- Build basic AI workflows for rare disease detection and anomaly identification.
- Create simple interactive dashboards for biomedical AI predictions.
- Recognize ethical, privacy, and bias-related challenges in AI-powered healthcare research.
- Identify future research and career opportunities in biomedical AI and computational biology.
Who Should Attend
- Graduate students pursuing B.Sc., B.Tech, B.E., B.Pharm, MBBS, BDS, or related degrees
- Postgraduate students pursuing M.Sc., M.Tech, M.E., M.Pharm, MPH, MBA Healthcare, or related programs
- PhD scholars and research fellows working in biology, biotechnology, bioinformatics, computational biology, AI, ML, data science, healthcare, medicine, pharmacy, or engineering
- Academicians, faculty members, trainers, and educators interested in AI applications in computational biology and medicine
- Industry professionals from biotechnology, pharmaceutical, healthcare, diagnostics, biomedical devices, medical AI, data science, and drug discovery sectors
- Early-career researchers and professionals aiming to build skills in biomedical AI, genomics data science, and computational medicine
Prof. Kumud Malhotra
Department of Biotechnology
Speciality: Python Programming for Biomedical AI, Biological Data Handling, Bioinformatics Data Analysis, Probabilistic Modeling, Dimensionality Reduction
