The goal of the “Machine Learning in Biology” course is to provide a gentle introduction to machine learning, neural networks, and their applications in biology. The course is designed for bachelor’s and master’s students from natural science faculties at Lomonosov Moscow State University. During the course, students will become familiar with the mathematical foundations of machine learning and will get hands‑on experience with real‑world problems.
During the course, students will:
All practical tasks in the course will be tailored to biological applications. Successful examples of machine learning in biology will be discussed, such as AlphaFold2, U‑Net, SpliceAI, iNaturalist, the generation of new antibiotics and anticancer drugs, and more.
Topics covered in the course:
Module 1. Fundamentals of Machine Learning
1. Machine learning. Types of tasks, examples. Machine learning in biology. Major breakthroughs: AlphaFold 2, SpliceAI, DeepVariant, U‑Net. Working with NumPy, Pandas, and Matplotlib libraries.
Software: Python, Pandas, NumPy, Seaborn + Matplotlib.
2. k‑Nearest Neighbors method. Its application when a good feature representation is available. Classification task. Accuracy metric. Overfitting, underfitting. Splitting into train and test sets.
Training on genomic embeddings. Working with compounds in Python. Classifying compounds and issues related to splitting them into training and test sets.
3. Confusion matrix. F‑score. Integral metrics for classification performance. Regression task. Linear regression. Regularization. Ridge regression. Lasso. ElasticNet regularization. Logistic regression.
4. Hyperparameters. Why they must not be tuned on the test set. Cross‑validation. Hyperparameter tuning methods: GridSearch, RandomGridSearch, Optuna. The challenge of correctly splitting biological data (TargetFinder — leakage problem).
Working with biological sequences. Classifying harmful vs. neutral mutations in non‑coding regions.
5. The curse of dimensionality. Dimensionality reduction methods. PCA. Choosing the number of components. Stochastic dimensionality reduction methods: t‑SNE, UMAP. Challenges of dimensionality reduction. Batch effect. Visualizing transcriptomic data, particularly single‑cell data. Learning embeddings, semi‑supervised UMAP, PLS‑DA. Issues with the approach. Using representations from trained models as a new feature space.
Visualizing transcriptomic data. Analyzing data, identifying outliers and batch effects. Visualizing research papers.
6. Clustering. K‑Means, DBSCAN, Affinity Propagation, hierarchical clustering. Clustering biological objects (structures, drugs).
Clustering chemical compounds. Classifying virus sequences.
Module 2. Tree‑based Models and Boosting as State‑of‑the‑Art Methods
7. Building decision rules using EDA. Decision trees. Categorical features and how to handle them (label encoding, one‑hot encoding, mean encoding). Categorical features in biology. Random Forest. The importance of Random Forest in biology. Disease diagnosis using Random Forest.
Distinguishing sick vs. healthy individuals using Random Forest. Encoding biological sequences.
8. Gradient descent. Gradient boosting. Modifications of gradient boosting: XGBoost, LightGBM, CatBoost. Predicting drug properties using gradient boosting. Assessing feature importance in decision trees (Gini impurity and mean accuracy decrease). Boruta. Permutation method. SHAP. Analyzing functional groups important for inhibiting a given enzyme using SHAP.
Analyzing the importance of parts of a chemical compound in interactions. Selecting genes important for breast cancer tumor development.
Module 3. Neural Networks
9. Neural networks. Introduction. PyTorch framework. Stochastic gradient descent (SGD). Backpropagation. Chain rule for differentiating complex functions. Batch. Dataset. Data loader. Optimizer. Loss function. Applying multilayer neural networks to biological data.
Using torchANI. Predicting diseases based on expression data.
10. Convolutional neural networks. Applications in biology. Diagnosing diseases caused by genome mutations. Predicting the effects of single‑nucleotide polymorphisms. Predicting ligand–protein binding energy.
Classifying medical images using convolutional neural networks. Transfer learning for biological images. Predicting chromatin accessibility using convolutional neural networks.
11. Autoencoders. Representation learning. U‑Net. Segmenting biological images. Handling noisy biological data with neural networks. Predicting harmful mutations in coding regions using VAE.
Separating sick and healthy tissues in latent space. Working with medical images.
12. Methods for training deep neural networks. Activation functions. Weight initialization. Batch normalization. Optimizers. Learning rate tuning. LR range test. Learning schedules. One‑cycle learning schedule.
Classifying necrosis and tumors on histological images.
13. Recurrent neural networks. Predicting protein secondary structure. GANs. Drug design. Denoising diffusion models. Stable Diffusion.
Predicting splice sites (data from Kaggle). Predicting RNA secondary structure stability based on sequence.
14. Attention mechanism. Attention transformers. GPT‑3. BERT. Analyzing biological texts using BioBERT.
Text mining of biological articles. Protein structure modeling. Unsupervised segmentation of biological images. Enformer.
15. Representation learning. Self‑supervised learning. ESM1. AlphaFold2. DINO. Graph neural networks. Message passing. PyTorch Geometric framework. Creating custom datasets. Node classification. Graph classification. Applying graph neural networks to analyze protein–disease interaction graphs.
Predicting the role of proteins in disease development based on graph data.
The course program includes 30 sessions: 15 lectures and 15 practical sessions
Requirements for students:
Teachers:
Arseniy O. Zinkevich
Education
Bioengineering and bioinformatics
MSU
Areas of Expertise
Bioinformatics, NGS data analysis, identification of DNA motifs
Dmitry D. Penzar
Education
Bioengineering and bioinformatics
MSU
Areas of Expertise
Bioinformatics, statistical data processing, machine learning in biology
Andrey I. Sigorskikh
Education
Bioengineering and bioinformatics
MSU
Areas of Expertise
Identification of drug targets, non-coding RNA research, machine learning in biology
Занятия проводятся на Факультете Биоинженерии и биоинформатики.
В программе курса 30 занятий: 15 лекций и 15 практикумов
Старт курса: с 15 сентября
Занятия будут проходить по четвергам:
Требования к студентам:
Набор на курс 2022 года закрыт
FRIDAYS
09:00–10:30 — period 1
10:40–12:10 — period 2
12:10–13.00 — break
13:00–14:30 — period 3
14:40–16:10 — period 4
Instructor: A. Ganichev
THURSDAYS
09:00–10:30 — period 1
10:40–12:10 — period 2
12:10–13.00 — break
13:00–14:30 — period 3
14:40–16:10 — period 4
Instructor: A. Marakulin
MONDAYS
09:00–10:30 — period 1
10:40–12:10 — period 2
12:10–13.00 — break
13:00–14:30 — period 3
14:40–16:10 — period 4
Instructor: D. Penzar
SATURDAYS
09:00–10:30 — period 1
10:40–12:10 — period 2
12:10–13.00 — break
13:00–14:30 — period 3
14:40–16:10 — period 4
Instructor: I. Konyushok