Natural language processing encompasses various tasks: information retrieval, sentiment analysis, machine translation, building question‑answering systems, emotion detection, and typo correction. The most modern approach to solving these tasks is the use of machine learning methods: classical algorithms (logistic regression, random forest), recurrent neural networks (long short‑term memory, gated recurrent unit), and models based on the Transformer architecture (autoencoder and generative models). Current research in text processing necessarily includes a comparison of several methods. Large language models are expected to demonstrate higher quality than classical machine learning algorithms. However, as quality increases, so does the required amount of computing power. The choice of the optimal method for solving a practical problem depends on many factors, such as the problem formulation, evaluation metrics, availability of computing resources, and the volume of training data.
Thus, the ability to comprehensively analyze a task and select the most suitable machine learning method is an indispensable skill for a specialist in automatic text processing.
The goal of the course is to develop practical skills in applying machine learning to text processing tasks. It consists of two modules, each of which will thoroughly examine a specific computational linguistics task. The syntactic module focuses on the automatic assessment of sentence acceptability, while the semantic module deals with automatic emotion recognition in text. We will analyze how these tasks are addressed in existing research, review the datasets and methods proposed in the papers. At the end of each module, students will need to complete a project assignment: propose and apply a new solution method for one of the studies discussed.
Topics covered in the course
- Tasks of automatic text analysis
- Machine learning algorithms for text processing
- Classification of sentences by acceptability
- Correlation between acceptability and probability
- Targeted syntactic evaluation of language models
- Probing language models
- Emotion classification in text
- Emotion recognition in the context of dialogue
- Emotion classification for the Russian language
- Text generation with emotional tone
The course programme includes 16 sessions (2 academic hours each): 8 lectures and 8 seminars
Requirements for students:
- knowledge of automatic text processing;
- Python programming;
- linear algebra and mathematical analysis.
Занятия проводятся в ауд. 953 (1 ГУМ) на филологическом факультете МГУ им. М. В. Ломоносова
В программе курса 16 занятий (по 2 ак/ч): 8 лекций и 8 семинаров
Формат проведения: возможно как очное, так и дистанционное участие
Старт курса: 13 сентября
Занятия будут проходить 1 раз в неделю по средам в 9:00
Записаться на курс и задать вопросы можно по почте xeanst@gmail.com (Ксения Андреевна Студеникина)
Набор на курс 2023 года закрыт