7) Text vectorization based on statistics for document classification and clustering within a text collection. The tf‑idf metric. Measures of text similarity. Machine learning for classification and clustering of texts and textual objects, methods for quality assessment.
8) Machine translation (MT) as a key applied problem: strategies and generations of systems. Statistical MT technologies. Application of the seq2seq neural network model, emergence of the Transformer architecture. Quality evaluation of MT systems.
9) Recurrent neural networks for building contextualized word vector representations and their limitations. Neural network language models based on the Transformer architecture. The BERT encoder model: specifics of training and application.
10) Automatic generation of document texts: strategies and capabilities. Template‑ and rule‑based text generation. Generative neural network models of the GPT family and their use for text generation.
11) Information extraction from texts: approaches and types of extracted data. Solving the task as a sequence labeling problem using machine learning. Specific features of sentiment analysis, aspect‑based opinion analysis.
12) Automatic summarization and annotation of documents. Types of annotations and summaries, strategies for their construction: extraction and abstraction. Statistical extraction methods. Abstraction based on neural network models.
13) Question‑answering systems and conversational agents (chatbots). Types of question‑answering systems, strategies for question analysis, dialogue management, and answer generation. Methods for building chatbots. ChatGPT: specifics of training and application.