Ilya Galushko: "Most of the "pure humanities" are able to master the key concepts of higher mathematics and appreciate its beauty"

12.05.2026

Ilya Galushko is a graduate student at the Faculty of History of Lomonosov Moscow State University, winner of the MSU-2024 Young Scientists Competition. In 2024, he completed the program "Neural Networks and their Application in Scientific Research", which was developed with the support of the Intellect Foundation, and then became one of the winners of the scientific publications competition 5 the course flow.

We are watching with interest how the life and scientific trajectory of graduates of our programs is developing, and we asked Ilya about what he is doing today and how he uses the acquired knowledge.

– Ilya, please tell us about yourself — what are you doing now, what are you working on?

– I am completing my postgraduate studies at the Faculty of History of Moscow State University, where I studied from the first year of my bachelor's degree. In parallel with the work on the grant project, I am completing my PhD thesis on the topic "Regulation of the stock market of the Russian Empire at the beginning of the 20th century: sources and research methods" by Leonid Iosifovich Borodkin. I think that my life was determined by my admission to the Department of Historical Informatics and the scientific leadership of Leonid Iosifovich. I have always been interested in socio-economic history and the history of the stock market in particular, but here I began to study all this in a specialized way.

As I immersed myself in the topic, I became interested in mathematical modeling and econometrics (there were cathedral courses on this topic), which gradually led me to programming. I studied by myself, went to inter-faculty courses. Then I was able to enroll in the DPO program from the Intellect Foundation "Neural Networks and their application in scientific research". There I studied specialized programming, mathematics and machine learning in its various forms (image recognition, NLP, modeling based on tabular data).

Right now, my main academic interest is related to this, and I'm trying to apply it meaningfully in my PhD thesis. In a broad sense, I try to do more applied research – I see great potential in the application of ML methods in historical science. By the way, taking this opportunity, I would like to say a huge thank you again to the teaching team of this course – it was an excellent education.

In 2024, I became the winner of the young scientists competition from the Intellect Foundation and received a two-year grant to conduct research on "Automating the process of annotating archival documents using Large Language Models (LLM)." My research is devoted to exploring the possibilities of using large language models in archival work. Then I published a short article analyzing the attention scales of models in the context of historical tasks, and now I am completing a large block of work related to the development of a system for using neural networks for post-correction of optical recognition of historical documents.

– When and how did you become interested in neural networks? How did you come up with the idea to go learn this and work with them?

– My fascination with socio-economic history led me to this. Specialized journals (such as, for example, Cliometrica) are filled with econometric models and classical ML (clustering is especially popular). There are even more simulations in the history of the stock market: This is the study of time series, the analysis of historical reporting, the assessment of volatility of securities, trends, correlations. I think we can generally say that specialization in this field is impossible without mastering econometrics. And here it is already quite easy to make the transition from interface statistical packages to programming, and then to neural networks. Although, interestingly, neural networks are rarely used in tabular data analysis, classical ML (especially decision trees) it turns out to be more efficient there.

– Your main specialty is humanities. Was it difficult to study? What was the most difficult/interesting/unexpected part of the learning process?

– The training was really difficult. I had big gaps in mathematics, and, as it turned out later, they gave us a lot at school at the level of memorizing the rules, so I had to form a holistic view on my own. But it was extremely interesting, and today you can find hundreds of high-quality educational videos online. With special gratitude, I recall the 3Blue1Brown courses. In principle, I am sure that most "pure humanities" are actually at least quite capable of mastering the key concepts of higher mathematics. And appreciate the beauty of this science. The issue is the correct presentation of the material.

– Tell us about your project. How did you come up with the idea, how did you work on it, were there any difficulties, and what additional resources were needed?

– The idea of my project arose from practical needs – I myself work a lot in the archives for the PhD program. And this work is extremely complicated by the fact that many archival collections do not have annotations on the website – apart from the name of the case, we have no clear idea what kind of materials are inside – will they be useful for our research? What organizations and historical figures are mentioned there? What data can be collected from there? I'll find out all this when the case comes to hand. And the number of cases in the order is limited, and you need to wait for them for several days. Globally, my project is devoted to a comprehensive study of the prospects for using neural networks in archival work, in particular, for streaming creation of archival annotations.

Optimization of optical recognition (OCR) scans of historical documents has now become one of the main directions of my project. Another problem is already being solved here – many electronic collections displayed on the websites of electronic libraries and archives contain only scans, which prevents the use of automatic text search by keywords. And that would be extremely convenient. But modern recognition software is still far from ideal, and some characters are almost always recognized incorrectly. And I am solving this problem as part of the current stage of the project.

  – Каковы Ваши научные планы на ближайшее будущее?

– What are your scientific plans for the near future?

– As part of my PhD thesis, I am completing a key chapter in which event analysis is applied - I am trying to statistically assess how the actions of the State Bank of the Russian Empire influenced the dynamics of stock prices of industrial companies (and whether this influence is observed at all). In general, I would like to work more with event-study in the framework of socio-economic history – an empirical assessment of the impact of events on the dynamics of time series looks extremely tempting for many historical tasks. As part of my grant work, I am currently conducting a social survey in which respondents are asked to assess how likely it is that the description of a historical document they provided was generated by AI. By the way, I invite everyone to take part!

– Please tell us, where is the scientific frontier in your field now?

– I think that one of the most important problems remains the explainability of models. And in the context of history, the question of how and to what extent models understand the temporal context remains vague; how can we influence this understanding in the process of primary education or subsequent further education. If we talk about the use of classical ML, then a whole expanse opens up here: ML models are still rarely used in socio-economic history. But there is, for example, a fairly obvious approach in which we train a model to predict the dynamics of a social process - for example, we teach a random forest to predict the results of parliamentary elections at the level of individual regions; as part of this task, we try to get the best quality, and then we look at the importance of features and try to interpret them in the context of the era.. Few similar studies show that this can be very productive (I emphasize that there are few studies specifically with ML models, the same logistic regressions are used quite often in a similar context). Certain hopes are pinned on the linguistic modeling of large narratives. In the last few years, many research teams have been trying to propose an effective approach to modeling discourses using text embeddings. And I must say that these attempts arouse reasonable skepticism from purely humanitarian specialists, but this controversy itself is also, of course, productive for both NLP and the methodology of humanitarian knowledge.

Educational programs are implemented with the support of the Oleg Deripaska Foundation "Volnoe Delo".