The winner of the 2024 Young Scientists Contest of the “Intellect” Foundation, Andrey Bykov, a PhD student at the Faculty of Chemistry, Lomonosov Moscow State University, took part in the 12th National Crystallochemical Conference, which was held at the Kola Science Centre of the Russian Academy of Sciences in Apatity. The conference is considered one of the leading Russian forums in the field of crystallochemistry, and this year its programme was expanded to include topics on machine learning and in silico prediction of material properties.

At the conference, Andrey presented a report titled “Machine Learning and a Data‑Driven Approach in the Chemistry of Hybrid Halogenometallates: From Dataset and Database Construction to Structure Prediction and Rational Material Design” and was awarded a diploma for the best poster presentation.
In the interview, he talks about how machine learning helps chemists move away from trial‑and‑error methods towards a rational search for new materials, as well as why data quality is becoming a key factor for modern science.
– Andrey, you presented a report at the conference on machine learning and the data‑driven approach in the chemistry of hybrid halogenometallates. What is your research about, if you could explain it in simple terms?
– The research aims to develop approaches for the rational design of hybrid halogenometallates, i.e., to find ways to purposefully synthesise organic–inorganic halide complexes of post‑transition metals with specified physicochemical properties. In my report, I presented solutions to problems related to determining the ability of cations to form a specific type of halogenometallate anion and establishing the key “composition–structure–property” relationship crucial for materials science.
To explain it in simple terms, I usually use an analogy with a construction set: organic cations and the building blocks of halogenometallate anions can be imagined as Lego pieces. The problem is that it is not yet fully understood by what rules they fit together. My research task, figuratively speaking, is to determine the shape of these pieces and thus understand how they can be combined to assemble a crystal structure with the desired properties.
The results of this work can be applied in various fields: in the development of optoelectronic materials, including materials for sensors, photodetectors and LEDs; in creating light‑absorbing components for solar cells; as well as in the chemical industry and applied research.
– How does machine learning help scientists search for and create new materials?
– Machine learning and the data‑driven approach in general are powerful tools for analysing large volumes of data and identifying patterns. Their key feature, compared to more conventional methods used in chemistry, is that they allow researchers to consider the influence of a large number of factors on the relationships under study simultaneously, rather than separately. This makes it possible to establish a hierarchy of such factors and identify their optimal combinations to achieve the target functional characteristics of materials.
In my opinion, the main advantage is the transition from the experimental search for chemical compounds via trial and error to the rational selection of a much smaller number of promising candidates for new materials. This significantly reduces both time and resource costs.
– Why, in your view, is it important today to combine chemistry, data handling and artificial intelligence?
– I wouldn’t say that such a combination is exceptionally important. Rather, research at the intersection of chemistry and data science has become truly feasible and in demand today thanks to the emergence of the big data paradigm, the rapid development of AI and the accumulation of substantial volumes of chemical information — albeit still more modest compared to some other types of data.
To a greater extent, this combination is driven by the current pace of scientific development and the need to obtain results within a relatively short timeframe given limited resources. Big data methods allow us to see a more comprehensive picture of fundamental patterns, rather than just isolated snapshots or relationships within a limited sample of objects.
At the same time, this approach also has a downside: its capabilities depend not only on the volume but also on the quality of the data. And data quality can deteriorate if the focus is primarily on quantity.
– Your report received high praise from the experts and the jury. What do you think particularly interested the experts in your work?
– Probably, it was the combination of traditional chemistry with data‑handling approaches that are relatively new to this field, as well as the scale and complexity of the tasks. In addition, the progress in moving from descriptive to predictive crystallochemistry might have drawn attention. I was very pleased to be able to discuss and get feedback from colleagues on the issues of relevance, correctness and distortion of the data presented in the literature.
– Please share your professional plans. How do you plan to further develop your topic?
– First of all, I plan to complete the publication of the obtained results. In the long term, I intend to carry out a larger‑scale experimental validation of the model predictions and the developed approaches. Ultimately, in addition to synthesising halogenometallates with target physicochemical properties, I plan to test them as materials in prototypes of optoelectronic devices and determine a broader range of their performance characteristics. I also intend to address the limitations identified during the research — to find a way to validate a large volume of literature data and try to highlight this issue in the scientific literature.