Kanygin A.V. Interpretation of BERT Internal Representations for the Classification of Texts Containing Reasoning

https://doi.org/10.15688/mpcm.jvolsu.2026.1.4

Alexander V. Kanygin
Postgraduate Student, Department of Computer Sciences and Experimental Mathematics, Volgograd State University
This email address is being protected from spambots. You need JavaScript enabled to view it. ,

Prosp. Universitetsky, 100, 400062 Volgograd, Russian Federation

 

Abstract. This article examines the problem of binary text classification for the presence of reasoning (logical connections, argumentation, and cause-andeffect relationships). This problem was solved using a fine-tuned ruBert model, achieving significantly higher performance than classical ML approaches (TF-IDF, model stacking, static embedding vectors, and an LSTM classifier). The goal of the article is not so much to demonstrate the quality of the solution using vectorization via the BERT model, but rather to interpret BERT’s internal representations, allowing to understand which features and/or structural elements of the text determine the model’s high performance. The article utilizes a combination of methods for analysis: visualization of [CLS] vectors using 2D dimensionality reduction via UMAP, silhouette score calculation for each of the BERT model’s 12 layers, separate classifiers for all hidden layers of the model, and an ablation study, which evaluates how removing key structures that define reasoning in the text affects the model’s confidence. The results show a significant drop in model confidence when removing key structures/markers of causal relationships and logical connectives. Furthermore, using the silhouette score metric, the model’s ability to distinguish between classes improves as it moves toward higher layers of the model. Taken together, these observations indirectly demonstrate that BERT generates text representations based on the logical structure of the text, not just the importance of individual tokens. This paper presents indirect but consistent evidence that the [CLS] vectors generated by the BERT model capture the logical structure of reasoning in the text, making it an effective tool for specific, non-trivial text analysis tasks. Furthermore, the results highlight the need for further study of Russian-language language models.

Key words: machine learning, interpretability of transformers, BERT, text classification, argumentation, hidden representations, fine-tuning, natural language processing.

Creative Commons License
Kanygin A.V. Interpretation of BERT Internal Representations for the Classification of Texts Containing Reasoning is licensed under a Creative Commons Attribution 4.0 International License.

Citation in English: Mathematical Physics and Computer Simulation. Vol. 29 No. 1 2026, pp. 38-57

Attachments:
Download this file (Kanygin.pdf) Kanygin.pdf
URL: https://mp.jvolsu.com/index.php/en/component/attachments/download/1452
100 Downloads