<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with OASIS Tables with MathML3 v1.4 20241031//EN" "https://jats.nlm.nih.gov/archiving/1.4/JATS-archive-oasis-article1-4-mathml3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" dtd-version="1.4" article-type="research-article" xml:lang="en"><front><journal-meta><journal-title-group><journal-title xml:lang="ru">Математическая физика и компьютерное моделирование</journal-title></journal-title-group><issn publication-format="print">2587-6325</issn><issn publication-format="electronic">2587-6902</issn></journal-meta><article-meta><article-id pub-id-type="doi">10.15688/mpcm.jvolsu.2025.1.3</article-id><article-categories><subj-group><subject>Other</subject></subj-group></article-categories><title-group><article-title xml:lang="ru">Построение модели для решения задачи классификации рассудительного текста.</article-title><trans-title-group xml:lang="en"><trans-title>Construction of a Model for the Task of Reasoning Text Classification</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Каныгин</surname><given-names>Александр Владимирович</given-names></name><name xml:lang="en"><surname>Kanygin</surname><given-names>Alexander V.</given-names></name></name-alternatives><xref ref-type="aff" rid="aff1"/></contrib><aff-alternatives id="aff1"><aff xml:lang="en"><institution>Volgograd State University (Volgograd, Russian Federation)</institution></aff><aff xml:lang="ru"><institution>Волгоградский государственный университет (Волгоград, Российская Федерация)</institution></aff></aff-alternatives></contrib-group><pub-date pub-type="epub" iso-8601-date="2025-05-29"><day>29</day><month>05</month><year>2025</year></pub-date><volume>28</volume><issue>1</issue><fpage>27</fpage><lpage>39</lpage><history><date date-type="received" iso-8601-date="2025-02-24"><day>24</day><month>02</month><year>2025</year></date><date date-type="accepted" iso-8601-date="2025-03-13"><day>13</day><month>03</month><year>2025</year></date></history><permissions><license xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:title="CC BY 4.0"><ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p xml:lang="ru">CC BY 4.0</license-p></license></permissions><abstract xml:lang="ru"><p>В статье рассмотрена задача классификации текстов на предмет наличия в них рассуждений (логических связок, аргументации, причинно-следственных отношений). Цель исследования — разработать метод, позволяющий с высокой точностью определять «рассудительный» характер фрагмента текста, используя современные алгоритмы машинного обучения. Особое внимание уделено ансамблевому подходу на основе стекинга: в качестве базовых классификаторов рассматриваются сильные модели (CatBoost, XGBoost, Random Forest и т. п.), а роль мета-модели выполняет логистическая регрессия. Для обоснования выбора стекинга приводятся результаты сравнительного анализа более десяти популярных алгоритмов (Logistic Regression, SVC, Random Forest, CatBoost, XGBoost и др.) по показателям Accuracy, Precision, Recall, F1-score, ROC AUC, PR AUC. Основные этапы исследования включают генерацию и разметку обучающего набора данных, предварительную обработку текстов (токенизацию, лемматизацию, исключение стоп-слов), векторизацию признаков (TF-IDF) и экспериментальное сравнение моделей на контрольной выборке. Предложенная модель стекинга показала лучшие результаты по совокупности метрик, что позволило повысить точность классификации рассудительных текстов до уровня F1, равного 0,905, при ROC AUC, равному 0,887. В заключении обсуждаются перспективы применения описанного подхода для текстов разной длины и стиля, а также потенциальные методы дальнейшего улучшения качества классификации.</p></abstract><abstract xml:lang="en" abstract-type="summary"><p>The article addresses the task of classifying texts for the presence of reasoning (logical links, argumentation, cause-and-effect relationships). The aim of the study is to develop a method that allows for highly accurate determination of the “reasoning” nature of a text fragment using modern machine learning algorithms. Particular attention is paid to an ensemble approach based on stacking: strong models (XGBoost, CatBoost, Random Forest, etc.) are considered as base classifiers, while logistic regression serves as the meta-model. To justify the choice of stacking, we present the results of a comparative analysis of more than ten popular algorithms (Logistic Regression, SVC, Random Forest, CatBoost, XGBoost, etc.) by Accuracy, Precision, Recall, F1-score, ROC AUC, and PR AUC. The main stages of the study include the generation and annotation of the training dataset, preliminary text processing (tokenization, lemmatization, stop-word removal), feature vectorization (TF-IDF), and experimental comparison of the models on a control sample. The proposed stacking model showed the best overall performance across all metrics, enabling us to increase the accuracy of reasoning text classification to F1 equal to 0.905 at ROC AUC equal to 0.887.</p></abstract><kwd-group xml:lang="ru"><kwd>машинное обучение</kwd><kwd>ансамблевые методы</kwd><kwd>стекинг</kwd><kwd>TF-IDF</kwd><kwd>аргументация</kwd><kwd>анализ текстовых данных</kwd></kwd-group><kwd-group xml:lang="en"><kwd>machine learning</kwd><kwd>ensemble methods</kwd><kwd>stacking</kwd><kwd>TF-IDF</kwd><kwd>argumentation</kwd><kwd>text processing</kwd></kwd-group></article-meta></front><back><ref-list><ref id="ref1"><mixed-citation publication-type="other" xml:lang="ru">Клячин, В. А. Атрибуция медийных текстов на основе обученной модели естественного языка и лингвистическая оценка качества идентификации / В. А. Клячин, Е. В. Хижнякова // Вестник Волгоградского государственного университета. Серия 2. Языкознание. - 2024. - Т. 23, № 5. - C. 31-46. -. DOI: 10.15688/jvolsu2.2024.5.3 EDN: TEFUFD</mixed-citation></ref><ref id="ref2"><mixed-citation publication-type="other" xml:lang="ru">ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations / Zh. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, R. Soricut // arXiv preprint arXiv:1909.11942. - 2020. - P. 1-17. -. DOI: 10.48550/arXiv.1909.11942</mixed-citation></ref><ref id="ref3"><mixed-citation publication-type="other" xml:lang="ru">BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding / J. Devlin, M.-W. Chang, K. Lee, K. Toutanova // arXiv preprint arXiv:1810.04805. - 2019. - P. 1-16. -. DOI: 10.48550/arXiv.1810.04805</mixed-citation></ref><ref id="ref4"><mixed-citation publication-type="other" xml:lang="ru">DeBERTa: Decoding-Enhanced BERT with Disentangled Attention / P. He, X. Liu, J. Gao, W. Chen // arXiv preprint arXiv:2006.03654. - 2021. - P. 1-23. -. DOI: 10.48550/arXiv.2006.03654</mixed-citation></ref><ref id="ref5"><mixed-citation publication-type="other" xml:lang="ru">RoBERTa: A Robustly Optimized BERT Pretraining Approach / Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov // arXiv preprint arXiv:1907.11692. - 2019. - P. 1-13. -. DOI: 10.48550/arXiv.1907.11692</mixed-citation></ref><ref id="ref6"><mixed-citation publication-type="other" xml:lang="ru">XLNet: Generalized Autoregressive Pretraining for Language Understanding / Zh. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, Q. V. Le // arXiv preprint arXiv:1906.08237. - 2020. - P. 1-18. -. DOI: 10.48550/arXiv.1906.08237</mixed-citation></ref><ref id="ref7"><mixed-citation publication-type="other" xml:lang="en">Klyachin V.A., Khizhnyakova E.V. Atributsiya mediynykh tekstov na osnove obuchennoy modeli estestvennogo yazyka i lingvisticheskaya otsenka kachestva identifikatsii [Attribution of Media Texts Based on a Trained Natural Language Model and Linguistic Assessment of Identification Quality]. Vestnik Volgogradskogo gosudarstvennogo universiteta. Seriya 2. Yazykoznanie [Science Journal of Volgograd State University. Linguistics], 2024, vol. 23, no. 5, pp. 31-46. DOI:https://doi.org/10.15688/jvolsu2.2024.5.3</mixed-citation></ref><ref id="ref8"><mixed-citation publication-type="other" xml:lang="en">Lan Zh., Chen M., Goodman S., Gimpel K., Sharma P., Soricut R. ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations. arXiv preprint arXiv:1909.11942, 2020, pp. 1-17. DOI:https://doi.org/10.48550/arXiv.1909.11942</mixed-citation></ref><ref id="ref9"><mixed-citation publication-type="other" xml:lang="en">Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. arXiv preprint arXiv:1810.04805, 2019, pp. 1-16. DOI:https://doi.org/10.48550/arXiv.1810.04805</mixed-citation></ref><ref id="ref10"><mixed-citation publication-type="other" xml:lang="en">He P., Liu X., Gao J., Chen W. DeBERTa: Decoding-Enhanced BERT with Disentangled Attention. arXiv preprint arXiv:2006.03654, 2021, pp. 1-23. DOI:https://doi.org/10.48550/arXiv.2006.03654</mixed-citation></ref><ref id="ref11"><mixed-citation publication-type="other" xml:lang="en">Liu Y., Ott M., Goyal N., Du J., Joshi M., Chen D., Levy O., Lewis M., Zettlemoyer L., Stoyanov V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692, 2019, pp. 1-13. DOI:https://doi.org/10.48550/arXiv.1907.11692</mixed-citation></ref><ref id="ref12"><mixed-citation publication-type="other" xml:lang="en">Yang Zh., Dai Z., Yang Y., Carbonell J., Salakhutdinov R., Le Q.V. XLNet: Generalized Autoregressive Pretraining for Language Understanding. arXiv preprt arXiv:1906.08237, 2020, pp. 1-18. DOI:https://doi.org/10.48550/arXiv.1906.08237</mixed-citation></ref></ref-list></back></article>
