<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with OASIS Tables with MathML3 v1.4 20241031//EN" "https://jats.nlm.nih.gov/archiving/1.4/JATS-archive-oasis-article1-4-mathml3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" dtd-version="1.4" article-type="research-article" xml:lang="en"><front><journal-meta><journal-title-group><journal-title xml:lang="ru">Математическая физика и компьютерное моделирование</journal-title></journal-title-group><issn publication-format="print">2587-6325</issn><issn publication-format="electronic">2587-6902</issn></journal-meta><article-meta><article-id pub-id-type="doi">10.15688/mpcm.jvolsu.2025.3.3</article-id><article-categories><subj-group><subject>Other</subject></subj-group></article-categories><title-group><article-title xml:lang="ru">О ВОЗМОЖНОСТИ ИСПОЛЬЗОВАНИЯ ИНДЕКСА ВИНЕРА ДЛЯ ВЫЧИСЛЕНИЯ ПРИЗНАКОВ ТЕКСТОВ НА ЕСТЕСТВЕННОМ ЯЗЫКЕ</article-title><trans-title-group xml:lang="en"><trans-title>ON THE POSSIBILITY OF USING THE WIENER INDEX TO CALCULATE FEATURES OF NATURAL LANGUAGE TEXTS</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Клячин</surname><given-names>Владимир Александрович</given-names></name><name xml:lang="en"><surname>Klyachin</surname><given-names>Vladimir</given-names></name></name-alternatives><xref ref-type="aff" rid="aff1"/><email>klyachin.va@volsu.ru</email><contrib-id contrib-id-type="orcid">0000-0003-1922-7849</contrib-id></contrib><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Хижнякова</surname><given-names>Екатерина Владимировна</given-names></name><name xml:lang="en"><surname>Khizhnyakova</surname><given-names>Ekaterina</given-names></name></name-alternatives><xref ref-type="aff" rid="aff1"/><email>yakovleva.e.v@volsu.ru</email><contrib-id contrib-id-type="orcid">0000-0002-7914-9988</contrib-id></contrib><aff-alternatives id="aff1"><aff xml:lang="en"><institution>Volgograd State University</institution></aff><aff xml:lang="ru"><institution>Волгоградский государственный университет</institution></aff></aff-alternatives></contrib-group><pub-date pub-type="epub" iso-8601-date="2025-10-14"><day>14</day><month>10</month><year>2025</year></pub-date><volume>28</volume><issue>3</issue><fpage>24</fpage><lpage>36</lpage><history><date date-type="received" iso-8601-date="2025-04-22"><day>22</day><month>04</month><year>2025</year></date><date date-type="accepted" iso-8601-date="2025-05-06"><day>06</day><month>05</month><year>2025</year></date></history><permissions><license xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:title="CC BY 4.0"><ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p xml:lang="ru">CC BY 4.0</license-p></license></permissions><self-uri xlink:href="https://mp.jvolsu.com/index.php/ru/archive-ru/520-mathematical-physics-and-computer-simulation-2025-vol-28-no-3/modelirovanie-informatika-i-upravlenie/1166-klyachin-v-a-khizhnyakova-e-v-o-vozmozhnosti-ispolzovaniya-indeksa-vinera-dlya-vychisleniya-priznakov-tekstov-na-estestvennom-yazyke" xlink:title="https://mp.jvolsu.com/index.php/ru/archive-ru/520-mathematical-physics-and-computer-simulation-2025-vol-28-no-3/modelirovanie-informatika-i-upravlenie/1166-klyachin-v-a-khizhnyakova-e-v-o-vozmozhnosti-ispolzovaniya-indeksa-vinera-dlya-vychisleniya-priznakov-tekstov-na-estestvennom-yazyke">https://mp.jvolsu.com/index.php/ru/archive-ru/520-mathematical-physics-and-computer-simulation-2025-vol-28-no-3/modelirovanie-informatika-i-upravlenie/1166-klyachin-v-a-khizhnyakova-e-v-o-vozmozhnosti-ispolzovaniya-indeksa-vinera-dlya-vychisleniya-priznakov-tekstov-na-estestvennom-yazyke</self-uri><abstract xml:lang="ru"><p>В статье показано применение индекса Винера к решению одной из задач обработки текстов на естественном языке. Индекс Винера определяется как сумма всех кратчайших расстояний во взвешенном связном графе. Эта величина характеризует сложность графа. В настоящей работе вводятся две нормализации этого индекса. В первом варианте обычный индекс Винера 𝑁 вершинного связного графа делится на (𝑁 − 1)2. Во втором варианте индекс Винера евклидова графа делится на сумму расстояний между любой парой не совпадающих вершин. Для применения к задачам обработки текста в статье вводится граф предложений текста: ребро образует пара слов, которые встречаются в тексте в каком-либо предложении. Чтобы вычислять величину индекса Винера для евклидова графа, применяется вложение слов. В статье вкратце описан алгоритм обучения вложению слов Т. Миколова. Дополнительно приводится алгоритм приближенного вычисления остовного дерева с минимальным индексом Винера. Алгоритм основан на минимизациинового слагаемого при добавлении ребра к построенной части дерева. С целью идентификации неинформативного текста вычисляются 4 признака на основе индекса Винера и его модификаций. Классификация осуществляется стандартными методами машинного обучения.</p></abstract><abstract xml:lang="en" abstract-type="summary"><p>The article demonstrates the application of the Wiener index to solving one of the problems of natural language text processing. The Wiener index is defined as the sum of all shortest distances in a weighted connected graph. This value characterizes the complexity of the graph. In this paper, two modifications of this index are introduced. In the first version, the usual Wiener index of an 𝑁 vertex connected graph is divided by (𝑁 − 1)2. In the second version, the Wiener index of a Euclidean graph is divided by the sum of the distances between any pair of non-coinciding vertices. For application to text processing problems, the article introduces a graph of text sentences: an edge is formed by a pair of words that occur in the text in some sentence. To calculate the value of the Wiener index for a Euclidean graph, word embedding is used. The article briefly describes the algorithm for learning word embeddings by T. Mikolova. In addition, the article provides an algorithm for approximate calculation of a spanning tree with a minimal Wiener index. The algorithm is based on minimizing the new term when adding an edge to the constructed part of the tree. In order to identify uninformative text, 4 features are calculated based on the Wiener index and its modifications. Classification is carried out using standard machine learning methods.</p></abstract><kwd-group xml:lang="ru"><kwd>граф</kwd><kwd>индекс Винера</kwd><kwd>остовное дерево</kwd><kwd>вложение слов</kwd><kwd>машинное обучение</kwd></kwd-group><kwd-group xml:lang="en"><kwd>graph</kwd><kwd>Wiener index</kwd><kwd>spanning tree</kwd><kwd>words embedding</kwd><kwd>machine learning</kwd></kwd-group></article-meta></front><back><ref-list><ref id="ref1"><mixed-citation publication-type="other" xml:lang="ru">Григорьева, Е. Г. Исследование статистических характеристик текста на основе графовой модели лингвистического корпуса / Е. Г. Григорьева, В. А. Клячин // Изв. Сарат. ун-та. Нов. сер. Сер.: Математика. Механика. Информатика. — 2020. — Т. 20, № 1. — C. 116–126.</mixed-citation></ref><ref id="ref2"><mixed-citation publication-type="other" xml:lang="ru">Карабулатова, И. С. Специфика лингвистической параметризации деструктивного массмедийного текста с обесцениванием исторической памяти / И. С. Карабулатова, Г. А. Копнина // Медиалингвистика. — 2023. — Т. 10, № 3. — C. 319–335.</mixed-citation></ref><ref id="ref3"><mixed-citation publication-type="other" xml:lang="ru">Медных, А. Д. Циклические накрытия графов. Перечисление отмеченных остовных лесов и деревьев, индекс Кирхгофа и якобианы / А. Д. Медных, И. А. Медных // УМН. — 2023. — Т. 78, № 3. — C. 115–164.</mixed-citation></ref><ref id="ref4"><mixed-citation publication-type="other" xml:lang="ru">Некрасов, Г. А. Разработка поискового робота для обнаружения веб-контента с фейковыми новостями / Г. А. Некрасов, И. И. Романова // Инновационные, информационные и коммуникационные технологии. — 2017. — № 1. — C. 128–130.</mixed-citation></ref><ref id="ref5"><mixed-citation publication-type="other" xml:lang="ru">Попов, В. В. Естественный текст: математические методы атрибуции / В. В. Попов, Т. В. Штельмах // Вестник Волгоградского государственного университета. Серия 2. Языкознание. — 2019. — Т. 18, № 2. — C. 147–158. — DOI: https://doi.org/10.15688/jvolsu2.2019.2.13</mixed-citation></ref><ref id="ref6"><mixed-citation publication-type="other" xml:lang="ru">Хижнякова, Е. В. NP-полнота задачи построения графа с минимальным коэффициентом непрямолинейности / Е. В. Хижнякова // Математическая физика и компьютерное моделирование. — 2023. — Т. 26, № 2. —C. 43–51. — DOI: https://doi.org/10.15688/mpcm.jvolsu.2023.2.4</mixed-citation></ref><ref id="ref7"><mixed-citation publication-type="other" xml:lang="ru">Anand, A. We Used Neural Networks to Detect Clickbaits: You Won’t Believe What Happened Next / A. Anand, T. Chakraborty, N. Park // 39-th European Conference on Information Retrieval (ECIR). Aberdeen, United Kingdom, 8–13 April 2017. Lecture Notes in Computer Science (LNCS). — 2017. — Vol. 10193. — P. 541–547.</mixed-citation></ref><ref id="ref8"><mixed-citation publication-type="other" xml:lang="ru">Azjargal, E. Minimum of Product of Wiener and Harary Indices / E. Azjargal, B. Horoldagva, I. Gutman // MATCH Commun. Math. Comput. Chem. — 2024. — № 92. —P. 65–71.</mixed-citation></ref><ref id="ref9"><mixed-citation publication-type="other" xml:lang="ru">Biyani, P. 8 Amazing Secrets for Getting More Clicks: Detecting Clickbaits in News Streams Using Article Informality / P. Biyani, K. Tsioutsiouliklis, J. Blackmer // Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016. — 2016. — Vol. 10193, № 2. — P. 541–547.</mixed-citation></ref><ref id="ref10"><mixed-citation publication-type="other" xml:lang="ru">Efficient Estimation of Word Representations in Vector Space / T. Mikolov, K. Chen, G. Corrado, J. Dean // arXiv preprint arXiv:1301.3781. — 2013. — P. 1–12.</mixed-citation></ref><ref id="ref11"><mixed-citation publication-type="other" xml:lang="ru">Geometric Spanning Trees Minimizing the Wiener Index / A. K. Abu-Affash, P. Carmi, O. Luwisch, J. Mitchell // Algorithms and Data Structures. Springer Nature Switzerland. — 2023. — P. 1–14.</mixed-citation></ref><ref id="ref12"><mixed-citation publication-type="other" xml:lang="ru">Halin, R. U‥ ber Simpliziale Zerfallungen Beliebiger / R. Halin // Math. Ann. — 1964. — № 156. — P. 216–225.</mixed-citation></ref><ref id="ref13"><mixed-citation publication-type="other" xml:lang="ru">Identifying Clickbait: A Multi-Strategy Approach Using Neural Networks / V. Kumar, D. Khattar, S. Gairola, L. Y. Kumar, V. Varma // Proceedings of the 41-st International ACM SIGIR Conference on Research &amp; Development in Information Retrieval, Ann Arbor, MI, USA, 8–12 July 2018. — 2018. — Vol. 10193. — P. 1225–1228.</mixed-citation></ref><ref id="ref14"><mixed-citation publication-type="other" xml:lang="ru">Knor, M. Selected Topics on Wiener Index / M. Knor, R. ˆSkrekovski, A. Tepeh // Ars mathematica contemporanea. — 2024. — Vol. 24, № 4. — P. 1–31.</mixed-citation></ref><ref id="ref15"><mixed-citation publication-type="other" xml:lang="ru">Wang, Hedi The Minimum Wiener Index of Halin Graphs With Characteristic Trees of Diameter 4 / Hedi Wang, Kexiang Xu // Electron. J. Math. — 2025. — № 9. — P. 11–22.</mixed-citation></ref><ref id="ref16"><mixed-citation publication-type="other" xml:lang="ru">Wiener, H. Structural Determination of Paraffin Boiling Points / H. Wiener // J. Amer. Chem. Soc. — 1947. — Vol. 69, № 1. — P. 17–20.</mixed-citation></ref></ref-list></back></article>
