Informatics and Automation
RUS  ENG    JOURNALS   PEOPLE   ORGANISATIONS   CONFERENCES   SEMINARS   VIDEO LIBRARY   PACKAGE AMSBIB  
General information
Latest issue
Archive

Search papers
Search references

RSS
Latest issue
Current issues
Archive issues
What is RSS



Informatics and Automation:
Year:
Volume:
Issue:
Page:
Find






Personal entry:
Login:
Password:
Save password
Enter
Forgotten password?
Register


Informatics and Automation, 2024, Issue 23, volume 4, Pages 1110–1138
DOI: https://doi.org/10.15622/ia.23.4.7
(Mi trspy1316)
 

This article is cited in 4 scientific papers (total in 4 papers)

Artificial Intelligence, Knowledge and Data Engineering

A combined term extraction method for the problem of monitoring thematic discussions in social media

V. Pimeshkova, M. Nikonorovaa, M. Shishaevab

a IIMM KSC RAS
b Apatity branch of MAU
Abstract: Term extraction is an important stage in the automated construction of knowledge systems based on natural language texts, since it provides the formation of a basic concept system, which is then used in applied problems of intellectual information processing. The article discusses the problem of automated extraction of terms from natural language texts for their further use in the construction of formalized knowledge systems (ontologies, thesauruses, knowledge graphs) within the problem of monitoring thematic discussions in social media. This problem is characterized by the need to include in the formed knowledge system both concepts from several different domains, and some general concepts used by the audience of social media within thematic discussions. In addition, the generated knowledge system is dynamic both in terms of the composition of the domains it covers and the composition of relevant concepts to be included in the system. The use of existing classical methods for term extraction in this case is difficult, since they are focused on extracting terms within one domain. Based on this, to solve the problem under consideration, a combined method is proposed, combining approaches based on dictionaries, NER tools and rules. The results of the experiments demonstrate the effectiveness of the proposed combination of approaches to term extraction, which makes it possible to extract terms for the problem of monitoring and analyzing thematic discussions in social media. The developed method significantly exceeds the precision of the considered term extraction tools. As a further direction of research, the possibility of developing a method for solving the problem of identifying nested terms or entities is considered.
Keywords: text mining, term extraction, social media, knowledge extraction.
Funding agency Grant number
Ministry of Science and Higher Education of the Russian Federation 122022800551-0, FMEZ-2022-0007
This work was supported by the Ministry of Science and Higher Education of the Russian Federation (No.122022800551-0, FMEZ-2022-0007).
Received: 14.11.2023
Document Type: Article
UDC: 004.912
Language: Russian
Citation: V. Pimeshkov, M. Nikonorova, M. Shishaev, “A combined term extraction method for the problem of monitoring thematic discussions in social media”, Informatics and Automation, 23:4 (2024), 1110–1138
Citation in format AMSBIB
\Bibitem{PimNikShi24}
\by V.~Pimeshkov, M.~Nikonorova, M.~Shishaev
\paper A combined term extraction method for the problem of monitoring thematic discussions in social media
\jour Informatics and Automation
\yr 2024
\vol 23
\issue 4
\pages 1110--1138
\mathnet{http://mi.mathnet.ru/trspy1316}
\crossref{https://doi.org/10.15622/ia.23.4.7}
Linking options:
  • https://www.mathnet.ru/eng/trspy1316
  • https://www.mathnet.ru/eng/trspy/v23/i4/p1110
  • This publication is cited in the following 4 articles:
    Citing articles in Google Scholar: Russian citations, English citations
    Related articles in Google Scholar: Russian articles, English articles
    Informatics and Automation
     
      Contact us:
     Terms of Use  Registration to the website  Logotypes © Steklov Mathematical Institute RAS, 2025