Language and Communicative Aspects of Disinformation

International interdisciplinary conference

Enikő Németh T.
Faculty of Humanities and Social Sciences, University of Szeged

According to the Global Risk Report of the United Nations published in 2024 (UN 2024), disinformation and misinformation are the second most important threats to sustainable development in North America and Europe, and the third one to all member states. The report also states that the international community is not prepared to defend against them. All this shows that recognizing and curbing disinformation, or fake news, which is mainly spread online on digital platforms faster than real news (Vosoughi – Roy – Aral 2018), is one of the great challenges of the 21st century. The flood of fake news causes significant damage at the individual and societal level, e.g. in healthcare, politics or economy. That is why it is vital to expose disinformation and stop its spread.

In the MTA-SZTE-DE Research Group for Theoretical Linguistics and Informatics we raised the questions: Can we, as linguists, do something to stem the tide of disinformation? Can we help to identify fake news with linguistic methods? Are there any linguistic features that can be used to detect fake news? In our project entitled Linguistic-based identification of fake news and pseudoscientific views supported by the Hungarian Academy of Sciences, we investigate Hungarian health-related fake and real news to answer these questions. We employ a “form to function” corpus pragmatic approach (Aijmer 2018) and combine qualitative and quantitative comparative methods applied to the texts randomly selected from the MedCollect corpus, created by the research group. The MedCollect corpus (Szécsényi et al. 2024) contains 1480 health-related texts published between 2007 and 2023, which are of different genres and diverse topics. The fake news texts were collected manually from listed fake news sites, while most of the real news was collected automatically, using random sampling from reliable health portals.

The analyses based on the manual annotation of randomly selected texts from the MedCollect corpus have revealed that there are significant differences between fake and real news regarding the frequency and function of certain linguistic elements and influencing strategies. These include the implicit pragmatic phenomena in headlines, imperative forms, indirect directives, social deixis, intensification, quoting etc. In the present lecture three distinctive elements will be discussed in detail, namely implicit pragmatic phenomena in headlines, directives, and social deixis. The results of the discussion will show that the theoretically based, fine-grained empirical identification of the linguistic differences between fake and real news can contribute to the detection of fake news.

Bibliography

  • AIJMER, Karin (2018): Corpus pragmatics: From form to function. In: A. H. Jucker – K. P. Schneider – W. Bublitz (eds.): Methods in Pragmatics. Berlin/Boston: De Gruyter Mouton. pp. 555–585.
  • SZÉCSÉNYI, Tibor – NAGY C., Katalin – NÉMETH T., Enikő (2024): Felszólításannotálás a MedCollect egészségügyi álhírkorpuszban [Annotation of the imperative forms in MedCollect healthcare fake news corpus]. In: G. Berend – G. Gosztolya – V. Vincze (eds.): XX. Magyar Számítógépes Nyelvészeti Konferencia [XX. conference of Hungarian computational linguistics]. Szeged: Szegedi Tudományegyetem. pp. 159–170.

  • UNITED NATIONS GLOBAL RISK RECORD. 2024. United Nations. Available at: https://unglobalriskreport.org/UNHQ-GlobalRiskReport-WEB-FIN.pdf

  • VOSOUGHI, Soroush – ROY, Deb – ARAL, Sinan 2018. The spread of true and false news online. In: Science Vol.359, Issue 6380, pp. 1146–1151. DOI: 10.1126/science.aap955