IRMA: the 335-million-word Italian coRpus for studying MisinformAtion.


Journal

Proceedings of the conference. Association for Computational Linguistics. Meeting
ISSN: 0736-587X
Titre abrégé: Proc Conf Assoc Comput Linguist Meet
Pays: United States
ID NLM: 101639983

Informations de publication

Date de publication:
May 2023
Historique:
medline: 24 11 2023
pubmed: 24 11 2023
entrez: 24 11 2023
Statut: ppublish

Résumé

The dissemination of false information on the internet has received considerable attention over the last decade. Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive. Therefore, there is an increasing need to develop methods for automatic detection of misinformation. Although resources for creating such methods are available in English, other languages are often underrepresented in this effort. With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as 'untrustworthy' by professional factcheckers. The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms. It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.e., keywords, topics at three different resolutions, and LIWC lexical features). IRMA also includes domainspecific information such as source type (e.g., political, health, conspiracy, etc.), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior. IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.

Identifiants

pubmed: 37997575
pmc: PMC7615326
mid: EMS190995

Types de publication

Journal Article

Langues

eng

Pagination

2339-2349

Subventions

Organisme : European Research Council
ID : 101020961
Pays : International

Références

Nat Hum Behav. 2022 Apr;6(4):495-505
pubmed: 35115677
Nat Hum Behav. 2021 Mar;5(3):337-348
pubmed: 33547453
Nature. 2021 Apr;592(7855):590-595
pubmed: 33731933
Science. 2019 Jan 25;363(6425):374-378
pubmed: 30679368
Online Soc Netw Media. 2021 May;23:100136
pubmed: 36570036
Nat Hum Behav. 2022 Aug;6(8):1069-1078
pubmed: 35606514
PNAS Nexus. 2023 Sep 02;2(9):pgad286
pubmed: 37719749
Sci Adv. 2022 Oct 28;8(43):eabq3668
pubmed: 36288312
Front Artif Intell. 2020 Aug 25;3:62
pubmed: 33733179
Proc Natl Acad Sci U S A. 2019 Feb 12;116(7):2521-2526
pubmed: 30692252
Behav Res Methods. 2022 Aug;54(4):1794-1817
pubmed: 34697754
Proc Natl Acad Sci U S A. 2017 Jan 24;114(4):E457-E465
pubmed: 28069962
EPJ Data Sci. 2021;10(1):34
pubmed: 34249599
PNAS Nexus. 2022 Sep 22;1(4):pgac186
pubmed: 36380855
R Soc Open Sci. 2020 Oct 14;7(10):201199
pubmed: 33204475
Science. 2018 Mar 9;359(6380):1146-1151
pubmed: 29590045
PLoS One. 2016 Mar 04;11(3):e0150989
pubmed: 26943909
PLoS One. 2015 Aug 14;10(8):e0134641
pubmed: 26275043
Glob Chall. 2017 Jan 23;1(2):1600008
pubmed: 31565263
PLoS One. 2019 Nov 18;14(11):e0225098
pubmed: 31738787
Science. 2018 Mar 9;359(6380):1094-1096
pubmed: 29590025
Nat Hum Behav. 2020 Dec;4(12):1285-1293
pubmed: 33122812

Auteurs

Fabio Carrella (F)

School of Psychological Science, University of Bristol.

Alessandro Miani (A)

Institute of Work and Organizational Psychology, University of Neuchâtel.

Stephan Lewandowsky (S)

School of Psychological Science, University of Bristol.

Classifications MeSH