IRMA: the 335-million-word Italian coRpus for studying MisinformAtion.
Journal
Proceedings of the conference. Association for Computational Linguistics. Meeting
ISSN: 0736-587X
Titre abrégé: Proc Conf Assoc Comput Linguist Meet
Pays: United States
ID NLM: 101639983
Informations de publication
Date de publication:
May 2023
May 2023
Historique:
medline:
24
11
2023
pubmed:
24
11
2023
entrez:
24
11
2023
Statut:
ppublish
Résumé
The dissemination of false information on the internet has received considerable attention over the last decade. Misinformation often spreads faster than mainstream news, thus making manual fact checking inefficient or, at best, labor-intensive. Therefore, there is an increasing need to develop methods for automatic detection of misinformation. Although resources for creating such methods are available in English, other languages are often underrepresented in this effort. With this contribution, we present IRMA, a corpus containing over 600,000 Italian news articles (335+ million tokens) collected from 56 websites classified as 'untrustworthy' by professional factcheckers. The corpus is freely available and comprises a rich set of text- and website-level data, representing a turnkey resource to test hypotheses and develop automatic detection algorithms. It contains texts, titles, and dates (from 2004 to 2022), along with three types of semantic measures (i.e., keywords, topics at three different resolutions, and LIWC lexical features). IRMA also includes domainspecific information such as source type (e.g., political, health, conspiracy, etc.), quality, and higher-level metadata, including several metrics of website incoming traffic that allow to investigate user online behavior. IRMA constitutes the largest corpus of misinformation available today in Italian, making it a valid tool for advancing quantitative research on untrustworthy news detection and ultimately helping limit the spread of misinformation.
Types de publication
Journal Article
Langues
eng
Pagination
2339-2349Subventions
Organisme : European Research Council
ID : 101020961
Pays : International
Références
Nat Hum Behav. 2022 Apr;6(4):495-505
pubmed: 35115677
Nat Hum Behav. 2021 Mar;5(3):337-348
pubmed: 33547453
Nature. 2021 Apr;592(7855):590-595
pubmed: 33731933
Science. 2019 Jan 25;363(6425):374-378
pubmed: 30679368
Online Soc Netw Media. 2021 May;23:100136
pubmed: 36570036
Nat Hum Behav. 2022 Aug;6(8):1069-1078
pubmed: 35606514
PNAS Nexus. 2023 Sep 02;2(9):pgad286
pubmed: 37719749
Sci Adv. 2022 Oct 28;8(43):eabq3668
pubmed: 36288312
Front Artif Intell. 2020 Aug 25;3:62
pubmed: 33733179
Proc Natl Acad Sci U S A. 2019 Feb 12;116(7):2521-2526
pubmed: 30692252
Behav Res Methods. 2022 Aug;54(4):1794-1817
pubmed: 34697754
Proc Natl Acad Sci U S A. 2017 Jan 24;114(4):E457-E465
pubmed: 28069962
EPJ Data Sci. 2021;10(1):34
pubmed: 34249599
PNAS Nexus. 2022 Sep 22;1(4):pgac186
pubmed: 36380855
R Soc Open Sci. 2020 Oct 14;7(10):201199
pubmed: 33204475
Science. 2018 Mar 9;359(6380):1146-1151
pubmed: 29590045
PLoS One. 2016 Mar 04;11(3):e0150989
pubmed: 26943909
PLoS One. 2015 Aug 14;10(8):e0134641
pubmed: 26275043
Glob Chall. 2017 Jan 23;1(2):1600008
pubmed: 31565263
PLoS One. 2019 Nov 18;14(11):e0225098
pubmed: 31738787
Science. 2018 Mar 9;359(6380):1094-1096
pubmed: 29590025
Nat Hum Behav. 2020 Dec;4(12):1285-1293
pubmed: 33122812