Quantum computing and machine learning for Arabic language sentiment classification in social media.
Journal
Scientific reports
ISSN: 2045-2322
Titre abrégé: Sci Rep
Pays: England
ID NLM: 101563288
Informations de publication
Date de publication:
12 10 2023
12 10 2023
Historique:
received:
29
04
2023
accepted:
03
10
2023
medline:
23
10
2023
pubmed:
13
10
2023
entrez:
12
10
2023
Statut:
epublish
Résumé
With the increasing amount of digital data generated by Arabic speakers, the need for effective and efficient document classification techniques is more important than ever. In recent years, both quantum computing and machine learning have shown great promise in the field of document classification. However, there is a lack of research investigating the performance of these techniques on the Arabic language. This paper presents a comparative study of quantum computing and machine learning for two datasets of Arabic language document classification. In the first dataset of 213,465 Arabic tweets, both classic machine learning (ML) and quantum computing approaches achieve high accuracy in sentiment analysis, with quantum computing slightly outperforming classic ML. Quantum computing completes the task in approximately 59 min, slightly faster than classic ML, which takes around 1 h. The precision, recall, and F1 score metrics indicate the effectiveness of both approaches in predicting sentiment in Arabic tweets. Classic ML achieves precision, recall, and F1 score values of 0.8215, 0.8175, and 0.8121, respectively, while quantum computing achieves values of 0.8239, 0.8199, and 0.8147, respectively. In the second dataset of 44,000 tweets, both classic ML (using the Random Forest algorithm) and quantum computing demonstrate significantly reduced processing times compared to the first dataset, with no substantial difference between them. Classic ML completes the analysis in approximately 2 min, while quantum computing takes approximately 1 min and 53 s. The accuracy of classic ML is higher at 0.9241 compared to 0.9205 for quantum computing. However, both approaches achieve high precision, recall, and F1 scores, indicating their effectiveness in accurately predicting sentiment in the dataset. Classic ML achieves precision, recall, and F1 score values of 0.9286, 0.9241, and 0.9249, respectively, while quantum computing achieves values of 0.92456, 0.9205, and 0.9214, respectively. The analysis of the metrics indicates that quantum computing approaches are effective in identifying positive instances and capturing relevant sentiment information in large datasets. On the other hand, traditional machine learning techniques exhibit faster processing times when dealing with smaller dataset sizes. This study provides valuable insights into the strengths and limitations of quantum computing and machine learning for Arabic document classification, emphasizing the potential of quantum computing in achieving high accuracy, particularly in scenarios where traditional machine learning techniques may encounter difficulties. These findings contribute to the development of more accurate and efficient document classification systems for Arabic data.
Identifiants
pubmed: 37828056
doi: 10.1038/s41598-023-44113-7
pii: 10.1038/s41598-023-44113-7
pmc: PMC10570340
doi:
Types de publication
Journal Article
Research Support, Non-U.S. Gov't
Langues
eng
Sous-ensembles de citation
IM
Pagination
17305Informations de copyright
© 2023. Springer Nature Limited.
Références
Nature. 2000 Apr 6;404(6778):579-81
pubmed: 10766235
Science. 2021 Apr 16;372(6539):
pubmed: 33859004
Arab J Sci Eng. 2022;47(2):2499-2511
pubmed: 34660170
Int Arch Occup Environ Health. 2021 Jul;94(5):1097-1111
pubmed: 33491101
Front Artif Intell. 2022 Mar 30;5:843038
pubmed: 35434606
Brief Bioinform. 2019 Sep 27;20(5):1878-1912
pubmed: 30084866
Lancet Digit Health. 2020 May;2(5):e221-e223
pubmed: 33328054
J Am Med Inform Assoc. 2021 Feb 15;28(2):360-364
pubmed: 33027509
Nature. 2022 Jul;607(7920):667-676
pubmed: 35896643
Entropy (Basel). 2023 Mar 21;25(3):
pubmed: 36981428
Entropy (Basel). 2023 Feb 03;25(2):
pubmed: 36832654