Detecting Ideological Framing in Geopolitical News: Fine-Tuning RoBERTa-Large with Algorithmic Resampling for Binary Bias Classification
Fine-Tuning RoBERTa-Large with Algorithmic Resampling for Binary Bias Classification
DOI:
https://doi.org/10.29304/jqcsm.2026.18.33020Keywords:
Ideological Framing Detection, RoBERTa-Large, Algorithmic Resampling, Class Imbalance, AG News, Computational Linguistics, Media Bias, Transformer ModelsAbstract
Automated identification of political frame effects in biased news texts continues to be a difficult NLP problem, motivated by issues of label distribution imbalance and contextual richness. Despite being successful at traditional text classification tasks, identification of complex geopolitical bias effects becomes hindered by imbalanced datasets with an overrepresentation of the majority class. In this paper, we propose a binary classification pipeline for identification of geopolitical framing biases within a collection of 30,000 short news articles in the AG News dataset "World" sub-category. In order to overcome imbalances in the minority class, we used stratified partitioning and random sampling techniques for achieving class balance in a training set with 43,698 examples. We fine-tuned RoBERTa-Large (355M parameters) model for three epochs with sequence classification task using FP16 precision strategy. The resulting classifier reaches 91.08% accuracy (95% CI: [90.46%, 91.66%], Wilson score method), macro-F1 equal to 0.91 and AUC score of 0.9641 (95% CI: [0.9601, 0.9681], Hanley-McNeil approximation). Pair-wise McNemar's significance test shows that RoBERTa-Large model significantly outperforms ALBERT-base (χ2 = 5.68, p = 0.017) and DistilBERT (χ2 = 155.42, p < 0.0001) transformers which were independently fine-tuned using the same strategy. Experimental findings show that data balancing combined with transformer representations provide a statistically sound solution for bias identification.
Downloads
References
G. Mallard and D. Eggel, “Narrative warfare in the digital age,” Grad. Inst. Int. Dev. Stud., 2023, doi: https://doi.org/10.71609/iheid-25s5-fw06.
T. Spinde et al., “The Media Bias Taxonomy: A Systematic Literature Review on the Forms and Automated Detection of Media Bias,” 2023, [Online]. Available: http://arxiv.org/abs/2312.16148
Y. Otmakhova, S. Khanehzar, and L. Frermann, “Media Framing: A Typology and Survey of Computational Approaches Across Disciplines,” Proc. Annu. Meet. Assoc. Comput. Linguist., vol. 1, no. Figure 1, pp. 15407–15428, 2024, doi: 10.18653/v1/2024.acl-long.822.
F. Hamborg, K. Donnay, and B. Gipp, “Automated identification of media bias in news articles: an interdisciplinary literature review,” Int. J. Digit. Libr., vol. 20, no. 4, pp. 391–415, 2019, doi: 10.1007/s00799-018-0261-y.
G. Vallejo, T. Baldwin, and L. Frermann, “Connecting the Dots in News Analysis: Bridging the Cross-Disciplinary Disparities in Media Bias and Framing,” NLP+CSS 2024 - 6thWorkshop Nat. Lang. Process. Comput. Soc. Sci. Proc. Work., pp. 16–31, 2024, doi: 10.18653/v1/2024.nlpcss-1.2.
F. J. Rodrigo-Ginés, J. Carrillo-de-Albornoz, and L. Plaza, “A systematic review on media bias detection: What is media bias, how it is expressed, and how to detect it,” Expert Syst. Appl., vol. 237, no. PC, p. 121641, 2024, doi: 10.1016/j.eswa.2023.121641.
M. Yahya et al., “Bias Detection in Media: Traditional Models vs. Transformers in Analyzing Social Media Coverage of the Israeli-Gaza Conflict,” Proc. - Int. Conf. Comput. Linguist. COLING, pp. 114–121, 2025.
M. Powers, U. Mavani, H. R. Jonala, A. Tiwari, and H. Wei, “GUS-Net: Social Bias Classification in Text with Generalizations, Unfairness, and Stereotypes,” pp. 1–30, 2024, [Online]. Available: http://arxiv.org/abs/2410.08388
K. Mohiuddin et al., “Attention Is All You Need,” Int. Conf. Inf. Knowl. Manag. Proc., no. Nips, pp. 4752–4758, 2023, doi: 10.1145/3583780.3615497.
P. Miłkowski, M. Gruza, K. Kanclerz, P. Kazienko, D. Grimling, and J. Kocó, “Improving Media Bias Detection with State-of-the-art Transformers,” pp. 248–259, 2023, [Online]. Available: https://gipplab.org/wp-content/papercite-data/pdf/wessel2022.pdf
A. A. Rai, “AG News Classification Dataset,” Kaggle. Accessed: Jan. 15, 2026. [Online]. Available: https://www.kaggle.com/datasets/amananandrai/ag-news-classification-dataset
A. Ortiz, T. Rodrigo, and J. Sicilia, “The BBVA Research Geopolitics Monitor : Tracking Geopolitical Sentiment and Events using Natural Language Techniques,” pp. 1–7, 2023.
T. Spinde, L. Rudnitckaia, K. Sinha, F. Hamborg, B. Gipp, and K. Donnay, “MBIC – A Media Bias Annotation Dataset Including Annotator Characteristics,” Proc. iConference 2021, pp. 1–8, 2021.
H. Ghosh, A. Mosharafa, and G. Groh, “To Bias or Not to Bias: Detecting bias in News with bias-detector,” pp. 1–7, 2025, [Online]. Available: http://arxiv.org/abs/2505.13010
S. Henning, W. Beluch, A. Fraser, and A. Friedrich, “A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language Processing,” EACL 2023 - 17th Conf. Eur. Chapter Assoc. Comput. Linguist. Proc. Conf., no. 2022, pp. 523–540, 2023, doi: 10.18653/v1/2023.eacl-main.38.
W. Albattah and R. U. Khan, “Impact of imbalanced features on large datasets,” Front. Big Data, vol. 8, 2025, doi: 10.3389/fdata.2025.1455442.
H. He and E. A. Garcia, “Learning from Imbalanced Data,” IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, 2009, doi: 10.1109/TKDE.2008.239.
L. Dube and T. Verster, “Enhancing classification performance in imbalanced datasets: A comparative analysis of machine learning models,” Data Sci. Financ. Econ., vol. 3, no. 4, pp. 354–379, 2023, doi: 10.3934/dsfe.2023021.
A. Moreo, A. Esuli, and F. Sebastiani, “Distributional random oversampling for imbalanced text classification,” SIGIR 2016 - Proc. 39th Int. ACM SIGIR Conf. Res. Dev. Inf. Retr., pp. 805–808, 2016, doi: 10.1145/2911451.2914722.
S. F. Taskiran, B. Turkoglu, E. Kaya, and T. Asuroglu, “A comprehensive evaluation of oversampling techniques for enhancing text classification performance,” Sci. Rep., vol. 15, no. 1, pp. 1–20, 2025, doi: 10.1038/s41598-025-05791-7.
Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” no. 1, 2019.
J. C. Timoneda and S. V. Vera, “BERT, RoBERTa, or DeBERTa? Comparing Performance Across Transformers Models in Political Science Text,” J. Polit., vol. 87, no. 1, pp. 347–364, 2025, doi: 10.1086/730737.
S. Liu, B. Wang, W. Xiang, H. Xu, and M. Xu, “Encoding Hierarchical Schema via Concept Flow for Multifaceted Ideology Detection,” Proc. Annu. Meet. Assoc. Comput. Linguist., pp. 2930–2942, 2024, doi: 10.18653/v1/2024.findings-acl.172.
L. Lin, L. Wang, X. Zhao, J. Li, and K. F. Wong, “IndiVec: An Exploration of Leveraging Large Language Models for Media Bias Detection with Fine-Grained Bias Indicators,” EACL 2024 - 18th Conf. Eur. Chapter Assoc. Comput. Linguist. Find. EACL 2024, pp. 1038–1050, 2024, doi: 10.18653/v1/2024.findings-eacl.70.
A. Elbouanani, E. Dufraisse, and A. Popescu, “Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification,” pp. 15476–15505, 2025, doi: 10.18653/v1/2025.findings-acl.799.
J. Yoo and Y. Shin, “Fair or Framed? Political Bias in News Articles Generated by LLMs,” EMNLP 2025 - 2025 Conf. Empir. Methods Nat. Lang. Process. Proc. Conf., pp. 16904–16930, 2025, doi: 10.18653/v1/2025.emnlp-main.856.
V. U. Gongane, M. V Munot, and A. D. Anuse, “A survey of explainable AI techniques for detection of fake news and hate speech on social media platforms,” J. Comput. Soc. Sci., vol. 7, no. 1, pp. 587–623, 2024, doi: 10.1007/s42001-024-00248-9.
J. Cohen, “A Coefficient of Agreement for Nominal Scales,” Educ. Psychol. Meas., vol. 20, no. 1, pp. 37–46, 1960.
T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006, doi: 10.1016/j.patrec.2005.10.010.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Hanaa Ali Alshaibani, Ibtihal A. Mustafa

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.








