Trusted or Tricked? A Multi-Layer Robustness Scoring Framework for Evaluating LLMs Against Public Health Misinformation
Razan El Mais, Nourhan Hamze, Chaymaa Abbass, Hadi Al Mubasher, Mariette Awad
DOI: http://dx.doi.org/10.15439/2026F3733
Citation: Razan El Mais, Nourhan Hamze, Chaymaa Abbass, Hadi Al Mubasher, Mariette Awad (2026). Trusted or Tricked? A Multi-Layer Robustness Scoring Framework for Evaluating LLMs Against Public Health Misinformation. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems. ACSIS, Vol. 48, pages 205–212.
Abstract. The increasing use of Large Language Models (LLMs) as accessible sources of public health information raises concerns about their potential to generate or reinforce medical misinformation. This study evaluates four state-of-the-art LLMs (LLaMA, DeepSeek, Gemini, and GPT-4.1-nano) using the U.S. Centers for Disease Control and Prevention (CDC)-verified medical questions to assess factual reliability, consistency, and robustness under neutral, provocative, and misinformation-seeking prompts. We propose a two-layer evaluation framework that combines quantitative behavioral metrics including factual alignment, semantic and topic drift, hallucination, toxicity, overconfidence, and refusal behavior with qualitative LLM-as-a-judge assessment into a unified robustness score. This enables a comprehensive evaluation of factual reliability, safety, tone stability, and prompt-specific robustness. Results reveal clear behavioral differences across models. DeepSeek and GPT-4.1-nano achieve the strongest factual consistency and safer responses, whereas LLaMA is more susceptible to prompt-induced drift and Gemini exhibits more conservative refusal behavior but lower overall robustness. These findings demonstrate the value of structured multi-dimensional evaluation for identifying failure modes and supporting the deployment of reliable, misinformation-aware LLMs in healthcare.
References
- L. Hanu and the Unitary Team, Detoxify. Unitary AI, 2020. [Online]. Available: https://github.com/unitaryai/detoxify
- Centers for Disease Control and Prevention, Artificial Intelligence and Public Health: Opportunities and Risks. Atlanta, GA, USA: CDC, 2023.
- C. Wardle and H. Derakhshan, Information Disorder: Toward an Interdisciplinary Framework for Research and Policymaking. Strasbourg, France: Council of Europe, 2017.
- S. Akhtar and A. Akhtar, “Instruction-tuning LLMs for bias and misinformation detection across diverse domains,” arXiv preprint, 2025.
- S. Kuntur, A. Wróblewska, M. Paprzycki, and M. Ganzha, “Under the influence: A survey of large language models in fake news detection,” IEEE Transactions on Artificial Intelligence, 2024, early access, https://dx.doi.org/10.1109/TAI.2024.3471735.
- A. Dharawat, I. Lourentzou, A. Morales, and C. Zhai, “Drink bleach or do what now? COVID-HERA: A study of risk-informed health decision making in the presence of COVID-19 misinformation,” in Proc. Int. AAAI Conf. Web and Social Media (ICWSM), vol. 16, 2022, pp. 1218– 1227.
- A. Bondielli and F. Marcelloni, “A survey on fake news and rumour detection techniques,” Information Sciences, vol. 497, pp. 38–55, 2019, https://dx.doi.org/10.1016/j.ins.2019.05.035.
- K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, “Fake news detection on social media: A data mining perspective,” ACM SIGKDD Explorations Newsletter, vol. 19, no. 1, pp. 22–36, 2017, https://dx.doi.org/10.1145/3137597.3137600.
- H. Ahmed, I. Traore, and S. Saad, “Detection of online fake news using n-gram analysis and machine learning techniques,” in Proc. IEEE Int. Conf. Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments (DESA), 2017, pp. 127–132, https://dx.doi.org/10.1109/DESA.2017.8267730.
- J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186.
- Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint https://arxiv.org/abs/1907.11692, 2019.
- K. Shu, S. Wang, and H. Liu, “Beyond news contents: The role of social context for fake news detection,” in Proc. ACM Int. Conf. Web Search and Data Mining (WSDM), 2019, pp. 312–320, https://dx.doi.org/10.1145/3289600.3290994.
- E. Dai, Y. Sun, and C. Zhai, “FakeHealth: A benchmark for health news credibility classification and explanation,” GitHub repository, 2020. [Online]. Available: https://github.com/EnyanDai/FakeHealth
- X. Sun et al., “Med-MMHL: A multimodal benchmark for medical misinformation detection,” GitHub repository, 2023. [Online]. Available: https://github.com/styxsys0927/Med-MMHL
- T. Huang, J. Yi, P. Yu, and X. Xu, “Unmasking digital falsehoods: A comparative analysis of LLM-based misinformation detection strategies,” arXiv preprint arXiv:2503.00724, 2025.
- M. Chen, L. Wei, H. Cao, W. Zhou, and S. Hu, “Explore the potential of LLMs in misinformation detection: An empirical study,” arXiv preprint arXiv:2311.12699, 2023.
- OpenAI, “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2024.
- A. Dubey et al., “The Llama 3 Herd of Models,” arXiv preprint arXiv:2407.21783, 2024.
- DeepSeek-AI, “DeepSeek LLM: Scaling Open-Source Language Models with Longtermism,” arXiv preprint arXiv:2401.02954, 2024.
- Z. Chen et al., “MEDITRON-70B: Scaling Medical Pretraining for Large Language Models,” arXiv preprint arXiv:2311.16079, 2023.
- World Health Organization, “Health topics and fact sheets,” 2024. [Online]. Available: https://www.who.int/health-topics
- Centers for Disease Control and Prevention, “Health information and disease prevention resources,” 2024. [Online]. Available: https://www.cdc.gov
- National Library of Medicine, “PubMed: Biomedical literature database,” 2024. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov
- N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” arXiv preprint arXiv:1908.10084, 2019.
- N. Dziri et al., “On the origin of hallucinations in conversational models: Is it the datasets or the models?,” in Proc. NAACL-HLT, 2022, pp. 5271– 5285.
- X. Yang et al., “Neural topic modeling with large language models in the loop,” in Proc. ACL, 2025, pp. 1377–1401.
- Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023.
- S. Gehman et al., “RealToxicityPrompts: Evaluating neural toxic degeneration in language models,” arXiv preprint arXiv:2009.11462, 2020.
- S. Kadavath et al., “Language models (mostly) know what they know,” arXiv preprint arXiv:2207.05221, 2022.
- F. Jiang et al., “SafeChain: Safety of language models with long chainof-thought reasoning capabilities,” arXiv preprint arXiv:2502.12025, 2025.
- S. Bowman, G. Angeli, C. Potts, and C. Manning, “A large annotated corpus for learning natural language inference,” in Proc. EMNLP, 2015, pp. 632–642.
- H. Li et al., “LLMs-as-judges: A comprehensive survey on LLM-based evaluation methods,” arXiv preprint arXiv:2412.05579, 2024.
- T. Kocmi and C. Federmann, “Large language models are state-of-theart evaluators of translation quality,” arXiv preprint arXiv:2302.14520, 2023.
- J. Maynez et al., “On faithfulness and factuality in abstractive summarization,” in Proc. ACL, 2020.
- L. Weidinger et al., “Taxonomy of risks from language models,” in Proc. ACM FAccT, 2022.
- P. Röttger et al., “HateCheck: Functional tests for robustness of hate speech detection models,” in Proc. ACL, 2021.
- H. Nori et al., “Capabilities of GPT-4 on medical challenge problems,” arXiv preprint arXiv:2303.13375, 2023.
- Y. Bai et al., “Constitutional AI: Harmlessness from AI feedback,” arXiv preprint arXiv:2212.08073, 2022.