Re-evaluating Polish LLMs in RAG: The Evolution of Bielik and PLLuM and the Limits of Open-Source Evaluators
Marcin Szewczyk, Tymon Pluciński, Krzysztof Maksim, Szymon Bartanowicz, Krzysztof Jassem
DOI: http://dx.doi.org/10.15439/2026F9874
Citation: Marcin Szewczyk, Tymon Pluciński, Krzysztof Maksim, Szymon Bartanowicz, Krzysztof Jassem (2026). Re-evaluating Polish LLMs in RAG: The Evolution of Bielik and PLLuM and the Limits of Open-Source Evaluators. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems. ACSIS, Vol. 48, pages 165–169.
Abstract. This paper extends a framework suggested for the evaluation of a RAG system in an industrial environment. We measure generational progress by comparing Polish LLMs against their latest iterations and analyze the ''LLM-as-a-Judge'' paradigm by evaluating the evaluators themselves. The evaluation was conducted using two distinct classes of open-weight judges: small (under 10B parameters) and large (over 20B parameters), employing both quantitative scoring and order-controlled pair- wise A/B testing. The results reveal that small judges exhibit severe position bias, making them unreliable. Furthermore, they confirm that Bielik-11B-v3.0-Instruct consistently outperforms PLLuM-12B-nc-chat-250715 in accuracy and relevance. We con- clude that robust local RAG evaluation requires judge models exceeding the 10B parameter threshold, while Bielik remains the superior choice for Polish tasks.
References
- S. Dadas, M. Perełkiewicz, and R. Poświata, “Pirb: A comprehensive benchmark of polish dense and hybrid text retrieval methods,” 2024. [Online]. Available: https://arxiv.org/abs/2402.13350
- K. Ociepa, Ł. Flis, K. Wróbel, A. Gwoździej, and R. Kinas, “Bielik 11b v2 technical report,” arXiv, Tech. Rep. https://arxiv.org/abs/2505.02410, 2025.
- S. Bartanowicz and K. Jassem, “Evaluation of two leading polish language models in a real-world rag scenario,” Adam Mickiewicz University, Tech. Rep., 2025.
- B. Orlik, “KartonBERT-USE-base-v1: A polish universal sentence encoder,” https://huggingface.co/OrlikB/KartonBERT-USE-base-v1, 2024, hugging Face model repository.
- R. Friel et al., “Ragbench: Explainable benchmark for retrievalaugmented generation,” in Proceedings of a 2024 venue (preprint: arXiv:2407.11005), 2024. https://dx.doi.org/10.48550/arXiv.2407.11005
- G. Xiong et al., “Benchmarking retrieval-augmented generation for medicine,” in Findings of ACL 2024, 2024. https://dx.doi.org/10.18653/v1/2024.findings-acl.372
- V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, “Dense passage retrieval for open-domain question answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), 2020. https://dx.doi.org/10.18653/v1/2020.emnlp-main.550 pp. 6769–6781. [Online]. Available: https://aclanthology.org/2020.emnlp-main.550/
- CYFRAGOVPL, “Pllum-12b-instruct: Model card,” https://huggingface.co/CYFRAGOVPL/PLLuM-12B-instruct, 2025.
- L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023. [Online]. Available: https://arxiv.org/abs/2306.05685
- J. Ye et al., “Justice or prejudice? quantifying biases in llm-as-a-judge,” in OpenReview (2024/2025), 2025. https://dx.doi.org/10.48550/arXiv.2410.02736
- Y. Zhang et al., “Llms-as-judges: A comprehensive survey on llm-based evaluation,” arXiv preprint arXiv:2412.05579, 2024. https://dx.doi.org/10.48550/arXiv.2412.05579
- H. Yin, S. Vardi, and V. Choudhary, “Fragile preferences: A deep dive into order effects in large language models,” in Proceedings of the 14th Conference on Computational Natural Language Learning (CoNLL 2025), 2025. https://dx.doi.org/10.48550/arXiv.2506.14092
- W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles, 2023. https://dx.doi.org/10.1145/3600006.3613165