Logo PTI Logo FedCSIS

Position Papers of the 21st Conference on Computer Science and Intelligence Systems

Annals of Computer Science and Information Systems, Volume 48

Re-evaluating Polish LLMs in RAG: The Evolution of Bielik and PLLuM and the Limits of Open-Source Evaluators

, , , ,

DOI: http://dx.doi.org/10.15439/2026F9874

Citation: Marcin Szewczyk, , , ,

Full text

Abstract. This paper extends a framework suggested for the evaluation of a RAG system in an industrial environment. We measure generational progress by comparing Polish LLMs against their latest iterations and analyze the ''LLM-as-a-Judge'' paradigm by evaluating the evaluators themselves. The evaluation was conducted using two distinct classes of open-weight judges: small (under 10B parameters) and large (over 20B parameters), employing both quantitative scoring and order-controlled pair- wise A/B testing. The results reveal that small judges exhibit severe position bias, making them unreliable. Furthermore, they confirm that Bielik-11B-v3.0-Instruct consistently outperforms PLLuM-12B-nc-chat-250715 in accuracy and relevance. We con- clude that robust local RAG evaluation requires judge models exceeding the 10B parameter threshold, while Bielik remains the superior choice for Polish tasks.

References

  1. S. Dadas, M. Perełkiewicz, and R. Poświata, “Pirb: A comprehensive benchmark of polish dense and hybrid text retrieval methods,” 2024. [Online]. Available: https://arxiv.org/abs/2402.13350
  2. K. Ociepa, Ł. Flis, K. Wróbel, A. Gwoździej, and R. Kinas, “Bielik 11b v2 technical report,” arXiv, Tech. Rep. https://arxiv.org/abs/2505.02410, 2025.
  3. S. Bartanowicz and K. Jassem, “Evaluation of two leading polish language models in a real-world rag scenario,” Adam Mickiewicz University, Tech. Rep., 2025.
  4. B. Orlik, “KartonBERT-USE-base-v1: A polish universal sentence encoder,” https://huggingface.co/OrlikB/KartonBERT-USE-base-v1, 2024, hugging Face model repository.
  5. R. Friel et al., “Ragbench: Explainable benchmark for retrievalaugmented generation,” in Proceedings of a 2024 venue (preprint: arXiv:2407.11005), 2024. https://dx.doi.org/10.48550/arXiv.2407.11005
  6. G. Xiong et al., “Benchmarking retrieval-augmented generation for medicine,” in Findings of ACL 2024, 2024. https://dx.doi.org/10.18653/v1/2024.findings-acl.372
  7. V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, “Dense passage retrieval for open-domain question answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), 2020. https://dx.doi.org/10.18653/v1/2020.emnlp-main.550 pp. 6769–6781. [Online]. Available: https://aclanthology.org/2020.emnlp-main.550/
  8. CYFRAGOVPL, “Pllum-12b-instruct: Model card,” https://huggingface.co/CYFRAGOVPL/PLLuM-12B-instruct, 2025.
  9. L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023. [Online]. Available: https://arxiv.org/abs/2306.05685
  10. J. Ye et al., “Justice or prejudice? quantifying biases in llm-as-a-judge,” in OpenReview (2024/2025), 2025. https://dx.doi.org/10.48550/arXiv.2410.02736
  11. Y. Zhang et al., “Llms-as-judges: A comprehensive survey on llm-based evaluation,” arXiv preprint arXiv:2412.05579, 2024. https://dx.doi.org/10.48550/arXiv.2412.05579
  12. H. Yin, S. Vardi, and V. Choudhary, “Fragile preferences: A deep dive into order effects in large language models,” in Proceedings of the 14th Conference on Computational Natural Language Learning (CoNLL 2025), 2025. https://dx.doi.org/10.48550/arXiv.2506.14092
  13. W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles, 2023. https://dx.doi.org/10.1145/3600006.3613165