Small Language Models for Pay Equity Compliance: A Fine-Tuning and Evaluation Study
Mikołaj Kuna, Marcin Kowalczyk
DOI: http://dx.doi.org/10.15439/2026F1202
Citation: Mikołaj Kuna, Marcin Kowalczyk (2026). Small Language Models for Pay Equity Compliance: A Fine-Tuning and Evaluation Study. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS). ACSIS, Vol. 47, pages 561–566.
Abstract. The EU Pay Transparency Directive (2023/970), applicable from June 2026, requires large employers to report gender pay gaps with documentation of how they were calculated. Turning raw statistical outputs into compliant, readable reports is a barrier for human-resources teams without data-science expertise. Cloud language models could help, but pay data is special-category personal data under the GDPR and risky to send to external services. We therefore ask whether small open-weight language models, fine-tuned and run entirely on local hardware, can do this work without the data leaving the organisation. Using a memory-efficient method (QLoRA, quantized low-rank adaptation), we adapt Llama 3.1 8B, Mistral 7B and Phi-3.5 Mini for three tasks: turning an analyst's question into statistical parameters, explaining pay-gap results faithfully, and checking a report against the Directive. All run on a mid-range workstation with an 8 GB GPU. Fine-tuning clearly helps the first and third tasks, whereas faithful explanation depends more on retrieving worked examples at inference time. Accuracy saturates at about 40 training examples in total (around 13 per task), indicating feasibility without dedicated machine-learning infrastructure.
References
- European Parliament and Council, "Directive (EU) 2023/970 of the European Parliament and of the Council on pay transparency and enforcement of the principle of equal pay," Official Journal of the European Union, L 132, pp. 21–44, 10 May 2023.
- L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al., "Training Language Models to Follow Instructions with Human Feedback," in Proc. 36th Conf. Neural Information Processing Systems (NeurIPS), 2022.
- J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai and Q. V. Le, "Finetuned Language Models Are Zero-Shot Learners," in Proc. 10th Int. Conf. Learning Representations (ICLR), 2022.
- T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., "Language Models are Few-Shot Learners," in Proc. 34th Conf. Neural Information Processing Systems (NeurIPS), 2020, pp. 1877–1901.
- S. G. Patil, T. Zhang, X. Wang and J. E. Gonzalez, "Gorilla: Large Language Model Connected with Massive APIs," arXiv preprint https://arxiv.org/abs/2305.15334, 2023. DOI: 10.48550/arXiv.2305.15334.
- Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, S. Zhao, L. Hong, R. Tian, R. Xie, J. Zhou, M. Gerstein, D. Li, Z. Liu and M. Sun, "ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs," arXiv preprint https://arxiv.org/abs/2307.16789, 2023. DOI: 10.48550/arXiv.2307.16789.
- E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang and W. Chen, "LoRA: Low-Rank Adaptation of Large Language Models," in Proc. 10th Int. Conf. Learning Representations (ICLR), 2022.
- T. Dettmers, A. Pagnoni, A. Holtzman and L. Zettlemoyer, "QLoRA: Efficient Finetuning of Quantized LLMs," in Proc. 37th Conf. Neural Information Processing Systems (NeurIPS), 2023.
- C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, M. Mao, X. Ma, A. Efrat, P. Yu, L. Yu, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer and O. Levy, "LIMA: Less Is More for Alignment," in Proc. 37th Conf. Neural Information Processing Systems (NeurIPS), 2023.
- T. Kojima, S. S. Gu, M. Reid, Y. Matsuo and Y. Iwasawa, "Large Language Models are Zero-Shot Reasoners," in Proc. 36th Conf. Neural Information Processing Systems (NeurIPS), 2022.
- Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi and H. Hajishirzi, "Self-Instruct: Aligning Language Models with Self-Generated Instructions," in Proc. 61st Annual Meeting of the Association for Computational Linguistics (ACL), 2023, pp. 13484–13508. DOI: 10.18653/v1/2023.acl-long.754.
- S. Es, J. James, L. Espinosa-Anke and S. Schockaert, "RAGAS: Automated Evaluation of Retrieval Augmented Generation," in Proc. 18th Conf. European Chapter of the Association for Computational Linguistics (EACL), 2024, pp. 150–158. DOI: 10.18653/v1/2024.eacl-demo.16.
- N. Reimers and I. Gurevych, "Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks," in Proc. 2019 Conf. on Empirical Methods in Natural Language Processing (EMNLP), 2019, pp. 3982–3992. DOI: 10.18653/v1/D19-1410.
- J. Johnson, M. Douze and H. Jégou, "Billion-Scale Similarity Search with GPUs," IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, 2021. DOI: 10.1109/TBDATA.2019.2921572.
- L. Xu, M. Skoularidou, A. Cuesta-Infante and K. Veeramachaneni, "Modeling Tabular Data using Conditional GAN" in Proc. 33rd Conf. Neural Information Processing Systems (NeurIPS), 2019.
- A. Vogelsang and M. Borg, "Requirements Engineering for Machine Learning: Perspectives from Data Scientists," in Proc. 6th Int. Workshop on Artificial Intelligence for Requirements Engineering (AIRE), IEEE RE, 2019, pp. 245–251. DOI: 10.1109/REW.2019.00050.
- S. Amershi, A. Begel, C. Bird, R. DeLine, H. Gall, E. Kamar, N. Nagappan, B. Nushi and T. Zimmermann, "Software Engineering for Machine Learning: A Case Study," in Proc. 41st IEEE/ACM Int. Conf. Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2019, pp. 291–300. DOI: 10.1109/ICSE-SEIP.2019.00042.
- A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, et al., "The Llama 3 Herd of Models," arXiv preprint https://arxiv.org/abs/2407.21783, 2024. DOI: 10.48550/arXiv.2407.21783.
- A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand and G. Lengyel, "Mistral 7B," arXiv preprint https://arxiv.org/abs/2310.06825, 2023. DOI: 10.48550/arXiv.2310.06825.
- M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, et al., "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone," arXiv preprint https://arxiv.org/abs/2404.14219, 2024.
- S. N. Wood, Generalized Additive Models: An Introduction with R, 2nd ed. Boca Raton, FL: CRC Press, 2017.
- I. Chalkidis, A. Jana, D. Hartung, M. Bommarito, I. Androutsopoulos, D. M. Katz and N. Aletras, "LexGLUE: A Benchmark Dataset for Legal Language Understanding in English," in Proc. 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022, pp. 4310–4330. DOI: 10.18653/v1/2022.acl-long.297.
- I. Chalkidis, T. Pasini, S. Zhang, L. Tomada, S. F. Schwemer and A. Søgaard, "FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing," in Proc. 60th Annual Meeting of the Association for Computational Linguistics (ACL), 2022, pp. 4389–4406. DOI: 10.18653/v1/2022.acl-long.301.