Logo PTI Logo FedCSIS

Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS)

Annals of Computer Science and Information Systems, Volume 47

High-Performance GEMV on AMD Zen 3 via Architecture-Specific AVX2/FMA Optimization

, ,

DOI: http://dx.doi.org/10.15439/2026F6278

Citation: Rafał Lenart, ,

Full text

Abstract. The General Matrix-Vector Multiply (GEMV) is a fundamental Level-2 BLAS operation appearing throughout scientific computing and numerical linear algebra. GEMV is inherently memory-bound, making effective cache use the dominant optimization factor. We present a hand-optimized GEMV kernel for AMD Zen~3, implemented using AVX2 and FMA intrinsics. Our single-threaded kernel (ZenGEMV) processes 4~rows simultaneously with dual accumulators per row, creating 8~independent FMA chains that fully saturate the Zen~3 pipeline. Benchmarked against OpenBLAS~0.3.31 and AOCL-BLAS~5.2.0 on an AMD Ryzen~5~5600, ZenGEMV achieves up to $1.37\times$ speedup over OpenBLAS and $1.20\times$ over AOCL-BLAS. Our parallel variant (ZenGEMV-P) achieves up to $3.22\times$ speedup over OpenBLAS on small matrices. We analyse why dispatch overhead, suboptimal accumulator strategies, and generic design choices in general-purpose BLAS libraries leave performance gains available in architecture-specific workloads.

References

  1. L. S. Blackford, A. Petitet, R. Pozo, et al., “An updated set of basic linear algebra subprograms (BLAS),” ACM Trans. Math. Softw., vol. 28, no. 2, pp. 135–151, 2002. https://doi.org/10.1145/567806.567807
  2. G. H. Golub and C. F. Van Loan, Matrix Computations, 4th ed. Baltimore: Johns Hopkins University Press, 2013.
  3. C. He, Y. Huang, P. Mu, Z. Miao, J. Xue, L. Ma, et al., “WaferLLM: Large language model inference at wafer scale,” https://arxiv.org/abs/2502.04563 [cs.LG], 2025. https://doi.org/10.48550/arXiv.2502.04563
  4. C. Holmes, D. Mawhirter, Y. He, F. Yan, and B. Wu, “GRNN: Lowlatency and scalable RNN inference on GPUs,” in Proc. 14th EuroSys Conf. (EuroSys’19), New York: ACM, 2019, pp. 41:1–41:16. https://doi. org/10.1145/3302424.3303949
  5. OpenBLAS: An optimized BLAS library. https://www.openblas.net, https://github.com/OpenMathLib/OpenBLAS
  6. AMD, “AOCL-BLAS (BLIS fork with AMD optimizations),” version 5.2.0. https://github.com/amd/blis
  7. P. Gepner, “Using AVX2 instruction set to increase performance of high performance computing code,” Computing and Informatics, vol. 36, no. 5, pp. 1001–1018, 2017. https://doi.org/10.4149/cai_2017_5_1001
  8. H. Zhang, D. Chen, and S. B. Ko, “Efficient multiple-precision floatingpoint fused multiply-add with mixed-precision support,” IEEE Trans. Comput., vol. 68, no. 7, pp. 1035–1048, 2019. https://doi.org/10.1109/ TC.2019.2895031
  9. M. Evers, L. Barnes, and M. Clark, “The AMD next-generation ‘Zen 3’ core,” IEEE Micro, vol. 42, no. 3, pp. 7–12, 2022. https://doi.org/10. 1109/MM.2022.3152788
  10. S. Williams, A. Waterman, and D. A. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Commun. ACM, vol. 52, pp. 65–76, 2009. https://doi.org/10.1145/1498765.1498785
  11. Advanced Micro Devices, “Software optimization guide for AMD EPYC 7003 series processors (family 19h models 00h–0Fh),” rev. 3.00. https://kib.kiev.ua/x86docs/AMD/Optimization/56665_3.00_ Software%20Optimization%20Guide%20for%20AMD%20Family% 2019h%20Processors%20%28PUB%29.pdf
  12. A. Fog, “Instruction tables: lists of instruction latencies, throughputs and micro-operation breakdowns for Intel, AMD, and VIA CPUs,” 2011, last updated Sep. 20, 2025. https://www.agner.org/optimize/instruction_ tables.pdf
  13. A. Fog, “The microarchitecture of Intel, AMD, and VIA CPUs: an optimization guide for assembly programmers and compiler makers,” 2017, last updated Dec. 15, 2025. https://www.agner.org/optimize/ microarchitecture.pdf
  14. FinlayDaG33k, “RAM bandwidth calculator.” https://edu.finlaydag33k. nl/calculating%20ram%20bandwidth/, https://gitlab.com/FinlayDaG33k/ ram-calculator
  15. Z. Smith, “bandwidth: a memory bandwidth benchmark.” https://zs3.me/ bandwidth
  16. K. Goto and R. A. van de Geijn, “Anatomy of high-performance matrix multiplication,” ACM Trans. Math. Softw., vol. 34, no. 3, p. 12, 2008. https://doi.org/10.1145/1356052.1356053
  17. F. G. Van Zee and R. A. van de Geijn, “BLIS: A framework for rapidly instantiating BLAS functionality,” ACM Trans. Math. Softw., vol. 41, no. 3, p. 14, 2015. https://doi.org/10.1145/2764454
  18. R. C. Whaley and J. J. Dongarra, “Automatically tuned linear algebra software,” in Proc. 1998 ACM/IEEE Conf. Supercomputing (SC’98). IEEE, 1998, p. 38. https://doi.org/10.1109/SC.1998.10004
  19. S. A. Hassan, M. M. Mahmoud, A. M. Hemeida, and M. A. Saber, “Effective implementation of matrix–vector multiplication on Intel’s AVX multicore processor,” Comput. Lang. Syst. Struct., vol. 51, pp. 158– 175, 2017. https://doi.org/10.1016/j.cl.2017.06.003
  20. A. Abdelfattah, D. Keyes, and H. Ltaief, “KBLAS: An optimized library for dense matrix-vector multiplication on GPU accelerators,” ACM Trans. Math. Softw., vol. 42, no. 3, p. 18, 2016. https://doi.org/ 10.1145/2818311
  21. J. Liang and Y. Zhang, “Optimization of GEMV on Intel AVX processor,” Int. J. Database Theory Appl., vol. 9, no. 2, pp. 47–60, 2016. https://doi.org/10.14257/ijdta.2016.9.2.06
  22. A. Singh and C. Bassoy, “High-Performance Level-1 and Level2 BLAS,” arXiv:2108.02025 [cs.MS], 2021. https://doi.org/10.48550/ arXiv.2108.02025
  23. B. Bylina, J. Bylina, and M. Piekarz, “Influence of loop transformations on performance and energy consumption of the multithreaded WZ factorization,” in Proc. 17th Conf. Computer Science and Intelligence Systems (FedCSIS), Ann. Comput. Sci. Inf. Syst., vol. 30, pp. 479–488, 2022. https://doi.org/10.15439/2022F251
  24. T. Jakobs and G. Rünger, “On the energy consumption of Load/Store AVX instructions,” in Proc. 2018 Federated Conf. Computer Science and Information Systems (FedCSIS), Ann. Comput. Sci. Inf. Syst., vol. 15, pp. 319–327, 2018. https://doi.org/10.15439/2018F28
  25. Free Software Foundation, “Using the GNU Compiler Collection (GCC),” version 15.2.0, 2025. https://gcc.gnu.org/onlinedocs/gcc-15.2. 0/gcc/
  26. R. Chandra, L. Dagum, D. Kohr, R. Menon, D. Maydan, and J. McDonald, Parallel Programming in OpenMP. Morgan Kaufmann, 2001.