High-Performance GEMV on AMD Zen 3 via Architecture-Specific AVX2/FMA Optimization
Rafał Lenart, Beata Bylina, Jarosław Bylina
DOI: http://dx.doi.org/10.15439/2026F6278
Citation: Rafał Lenart, Beata Bylina, Jarosław Bylina (2026). High-Performance GEMV on AMD Zen 3 via Architecture-Specific AVX2/FMA Optimization. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS). ACSIS, Vol. 47, pages 357–364.
Abstract. The General Matrix-Vector Multiply (GEMV) is a fundamental Level-2 BLAS operation appearing throughout scientific computing and numerical linear algebra. GEMV is inherently memory-bound, making effective cache use the dominant optimization factor. We present a hand-optimized GEMV kernel for AMD Zen~3, implemented using AVX2 and FMA intrinsics. Our single-threaded kernel (ZenGEMV) processes 4~rows simultaneously with dual accumulators per row, creating 8~independent FMA chains that fully saturate the Zen~3 pipeline. Benchmarked against OpenBLAS~0.3.31 and AOCL-BLAS~5.2.0 on an AMD Ryzen~5~5600, ZenGEMV achieves up to $1.37\times$ speedup over OpenBLAS and $1.20\times$ over AOCL-BLAS. Our parallel variant (ZenGEMV-P) achieves up to $3.22\times$ speedup over OpenBLAS on small matrices. We analyse why dispatch overhead, suboptimal accumulator strategies, and generic design choices in general-purpose BLAS libraries leave performance gains available in architecture-specific workloads.
References
- L. S. Blackford, A. Petitet, R. Pozo, et al., “An updated set of basic linear algebra subprograms (BLAS),” ACM Trans. Math. Softw., vol. 28, no. 2, pp. 135–151, 2002. https://doi.org/10.1145/567806.567807
- G. H. Golub and C. F. Van Loan, Matrix Computations, 4th ed. Baltimore: Johns Hopkins University Press, 2013.
- C. He, Y. Huang, P. Mu, Z. Miao, J. Xue, L. Ma, et al., “WaferLLM: Large language model inference at wafer scale,” https://arxiv.org/abs/2502.04563 [cs.LG], 2025. https://doi.org/10.48550/arXiv.2502.04563
- C. Holmes, D. Mawhirter, Y. He, F. Yan, and B. Wu, “GRNN: Lowlatency and scalable RNN inference on GPUs,” in Proc. 14th EuroSys Conf. (EuroSys’19), New York: ACM, 2019, pp. 41:1–41:16. https://doi. org/10.1145/3302424.3303949
- OpenBLAS: An optimized BLAS library. https://www.openblas.net, https://github.com/OpenMathLib/OpenBLAS
- AMD, “AOCL-BLAS (BLIS fork with AMD optimizations),” version 5.2.0. https://github.com/amd/blis
- P. Gepner, “Using AVX2 instruction set to increase performance of high performance computing code,” Computing and Informatics, vol. 36, no. 5, pp. 1001–1018, 2017. https://doi.org/10.4149/cai_2017_5_1001
- H. Zhang, D. Chen, and S. B. Ko, “Efficient multiple-precision floatingpoint fused multiply-add with mixed-precision support,” IEEE Trans. Comput., vol. 68, no. 7, pp. 1035–1048, 2019. https://doi.org/10.1109/ TC.2019.2895031
- M. Evers, L. Barnes, and M. Clark, “The AMD next-generation ‘Zen 3’ core,” IEEE Micro, vol. 42, no. 3, pp. 7–12, 2022. https://doi.org/10. 1109/MM.2022.3152788
- S. Williams, A. Waterman, and D. A. Patterson, “Roofline: an insightful visual performance model for multicore architectures,” Commun. ACM, vol. 52, pp. 65–76, 2009. https://doi.org/10.1145/1498765.1498785
- Advanced Micro Devices, “Software optimization guide for AMD EPYC 7003 series processors (family 19h models 00h–0Fh),” rev. 3.00. https://kib.kiev.ua/x86docs/AMD/Optimization/56665_3.00_ Software%20Optimization%20Guide%20for%20AMD%20Family% 2019h%20Processors%20%28PUB%29.pdf
- A. Fog, “Instruction tables: lists of instruction latencies, throughputs and micro-operation breakdowns for Intel, AMD, and VIA CPUs,” 2011, last updated Sep. 20, 2025. https://www.agner.org/optimize/instruction_ tables.pdf
- A. Fog, “The microarchitecture of Intel, AMD, and VIA CPUs: an optimization guide for assembly programmers and compiler makers,” 2017, last updated Dec. 15, 2025. https://www.agner.org/optimize/ microarchitecture.pdf
- FinlayDaG33k, “RAM bandwidth calculator.” https://edu.finlaydag33k. nl/calculating%20ram%20bandwidth/, https://gitlab.com/FinlayDaG33k/ ram-calculator
- Z. Smith, “bandwidth: a memory bandwidth benchmark.” https://zs3.me/ bandwidth
- K. Goto and R. A. van de Geijn, “Anatomy of high-performance matrix multiplication,” ACM Trans. Math. Softw., vol. 34, no. 3, p. 12, 2008. https://doi.org/10.1145/1356052.1356053
- F. G. Van Zee and R. A. van de Geijn, “BLIS: A framework for rapidly instantiating BLAS functionality,” ACM Trans. Math. Softw., vol. 41, no. 3, p. 14, 2015. https://doi.org/10.1145/2764454
- R. C. Whaley and J. J. Dongarra, “Automatically tuned linear algebra software,” in Proc. 1998 ACM/IEEE Conf. Supercomputing (SC’98). IEEE, 1998, p. 38. https://doi.org/10.1109/SC.1998.10004
- S. A. Hassan, M. M. Mahmoud, A. M. Hemeida, and M. A. Saber, “Effective implementation of matrix–vector multiplication on Intel’s AVX multicore processor,” Comput. Lang. Syst. Struct., vol. 51, pp. 158– 175, 2017. https://doi.org/10.1016/j.cl.2017.06.003
- A. Abdelfattah, D. Keyes, and H. Ltaief, “KBLAS: An optimized library for dense matrix-vector multiplication on GPU accelerators,” ACM Trans. Math. Softw., vol. 42, no. 3, p. 18, 2016. https://doi.org/ 10.1145/2818311
- J. Liang and Y. Zhang, “Optimization of GEMV on Intel AVX processor,” Int. J. Database Theory Appl., vol. 9, no. 2, pp. 47–60, 2016. https://doi.org/10.14257/ijdta.2016.9.2.06
- A. Singh and C. Bassoy, “High-Performance Level-1 and Level2 BLAS,” arXiv:2108.02025 [cs.MS], 2021. https://doi.org/10.48550/ arXiv.2108.02025
- B. Bylina, J. Bylina, and M. Piekarz, “Influence of loop transformations on performance and energy consumption of the multithreaded WZ factorization,” in Proc. 17th Conf. Computer Science and Intelligence Systems (FedCSIS), Ann. Comput. Sci. Inf. Syst., vol. 30, pp. 479–488, 2022. https://doi.org/10.15439/2022F251
- T. Jakobs and G. Rünger, “On the energy consumption of Load/Store AVX instructions,” in Proc. 2018 Federated Conf. Computer Science and Information Systems (FedCSIS), Ann. Comput. Sci. Inf. Syst., vol. 15, pp. 319–327, 2018. https://doi.org/10.15439/2018F28
- Free Software Foundation, “Using the GNU Compiler Collection (GCC),” version 15.2.0, 2025. https://gcc.gnu.org/onlinedocs/gcc-15.2. 0/gcc/
- R. Chandra, L. Dagum, D. Kohr, R. Menon, D. Maydan, and J. McDonald, Parallel Programming in OpenMP. Morgan Kaufmann, 2001.