TY - ECHAP AU - M. Karwacki AU - Przemysław Stpiczyński AU - K. De Bosschere AU - E.H. D`Hollander AU - G.R. Joubert AU - D. Padna AU - F. Peters AU - M. Sawyer AB - CUBLAS is a widely used implementation of BLAS (Basic Linear Algebra Subprograms) for NVIDIA CUDA Graphical Processing Units (GPUs). The aim of this paper is to show that the performance of the selected Level 2 BLAS routines for working with triangular matrices can be improved using some optimization techniques suitable for GPUs like using shared memory and coalesced memory access. We present new implementation of the routines xTRMV and xTRSV. The results of experiments carried out on two GPU architectures: Tesla M2050 and GeForce GTX 260 show that these new implementations are up to 500% faster than corresponding routines from CUBLAS Library. BT - Advances in Parallel Computing LA - eng N1 - DOI: 10.3233/978-1-61499-041-3-405 N2 - CUBLAS is a widely used implementation of BLAS (Basic Linear Algebra Subprograms) for NVIDIA CUDA Graphical Processing Units (GPUs). The aim of this paper is to show that the performance of the selected Level 2 BLAS routines for working with triangular matrices can be improved using some optimization techniques suitable for GPUs like using shared memory and coalesced memory access. We present new implementation of the routines xTRMV and xTRSV. The results of experiments carried out on two GPU architectures: Tesla M2050 and GeForce GTX 260 show that these new implementations are up to 500% faster than corresponding routines from CUBLAS Library. PB - IOS Press PY - 2012 SE - Improving performance of triangular Matrix-Vector BLAS routines on GPUs. SN - 978-1-61499-040-6 T2 - Advances in Parallel Computing TI - Applications and Techniques on the Road to Exascale Computing VL - 22 ER -