01170nas a2200217 4500000000100000000000100001008004100002260001400043100001600057700002900073700002000102700002100122700001700143700001300160700001400173700001400187245006600201490000700267520065600274020002200930 2012 d bIOS Press1 aM. Karwacki1 aPrzemysław Stpiczyński1 aK. De Bosschere1 aE.H. D`Hollander1 aG.R. Joubert1 aD. Padna1 aF. Peters1 aM. Sawyer00aApplications and Techniques on the Road to Exascale Computing0 v223 aCUBLAS is a widely used implementation of BLAS (Basic Linear Algebra Subprograms) for NVIDIA CUDA Graphical Processing Units (GPUs). The aim of this paper is to show that the performance of the selected Level 2 BLAS routines for working with triangular matrices can be improved using some optimization techniques suitable for GPUs like using shared memory and coalesced memory access. We present new implementation of the routines xTRMV and xTRSV. The results of experiments carried out on two GPU architectures: Tesla M2050 and GeForce GTX 260 show that these new implementations are up to 500% faster than corresponding routines from CUBLAS Library. a978-1-61499-040-6