The Sriwijaya University Library

  • Home
  • Information
  • News
  • Help
  • Login
  • Librarian
  • Member Area
  • Select Language :
    Arabic Bengali Brazilian Portuguese English Espanol German Indonesian Japanese Malay Persian Russian Thai Turkish Urdu

Search by :

ALL Author Subject ISBN/ISSN Advanced Search

Last search:

{{tmpObj[k].text}}
Image of OPTIMASI KERNEL CUDA BERBASIS MEMORY-AWARE UNTUK SELF-ATTENTION PADA INFERENSI LLM
Bookmark Share

Skripsi

OPTIMASI KERNEL CUDA BERBASIS MEMORY-AWARE UNTUK SELF-ATTENTION PADA INFERENSI LLM

Sofhiea, Defhanaya - Personal Name;

The self-attention mechanism in Large Language Models (LLMs) is computationally intensive and memory-bound, posing significant challenges for inference on consumer-grade hardware. This research proposes a memory-aware optimization of CUDA kernels for the self-attention mechanism within the GPT-2 architecture, integrating Shared Memory Tiling, Streaming Softmax, and Kernel Fusion to minimize global memory traffic and reduce kernel launch overhead. The implementation was developed using CUDA C++ and Python, validated against a PyTorch baseline on an NVIDIA RTX 3050 GPU. Experimental results on matrices of varying sizes show consistent optimization gains, with the 4096 × 4096 configuration yielding the most pronounced improvements: the fully fused kernel achieves a 3.44× speedup over the naive baseline (mean execution time reduced from 24.86 ms to 7.22 ms, averaged over 100 runs) and reduces peak memory usage by 89.3% (from 141.12 MB to 15.12 MB) by eliminating intermediate attention matrices. Numerical stability was confirmed with a negligible Mean Absolute Error (MAE) of approximately 10⁻⁸. These findings demonstrate that low-level kernel fusion significantly enhances memory efficiency and inference speed on consumer-grade GPUs.


Availability
#
Central Library (Reference) T1955422026
T195542
Available but not for loan - Not for Loan
Detail Information
Series Title
-
Call Number
T1955422026
Publisher
Indralaya : Prodi Teknik Informatika, Fakultas Ilmu Komputer Universitas Sriwijaya., 2026
Collation
xviii, 68 hlm.; ilus.; 29 cm
Language
Indonesia
ISBN/ISSN
-
Classification
005.180 7
Content Type
Text
Media Type
unmediated
Carrier Type
-
Edition
-
Subject(s)
Prodi Teknik Informatika
CUDA (Arsitektur Komputer)
Specific Detail Info
-
Statement of Responsibility
KA
Other version/related
TitleEditionLanguage
PENGARUH PEMBERIAN ZAT PENGATUR TUMBUH ALAMI PADA BERBAGAI MACAM ASAL BAHAN STEK TERHADAP PERTUMBUHAN TANAMAN CHAYA (Cnidoscolus aconitifolius) VAR. PICUDAid
RESPON TANAMAN CHAYA VARIETAS PICUDA (CNIDOSCOLUS ACONITIFOLIUS VAR.PICUDA) TERHADAP PUPUK KANDANG SAPI DAN PUPUK ORGANIK CAIR MEGA RHIZO PADA TANAH RAWA LEBAK-id
File Attachment
  • OPTIMASI KERNEL CUDA BERBASIS MEMORY-AWARE UNTUK SELF-ATTENTION PADA INFERENSI LLM
Comments

You must be logged in to post a comment

The Sriwijaya University Library
  • Information
  • Services
  • Librarian
  • Member Area

About Us

As a complete Library Management System, SLiMS (Senayan Library Management System) has many features that will help libraries and librarians to do their job easily and quickly. Follow this link to show some features provided by SLiMS.

Search

start it by typing one or more keywords for title, author or subject

Keep SLiMS Alive Want to Contribute?

© 2026 — Senayan Developer Community

Powered by SLiMS
Select the topic you are interested in
  • Computer Science, Information & General Works
  • Philosophy & Psychology
  • Religion
  • Social Sciences
  • Language
  • Pure Science
  • Applied Sciences
  • Art & Recreation
  • Literature
  • History & Geography
Icons made by Freepik from www.flaticon.com
Advanced Search
Where do you want to share?