PagedAttention is an attention algorithm for serving large language models efficiently, which manages the key-value cache used during autoregressive decoding by storing it in fixed-size blocks that can be mapped to non-contiguous physical memory, an approach inspired by virtual memory paging in operating systems. It was introduced by Woosuk Kwon and colleagues in the 2023 paper Efficient Memory Management for Large Language Model Serving with PagedAttention, alongside the vLLM serving engine built around it. Earlier LLM serving systems reserved a contiguous block of cache memory for each request in advance, which the researchers found left effective memory utilization as low as 20.4 percent because of reserved but unused space and fragmentation. The block-based allocation used by PagedAttention allows the cache to be shared across related decoding paths and improved throughput by two to four times while cutting memory overhead by 6 to 55 percent, depending on the workload, compared with earlier systems. This description is adapted from Wikipedia contributors under CC BY-SA 4.0; changes were made. https://creativecommons.org/licenses/by-sa/4.0/
Sources
Wikipedia: PagedAttention
Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.