Skip to main content

#lmcache

1

문서

0

카테고리

1k

단어

4

분 읽기

관련 카테고리: 없음

문서 목록

The concept of LMCache — offloading KV cache beyond GPU memory to CPU and disk and sharing it across inference instances — and its relationship to vLLM prefix cache, NIXL, and kvaware routing.