[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:kvcached":3},"\u003Cdiv align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fovg-project\u002Fkvcached\u002Frefs\u002Fheads\u002Fmain\u002Fassets\u002Flogo-v2.svg\" alt=\"kvcached logo\" height=\"96\" \u002F>  \u003Cbr \u002F>\n  \u003Cbr \u002F>\n  \u003Cp>\n    \u003Ca href=\"https:\u002F\u002Fwww.python.org\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"Python\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPython-3.9%E2%80%933.13-blue\" \u002F>\u003C\u002Fa>\n    \u003Cimg alt=\"Engines\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FEngines-SGLang%20%7C%20vLLM-blueviolet\" \u002F>\n    \u003Ca href=\"https:\u002F\u002Fyifanqiao.notion.site\u002FSolve-the-GPU-Cost-Crisis-with-kvcached-289da9d1f4d68034b17bf2774201b141\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"Blog\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FBlog-Read-FF5722?logo=rss&amp;logoColor=white&amp;labelColor=555555\" \u002F>\u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2508.08448\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"arXiv: GPU OS vision\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-GPU%20OS%20vision-b31b1b?logo=arxiv&amp;logoColor=white&amp;labelColor=555555\" \u002F>\u003C\u002Fa>\n    \u003Cbr \u002F>\n    \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2505.04021\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"arXiv: Multi LLM Serving\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-Multi%20LLM%20Serving-b31b1b?logo=arxiv&amp;logoColor=white&amp;labelColor=555555\" \u002F>\u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Fjoin.slack.com\u002Ft\u002Fovg-project\u002Fshared_invite\u002Fzt-3fr01t8s7-ZtDhHSJQ00hcLHgwKx3Dmw\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"Slack Join\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FSlack-Join-4A154B?logo=slack&amp;logoColor=white&amp;labelColor=555555\" \u002F>\u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Fdeepwiki.com\u002Fovg-project\u002Fkvcached\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"DeepWiki\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDeepWiki-Docs-6B46C1?logo=book&amp;logoColor=white&amp;labelColor=555555\" \u002F>\u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Fkvcached.org\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"Homepage\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FHomepage-kvcached.org-0A66C2?logo=internetexplorer&amp;logoColor=white&amp;labelColor=555555\" \u002F>\u003C\u002Fa>\n    \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fovg-project\u002Fkvcached\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg alt=\"License\" src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache_2.0-blue.svg\" \u002F>\u003C\u002Fa>\n  \u003C\u002Fp>\u003C\u002Fdiv>\u003Ch2>Make GPU Sharing Flexible and Easy \u003C\u002Fh2>\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fovg-project\u002Fkvcached\u002Frefs\u002Fheads\u002Fmain\u002Fassets\u002Fads.jpg\" alt=\"Make GPU Sharing Flexible and Easy\" width=\"500\" \u002F>\n\u003C\u002Fp>\u003Cp>kvcached (KV cache daemon) is a KV cache library for LLM serving\u002Ftraining on \u003Cstrong>shared GPUs\u003C\u002Fstrong>.  By bringing OS-style \u003Cstrong>virtual memory\u003C\u002Fstrong> abstraction to LLM systems, it enables \u003Cstrong>elastic and demand-driven\u003C\u002Fstrong> KV cache allocation, improving GPU utilization under dynamic workloads.\u003C\u002Fp>\n\u003Cp>kvcached achieves this by decoupling GPU virtual addressing from physical memory allocation for KV caches. It allows serving engines to initially reserve virtual memory only and later back it with physical GPU memory when the cache is actively used. This decoupling enables on-demand allocation and flexible sharing, bringing better GPU memory utilization under dynamic and mixed workloads. Check out more details in the \u003Ca href=\"https:\u002F\u002Fyifanqiao.notion.site\u002FSolve-the-GPU-Cost-Crisis-with-kvcached-289da9d1f4d68034b17bf2774201b141\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>.\u003C\u002Fp>\n\u003Ch3>Key Features\u003C\u002Fh3>\u003Cul>\n\u003Cli>\u003Cstrong>Elastic KV cache\u003C\u002Fstrong>: allocate and reclaim KV memory dynamically to match live load.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>GPU virtual memory\u003C\u002Fstrong>: decouple logical KV from physical GPU memory via runtime mapping.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Memory control CLI\u003C\u002Fstrong>: enforce memory limits with kvcached CLI.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Frontend router and sleep mode\u003C\u002Fstrong>: route requests to the target models and put models to sleep when idle.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Support mainstream serving engines\u003C\u002Fstrong>: integrate with SGLang and vLLM.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Prefix caching\u003C\u002Fstrong>: support automatic prefix caching (APC) with a configurable memory bound. See \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fovg-project\u002Fkvcached\u002Fblob\u002FHEAD\u002Fexamples\u002F09_prefix_caching\" rel=\"nofollow ugc noopener\">the example doc\u003C\u002Fa> for details.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>📢 Updates\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>[2026-04]\u003C\u002Fstrong> kvcached is \u003Cstrong>featured by Red Hat\u003C\u002Fstrong> for running LLMs dynamically in production under limited resources! Red Hat's \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Frh-aiservices-bu\u002Fsardeenz\" rel=\"nofollow ugc noopener\">Sardeenz\u003C\u002Fa> builds on kvcached to provide dynamic multi-model serving with Kubernetes and OpenShift support. See the [blog post](\u003Ca href=\"https:\u002F\u002Fwww.redhat.com\u002Fen\u002Fblog\u002Frunning-llms-dynamically-produc\" rel=\"nofollow ugc noopener\">https:\u002F\u002Fwww.redhat.com\u002Fen\u002Fblog\u002Frunning-llms-dynamically-produc\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n",1788652555457]