[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:lmcache":3},"\u003Cdiv align=\"center\">\n  \u003Cp align=\"center\">\n    \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FLMCache\u002Flmcache\u002FHEAD\u002Fasset\u002Flogo.png\" alt=\"lmcache logo\" width=\"45%\" \u002F>\n  \u003C\u002Fp>\n  \u003Ch3>\n    A KV Cache Management Layer for Scalable LLM Inference\n  \u003C\u002Fh3>\n    \u003Chr \u002F>  \u003Ch3>\n    \u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002F\" rel=\"nofollow ugc noopener\">Blog\u003C\u002Fa> |\n    \u003Ca href=\"https:\u002F\u002Fdocs.lmcache.ai\u002F\" rel=\"nofollow ugc noopener\">Documentation\u003C\u002Fa> |\n    \u003Ca href=\"https:\u002F\u002Fjoin.slack.com\u002Ft\u002Flmcacheworkspace\u002Fshared_invite\u002Fzt-3zxjao8h0-lRfBfnLqbALOtLsWn2ITxA\" rel=\"nofollow ugc noopener\">Join Slack\u003C\u002Fa> |\n    \u003Ca href=\"https:\u002F\u002Fdocs.lmcache.ai\u002Fcommunity\u002Fmeetings.html\" rel=\"nofollow ugc noopener\">Community Meeting\u003C\u002Fa> |\n    \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FLMCache\u002FLMCache\u002Fissues\u002F2923\" rel=\"nofollow ugc noopener\">Roadmap\u003C\u002Fa>\n  \u003C\u002Fh3>\u003Cp>  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FLMCache\u002FLMCache\u002Fstargazers\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fstars\u002FLMCache\u002FLMCache?style=flat&amp;logo=github\" alt=\"GitHub Repo stars\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Flmcache\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Flmcache\" alt=\"PyPI\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Flmcache\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fdm\u002Flmcache\" alt=\"PyPI - Downloads\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FLMCache\u002FLMCache\u002Fgraphs\u002Fcommit-activity\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fcommit-activity\u002Fw\u002FLMCache\u002FLMCache\" alt=\"GitHub commit activity\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fdeepwiki.com\u002FLMCache\u002FLMCache\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fdeepwiki.com\u002Fbadge.svg\" alt=\"Ask DeepWiki\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>  ⭐ \u003Cstrong>If LMCache helps you serve LLMs faster and cheaper, \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FLMCache\u002FLMCache\" rel=\"nofollow ugc noopener\">give us a star\u003C\u002Fa> — it helps more teams discover the project.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003C\u002Fdiv>\u003Ch2>Updates\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>[2026\u002F05] 🔥 Agentic workload benchmark on AMD MI300X (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2026\u002F05\u002F12\u002Fbenchmarking-lmcache-for-multi-turn-agentic-workloads-on-amd-mi300x\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2026\u002F04] 🔥 LMCache's new multiprocess (MP) architecture release (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2026\u002F04\u002F03\u002Flmcaches-new-architecture-boosts-moe-inference-performance-by-10x\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2026\u002F03] LMCache at GTC 2026 (\u003Ca href=\"https:\u002F\u002Fwww.linkedin.com\u002Fposts\u002Flmcache-lab_llm-opensource-nvidiagtc-activity-7442721875664826369-pMAu?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAADkIIvQBTyG53kXXX70OZdE5rhpllYQqmIA\" rel=\"nofollow ugc noopener\">post\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2026\u002F01] LMCache multi-node P2P CPU memory sharing, from experimental feature to production (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2026\u002F01\u002F21\u002Fp2p-1\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cdetails>\n\u003Csummary>More\u003C\u002Fsummary>\u003Cul>\n\u003Cli>[2025\u002F11] LMCache x CoreWeave accelerate efficient LLM inference for Cohere (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F10\u002F29\u002Fbreaking-the-memory-barrier-how-lmcache-and-coreweave-power-efficient-llm-inference-for-cohere\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F10] LMCache joins the PyTorch Foundation and Tensormesh unveiled (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F10\u002F31\u002Ftensormesh-unveiled-and-lmcache-joins-the-pytorch-foundation\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fpytorch.org\u002Fblog\u002Flmcache-joins-pytorch-ecosystem\u002F\" rel=\"nofollow ugc noopener\">PyTorch\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F09] NVIDIA Dynamo integrates LMCache, accelerating LLM inference (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F09\u002F18\u002Fnvidia-dynamo-integrates-lmcache-accelerating-llm-inference\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F08] 🎉 LMCache hits 5,000+ GitHub stars (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F08\u002F28\u002F%f0%9f%8e%89-lmcache-hits-5000-github-stars-thank-you-community\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F08] LMCache supports gpt-oss (20B\u002F120B) on day 1 (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F08\u002F05\u002Flmcache-supports-gpt-oss-20b-120b-on-day-1\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F07] Get faster LLM inference and cheaper responses with LMCache and Redis (\u003Ca href=\"https:\u002F\u002Fredis.io\u002Fblog\u002Fget-faster-llm-inference-and-cheaper-responses-with-lmcache-and-redis\u002F\" rel=\"nofollow ugc noopener\">Redis blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F07] LMCache extends its turbo-boost to multimodal models in vLLM V1 (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F07\u002F03\u002Flmcache-extends-its-turbo-boost-to-multimodal-models-in-vllm-v1\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003Cli>[2025\u002F06] LLM Production Stack goes cross-hardware: AMD, Arm and Ascend (\u003Ca href=\"https:\u002F\u002Fblog.lmcache.ai\u002Fen\u002F2025\u002F06\u002F20\u002Fllm-production-stack-goes-cross-hardware-ascend-arm-and-amd-support-incoming\u002F\" rel=\"nofollow ugc noopener\">blog\u003C\u002Fa>).\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fdetails>\u003Ch2>About\u003C\u002Fh2>\n\u003Cp>LMCache is a \u003Cstrong>KV cache management layer\u003C\u002Fstrong> for LLM inference. It turns KV cache from a temporary state into reusable \u003Cem>AI-native knowledge\u003C\u002Fem> that can be \u003Cem>stored\u003C\u002Fem> persistently, \u003Cem>reused\u003C\u002Fem> across multiple serving engines, \u003Cem>monitored\u003C\u002Fem> with an observabi\u003C\u002Fp>\n",1784564568165]