[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:agentic-api":3},"\u003Cdiv align=\"center\">\u003Cpicture>\n  \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\"assets\u002FWhite-Main-Logo.svg\">\u003C\u002Fsource>\n  \u003Csource media=\"(prefers-color-scheme: light)\" srcset=\"assets\u002FBlack-Main-Logo.svg\">\u003C\u002Fsource>\n  \u003Cimg alt=\"Agentic API\" src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fvllm-project\u002Fagentic-api\u002FHEAD\u002Fassets\u002FBlack-Main-Logo.svg\" width=\"600\" \u002F>\n\u003C\u002Fpicture>\u003Cp>\u003Cstrong>The stateful, agentic API layer for \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fvllm\" rel=\"nofollow ugc noopener\">vLLM\u003C\u002Fa>, written in Rust 🦀\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>\u003Cem>Run OpenAI-grade agentic workloads (Responses API, server-side tools, Codex) on your own GPUs.\u003C\u002Fem>\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fagentic-api\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-Apache%202.0-blue.svg\" alt=\"License\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fagentic-api\u002Fblob\u002FHEAD\u002FCargo.toml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FRust-1.85%2B-orange.svg?logo=rust\" alt=\"Rust\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fagentic-api\u002Factions\u002Fworkflows\u002Frust.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fagentic-api\u002Factions\u002Fworkflows\u002Frust.yml\u002Fbadge.svg\" alt=\"CI\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fagentic-api\u002Factions\u002Fworkflows\u002Fpre-commit.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fvllm-project\u002Fagentic-api\u002Factions\u002Fworkflows\u002Fpre-commit.yml\u002Fbadge.svg\" alt=\"pre-commit\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fdiv>\u003Chr \u002F>\n\u003Ch2>🧠 Overview\u003C\u002Fh2>\n\u003Cp>vLLM gives you state-of-the-art inference throughput. But real agentic applications need more than raw tokens: they need \u003Cstrong>conversation state, tool-call loops, and multi-turn orchestration\u003C\u002Fstrong>. Today, all of that complexity lives in your client code.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Agentic API moves it server-side.\u003C\u002Fstrong> It is a Rust-native gateway that sits in front of vLLM and owns the stateful agentic APIs, starting with an OpenAI-compatible \u003Ca href=\"https:\u002F\u002Fplatform.openai.com\u002Fdocs\u002Fapi-reference\u002Fresponses\" rel=\"nofollow ugc noopener\">Responses API\u003C\u002Fa>. vLLM is one supported backend, not part of the Agentic API product name. Your application makes \u003Cem>one API call\u003C\u002Fem> and the server handles the rest: state hydration, tool execution, streaming, and continuation.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-mermaid\">flowchart LR\n    C([\"🧑‍💻 Client&lt;br\u002F&gt;Codex · SDKs · curl\"]) --&gt;|\"📮 &lt;code&gt;POST \u002Fv1\u002Fresponses&lt;\u002Fcode&gt;&lt;br\u002F&gt;🌐 HTTP&amp;nbsp;&amp;nbsp;📡 SSE&amp;nbsp;&amp;nbsp;🔌 WebSocket\"| A\n    subgraph A [\"⚡ Agentic API (Rust 🦀)\"]\n        direction TB\n        S[\"🔄 State hydration&lt;br\u002F&gt;&lt;code&gt;previous_response_id&lt;\u002Fcode&gt;\"]\n        T[\"🛠️ Server-side tools&lt;br\u002F&gt;web search · functions\"]\n        P[\"💾 Persistence&lt;br\u002F&gt;SQLite response store\"]\n    end\n    A --&gt;|\"🚀 &lt;code&gt;POST \u002Fv1\u002Fresponses&lt;\u002Fcode&gt;&lt;br\u002F&gt;⚙️ stateless&amp;nbsp;&amp;nbsp;🤝 OpenAI-compatible\"| V([\"🚀 vLLM core&lt;br\u002F&gt;inference engine\"])\n\n    classDef client fill:#FFE8B3,stroke:#F59E0B,stroke-width:2px,color:#7C2D12\n    classDef inner fill:#E0E7FF,stroke:#6366F1,stroke-width:2px,color:#312E81\n    classDef engine fill:#DCFCE7,stroke:#22C55E,stroke-width:2px,color:#14532D\n\n    class C client\n    class S,T,P inner\n    class V engine\n    style A fill:#F5F3FF,stroke:#8B5CF6,stroke-width:2px,color:#5B21B6\n    linkStyle 0 stroke:#F59E0B,stroke-width:2px\n    linkStyle 1 stroke:#22C55E,stroke-width:2px\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cblockquote>\n\u003Cp>[!TIP]\nPoint \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fopenai\u002Fcodex\" rel=\"nofollow ugc noopener\">OpenAI Codex\u003C\u002Fa> at Agentic API and drive it entirely with open models served by vLLM. No OpenAI account required.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Ch2>✨ Key Features\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>🔄 \u003Cstrong>Stateful conversations\u003C\u002Fstrong>: the server manages history via \u003Ccode>previous_response_id\u003C\u002Fcode>. No client-side message tracking, no replaying full transcripts.\u003C\u002Fli>\n\u003Cli>🛠️ \u003Cstrong>Server-side tool execution\u003C\u002Fstrong>: an explicit tool-ownership model (gateway \u002F client \u002F provider) decides exactly what runs where. Web search ships today via \u003Ca href=\"https:\u002F\u002Fyou.com\" rel=\"nofollow ugc noopener\">You.com\u003C\u002Fa>, and the model executes multi-step tool chains automatically.\u003C\u002Fli>\n\u003Cli>📡 \u003Cstrong>Every transport\u003C\u002Fstrong>: non-streaming HTTP, server-sent events for token streaming, and full \u003Cstrong>WebSocket\u003C\u002Fstrong> support for interactive clients.\u003C\u002Fli>\n\u003Cli>🧰 \u003Cstrong>Codex-ready\u003C\u002Fstrong>: accepts Codex-shaped Responses traffic out of the box, preserving the tool declarations and response item shapes Codex depends on.\u003C\u002Fli>\n\u003Cli>🏃 \u003Cstrong>Background execution\u003C\u002Fstrong>: fire-and-forget requests that keep processing server-side.\u003C\u002Fli>\n\u003Cli>✅ \u003Cstrong>Compatibility tested\u003C\u002Fstrong>: validated against the \u003Ca href=\"https:\u002F\u002Fwww.openresponses.org\u002F\" rel=\"nofollow ugc noopener\">Open Responses\u003C\u002Fa> compatibility suite, with replay-cassette tests for real OpenAI and vLLM traffic.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>🧭 API Surface\u003C\u002Fh2>\n",1788652552322]