[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:speech-to-speech":3},"\u003Cdiv align=\"center\">\n  \u003Cdiv> \u003C\u002Fdiv>\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhuggingface\u002Fspeech-to-speech\u002Fmain\u002Flogo.png\" width=\"600\" \u002F>\u003Ch1>Speech To Speech: Build voice agents with open-source models\u003C\u002Fh1>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fspeech-to-speech\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fspeech-to-speech\" alt=\"PyPI\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fspeech-to-speech\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fpyversions\u002Fspeech-to-speech\" alt=\"Python\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fhuggingface\u002Fspeech-to-speech\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-Apache%202.0-blue\" alt=\"License\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Ftrendshift.io\u002Frepositories\u002F20645\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FGitHub%20Trending-%231%20Repository%20of%20the%20Day-7B2CBF?logo=github&amp;logoColor=white\" alt=\"GitHub Trending: #1 Repository of the Day\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fdiv>\u003Cp>A low-latency, fully modular voice-agent pipeline: \u003Cstrong>VAD -&gt; STT -&gt; LLM -&gt; TTS\u003C\u002Fstrong>, exposed through an \u003Cstrong>OpenAI Realtime-compatible WebSocket API\u003C\u002Fstrong>. Every component is swappable. The LLM slot speaks OpenAI-compatible protocols, so you can point it at a hosted provider, at \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Finference-providers\" rel=\"nofollow ugc noopener\">HF Inference Providers\u003C\u002Fa>, or at a vLLM or llama.cpp server on your own hardware for a fully local, fully open stack.\u003C\u002Fp>\n\u003Cp>This pipeline runs in production as the conversation backend for thousands of \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fblog\u002Freachy-mini\" rel=\"nofollow ugc noopener\">Reachy Mini\u003C\u002Fa> robots.\u003C\u002Fp>\n\u003Cp align=\"center\">\n  \u003Cpicture>\n    \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\".\u002Fdocs\u002Fassets\u002Fendpoint-swap-dark.gif\">\u003C\u002Fsource>\n    \u003Csource media=\"(prefers-color-scheme: light)\" srcset=\".\u002Fdocs\u002Fassets\u002Fendpoint-swap-light.gif\">\u003C\u002Fsource>\n    \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhuggingface\u002Fspeech-to-speech\u002FHEAD\u002Fdocs\u002Fassets\u002Fendpoint-swap-light.gif\" alt=\"Switching an OpenAI Realtime client endpoint from hosted OpenAI to a self-hosted speech-to-speech server\" width=\"640\" \u002F>\n  \u003C\u002Fpicture>\n\u003C\u002Fp>\u003Ch2>Quickstart\u003C\u002Fh2>\n\u003Cpre>\u003Ccode class=\"language-bash\">pip install speech-to-speech\nexport OPENAI_API_KEY=...\nspeech-to-speech\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This starts an OpenAI Realtime-compatible server at \u003Ccode>ws:\u002F\u002Flocalhost:8765\u002Fv1\u002Frealtime\u003C\u002Fcode> using Parakeet TDT for local STT, an OpenAI-compatible LLM, and Qwen3-TTS for local speech output.\u003C\u002Fp>\n\u003Cp>From a source checkout, talk to it from a second terminal:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">python scripts\u002Flisten_and_play_realtime.py --host 127.0.0.1 --port 8765\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Prefer to keep the LLM on your own machine? Serve Gemma 4 with llama.cpp:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">llama-server -hf ggml-org\u002Fgemma-4-E4B-it-GGUF -np 2 -c 65536 -fa on --swa-full\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Then point the OpenAI-compatible LLM backend at it:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">speech-to-speech \\\n    --model_name \"ggml-org\u002Fgemma-4-E4B-it-GGUF\" \\\n    --responses_api_base_url \"http:\u002F\u002F127.0.0.1:8080\u002Fv1\" \\\n    --responses_api_api_key \"\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Any OpenAI Realtime-compatible client can connect. See \u003Ca href=\"#realtime-api\" rel=\"nofollow ugc noopener\">Realtime API\u003C\u002Fa> for the protocol and \u003Ca href=\"#llm-backends\" rel=\"nofollow ugc noopener\">LLM backends\u003C\u002Fa> for provider and local-server options.\u003C\u002Fp>\n\u003Ch2>Index\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"#how-it-works\" rel=\"nofollow ugc noopener\">How it works\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#installation\" rel=\"nofollow ugc noopener\">Installation\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#supported-components\" rel=\"nofollow ugc noopener\">Supported components\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#run-modes\" rel=\"nofollow ugc noopener\">Run modes\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#realtime-api\" rel=\"nofollow ugc noopener\">Realtime API\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#llm-backends\" rel=\"nofollow ugc noopener\">LLM backends\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#multi-language-support\" rel=\"nofollow ugc noopener\">Multi-language support\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#pocket-tts\" rel=\"nofollow ugc noopener\">Pocket TTS\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#cli-reference\" rel=\"nofollow ugc noopener\">CLI reference\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#contributing\" rel=\"nofollow ugc noopener\">Contributing\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#star-history\" rel=\"nofollow ugc noopener\">Star history\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#citations\" rel=\"nofollow ugc noopener\">Citations\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>How it works\u003C\u002Fh2>\n\u003Cp>The pipeline is a cascade of four components, each running in its own thread and connected by queues:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Voice Activity Detection (VAD)\u003C\u002Fstrong>: \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fsnakers4\u002Fsilero-vad\" rel=\"nofollow ugc noopener\">Silero VAD v5\u003C\u002Fa> detects speech boundaries and turn-taking.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Speech to Text (STT)\u003C\u002Fstrong>: transcribes the user's turn, with optional live partial transcripts.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Language Model (LLM)\u003C\u002Fstrong>: generates the response, streaming text and tool calls.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Text to Speech (TTS)\u003C\u002Fstrong>: synthesizes audio and streams it back to the client.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>Every stage has multiple interchangeable backends, selected via CLI flags. The code is designed for easy modification, with a focus on models available through Transformers and the Hugging Face Hub.\u003C\u002Fp>\n\u003Ch2>Installation\u003C\u002Fh2>\n\u003Cp>Requires Python 3.10+.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">pip install speech-to-speech\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The default install covers the standard realtime path:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>Parak\u003C\u002Fli>\n\u003C\u002Ful>\n",1785888997320]