[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:search-r1":3},"\u003Ch1>Search-R1: Train your LLMs to reason and call a search engine with reinforcement learning\u003C\u002Fh1>\n\u003Cdiv align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FPeterGriffinJin\u002FSearch-R1\u002Fmain\u002Fpublic\u002Flogo.png\" alt=\"logo\" width=\"300\" \u002F>\n\u003C\u002Fdiv>\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2503.09516\" rel=\"nofollow ugc noopener\">\n    \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPaper1-blue?style=for-the-badge\" alt=\"Button1\" \u002F>\n  \u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2505.15117\" rel=\"nofollow ugc noopener\">\n    \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPaper2-green?style=for-the-badge\" alt=\"Button2\" \u002F>\n  \u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fcollections\u002FPeterJinGo\u002Fsearch-r1-67d1a021202731cb065740f5\" rel=\"nofollow ugc noopener\">\n    \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FResources-orange?style=for-the-badge\" alt=\"Button3\" \u002F>\n  \u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fx.com\u002FBowenJin13\u002Fstatus\u002F1895544294473109889\" rel=\"nofollow ugc noopener\">\n    \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FTweet-red?style=for-the-badge\" alt=\"Button4\" \u002F>\n  \u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fwandb.ai\u002Fpeterjin\u002FSearch-R1-v0.2\" rel=\"nofollow ugc noopener\">\n    \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLogs-purple?style=for-the-badge\" alt=\"Button5\" \u002F>\n  \u003C\u002Fa>\n\u003C\u002Fp>\n\n\u003Cp>\u003Cstrong>Search-R1\u003C\u002Fstrong> is a reinforcement learning framework designed for training \u003Cstrong>reasoning-and-searching interleaved LLMs\u003C\u002Fstrong>—language models that learn to reason and make tool calls (e.g., to search engines) in a coordinated manner.\u003C\u002Fp>\n\n\u003Cp>Built upon \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fvolcengine\u002Fverl\" rel=\"nofollow ugc noopener\">veRL\u003C\u002Fa>, Search-R1 extends the ideas of \u003Cstrong>DeepSeek-R1(-Zero)\u003C\u002Fstrong> by incorporating interleaved search engine access and provides a fully open-source RL training pipeline. It serves as an alternative and open solution to \u003Cstrong>OpenAI DeepResearch\u003C\u002Fstrong>, enabling research and development in tool-augmented LLM reasoning.\u003C\u002Fp>\n\u003Cp>We support different RL methods (e.g., PPO, GRPO, reinforce), different LLMs (e.g., llama3, Qwen2.5, etc) and different search engines (e.g., local sparse\u002Fdense retrievers and online search engines).\u003C\u002Fp>\n\u003Cp>Paper: \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fpdf\u002F2503.09516\" rel=\"nofollow ugc noopener\">link1\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2505.15117\" rel=\"nofollow ugc noopener\">link2\u003C\u002Fa>; Model and data: \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002Fcollections\u002FPeterJinGo\u002Fsearch-r1-67d1a021202731cb065740f5\" rel=\"nofollow ugc noopener\">link\u003C\u002Fa>; Twitter thread: \u003Ca href=\"https:\u002F\u002Fx.com\u002FBowenJin13\u002Fstatus\u002F1895544294473109889\" rel=\"nofollow ugc noopener\">link\u003C\u002Fa>; Full experiment log: \u003Ca href=\"https:\u002F\u002Fwandb.ai\u002Fpeterjin\u002FSearch-R1-open\" rel=\"nofollow ugc noopener\">prelim\u003C\u002Fa>; \u003Ca href=\"https:\u002F\u002Fwandb.ai\u002Fpeterjin\u002FSearch-R1-nq_hotpotqa_train\" rel=\"nofollow ugc noopener\">v0.1\u003C\u002Fa>; \u003Ca href=\"https:\u002F\u002Fwandb.ai\u002Fpeterjin\u002FSearch-R1-v0.2\" rel=\"nofollow ugc noopener\">v0.2\u003C\u002Fa>; \u003Ca href=\"https:\u002F\u002Fwandb.ai\u002Fpeterjin\u002FSearch-R1-v0.3\" rel=\"nofollow ugc noopener\">v0.3\u003C\u002Fa>. Details about these logs and methods can be find \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FPeterGriffinJin\u002FSearch-R1\u002Fblob\u002Fmain\u002Fdocs\u002Fexperiment_log.md\" rel=\"nofollow ugc noopener\">here\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FPeterGriffinJin\u002Fsearch-r1\u002FHEAD\u002Fpublic\u002Fmain.png\" alt=\"single-turn\" \u002F>\u003C\u002Fp>\n\u003Ch2>News\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>[2025.10] Search-R1 is featured by Thinking Machines Lab's first product \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fthinking-machines-lab\u002Ftinker-cookbook\" rel=\"nofollow ugc noopener\">Tinker\u003C\u002Fa>! Details: \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fthinking-machines-lab\u002Ftinker-cookbook\u002Ftree\u002Fmain\u002Ftinker_cookbook\u002Frecipes\u002Ftool_use\u002Fsearch\" rel=\"nofollow ugc noopener\">Document\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>[2025.7] Search-R1 is supported by \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNovaSky-AI\u002FSkyRL\" rel=\"nofollow ugc noopener\">SkyRL\u003C\u002Fa>! Detailed instructions: \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FNovaSky-AI\u002FSkyRL\u002Ftree\u002Fmain\u002Fskyrl-train\u002Fexamples\u002Fsearch\" rel=\"nofollow ugc noopener\">code\u003C\u002Fa>, \u003Ca href=\"https:\u002F\u002Fnovasky-ai.notion.site\u002Fskyrl-searchr1\" rel=\"nofollow ugc noopener\">Document\u003C\u002Fa>.\u003C\u002Fli>\n\u003Cli>[2025.6] Search-R1 is now integrated into the latest version of veRL and can take advantage of its most up-to-date features! Detailed instructions: \u003Ca href=\"https:\u002F\u002Fverl.readthedocs.io\u002Fen\u002Flatest\u002Fsglang_multiturn\u002Fsearch_tool_example.html\" rel=\"nofollow ugc noopener\">veRL\u003C\u002Fa>, [English Document](\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fzhaochenyang20\u002FAwesome-ML-SYS-Tutorial\u002Fblob\u002Fmai\" rel=\"nofollow ugc noopener\">https:\u002F\u002Fgithub.com\u002Fzhaochenyang20\u002FAwesome-ML-SYS-Tutorial\u002Fblob\u002Fmai\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n",1786315278202]