[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:clm":3},"\n\u003Cp align=\"center\">\n  \u003Cpicture>\n    \u003Cimg alt=\"CLM v0.1\" src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FContrastive-LM\u002Fclm\u002FHEAD\u002Fassets\u002Flogo.png\" width=\"45%\" \u002F>\n  \u003C\u002Fpicture>\n\u003C\u002Fp>\u003Ch3>\nContrastive Language Models\n\u003C\u002Fh3>\u003Cp align=\"center\">\n\u003Ci>A System One Model for Fast and Generalizable Decision-Making\u003C\u002Fi>\n\u003C\u002Fp>\u003Cp align=\"center\">\n| 📄 \u003Ca href=\"https:\u002F\u002Fcontrastive-lm.notion.site\" rel=\"nofollow ugc noopener\">\u003Cb>Blog\u003C\u002Fb>\u003C\u002Fa> | 🗣️ \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002F5dAQEDJBs\" rel=\"nofollow ugc noopener\">\u003Cb>Discord\u003C\u002Fb>\u003C\u002Fa> | 🤗 \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FContrastive-LM\" rel=\"nofollow ugc noopener\">\u003Cb>Data &amp; Models\u003C\u002Fb>\u003C\u002Fa> | 📚 \u003Ca href=\"#api-reference\" rel=\"nofollow ugc noopener\">\u003Cb>API Reference\u003C\u002Fb>\u003C\u002Fa> | 🛠️ \u003Ca href=\"#fine-tuning-clm-on-your-own-data\" rel=\"nofollow ugc noopener\">\u003Cb>Fine-Tuning Tutorial\u003C\u002Fb>\u003C\u002Fa> |\n\u003C\u002Fp>\u003Cp>🔥 \u003Cstrong>Contrastive Language Models (CLMs)\u003C\u002Fstrong> are a new class of \u003Cstrong>System One\nmodel\u003C\u002Fstrong> trained with a \u003Cstrong>contrastive learning\u003C\u002Fstrong> objective that connects\n\u003Cstrong>states and actions\u003C\u002Fstrong>. This repo serves \u003Cstrong>CLM-8B\u003C\u002Fstrong> behind a\nTypeSafe-compatible API.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>CLM-8B\u003C\u002Fstrong> is pre-trained on \u003Cstrong>60M Nemotron Q&amp;A pairs\u003C\u002Fstrong>, mid-trained on\n\u003Cstrong>30M synthetic hard negatives\u003C\u002Fstrong>, and post-trained on \u003Cstrong>1M agentic\ntrajectories\u003C\u002Fstrong>.\u003C\u002Fli>\n\u003Cli>It performs on par with \u003Cstrong>Jev\u003C\u002Fstrong> across computer-use, gaming and tool-calling\ntasks with up to \u003Cstrong>9× lower latency\u003C\u002Fstrong>. With lightweight fine-tuning it sets a\nnew SOTA as a verifier on agentic coding benchmarks: \u003Cstrong>Terminal-Bench 2.1\n(87.6%)\u003C\u002Fstrong> and \u003Cstrong>DeepSWE (81.6%)\u003C\u002Fstrong>.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>States and actions are disaggregated\u003C\u002Fstrong>, so their embeddings are cached and\nreused independently, which makes training and serving cheap and blazing fast!\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>We invite the community to plug it into their own agents and benchmarks!\u003C\u002Fp>\n\u003Chr \u002F>\n\u003Ch2>Installation\u003C\u002Fh2>\n\u003Cpre>\u003Ccode class=\"language-bash\">pip install contrastive-lm\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>To install the latest from a clone:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">pip install -e .\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Chr \u002F>\n\u003Ch2>Quickstart\u003C\u002Fh2>\n\u003Ch3>Serve\u003C\u002Fh3>\n\u003Cpre>\u003Ccode class=\"language-bash\"># 1. encoder (Qwen3-8B embeddings)\nvllm serve Qwen\u002FQwen3-8B --served-model-name qwen3-8b --runner pooling --max-model-len 2048 --port 8090 &amp;\n\n# 2. CLM API on :8700 (downloads the 75 MB reference head on first run)\nclm-serve\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>States longer than 2048 tokens are truncated. For longer states, raise both limits\ntogether, e.g. \u003Ccode>--max-model-len 8192\u003C\u002Fcode> on \u003Ccode>vllm serve\u003C\u002Fcode> and \u003Ccode>clm-serve --max-tokens 8192\u003C\u002Fcode>\n(needs more GPU memory).\u003C\u002Fp>\n\u003Ch3>Ask typed questions about a state\u003C\u002Fh3>\n\u003Cpre>\u003Ccode class=\"language-python\">from clm import CLMClient, Choice, Noul, Score\n\nclient = CLMClient()                          # CLM_BASE_URL (default http:\u002F\u002F127.0.0.1:8700), CLM_API_KEY\nr = client.system_one(\n    state=\"Customer: my invoice was charged twice and nobody answers the phone!\",\n    questions={\n        \"urgency\": Noul(instructions=\"Is this urgent?\"),\n        \"department\": Choice(instructions=\"Which team should handle this?\",\n                             criteria={\"billing\": \"Charges, invoices, refunds\",\n                                       \"technical\": \"Bugs and outages\"}),\n        \"frustration\": Score(instructions=\"How frustrated is the customer?\",\n                             criteria=[\"Calm\", \"Frustrated\", \"Very angry\"]),\n    },\n)\nprint(r.answers[\"urgency\"].noul)                # 0.41022     probability the statement is true\nprint(r.answers[\"department\"].choice)           # billing\nprint(r.answers[\"department\"].probabilities)    # {'billing': 0.93878, 'technical': 0.06122}\nprint(r.answers[\"frustration\"].score)           # 1.98386     expected level, 0..2\nprint(r.usage.input_tokens, r.latency_ms)       # 38 58.1     (106 tokens on a cold cache: option texts are embedded once)\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Questions may be \u003Ccode>Noul\u003C\u002Fcode> \u002F \u003Ccode>Choice\u003C\u002Fcode> \u002F \u003Ccode>Score\u003C\u002Fcode> objects or plain wire-format\ndicts, so a request written for TypeSafe replays as\n\u003Ccode>client.system_one(state, questions)\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Ch3>Rank candidates directly\u003C\u002Fh3>\n\u003Cp>\u003Ccode>system_one\u003C\u002Fcode> is built on one primitive: score a candidate against a state.\nFor free-form candidates (best-of-N answers, tool names, next moves) use the\nin-process engine's \u003Ccode>rank\u003C\u002Fcode>:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-python\">from clm import Engine\n\nengine = Engine(emb_url=\"http:\u002F\u002F127.0.0.1:8090\u002Fv1\u002Fembeddings\")     # reference head, downloaded if missing\nengine.rank(\"What causes tides on Earth?\",\n            [\"The Moon's gravitational\n\u003C\u002Fcode>\u003C\u002Fpre>\n",1791678831787]