[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:needle":3},"\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fcactus-compute\u002Fneedle\u002FHEAD\u002Fassets\u002Fbanner.png\" alt=\"Needle\" \u002F>\u003C\u002Fp>\n\u003Ch1>Needle 2\u003C\u002Fh1>\n\u003Cp>Needle 2 is an open 45M-parameter model for tool calling, device use and structured extraction. The whole model is a single 14MB binary that runs a full session in about 28MB of RAM. It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants, and baked into its own engine. On the benchmarks below, Needle 2 trades wins with other small models like FunctionGemma 270M, LFM2.5 230M and Apple FM, at 5x to 70x smaller, and 2 bits against their f16.\u003C\u002Fp>\n\u003Cp>This repository is the Python package: inference, LoRA fine-tuning, and export. \u003Ccode>pip install cactus-needle\u003C\u002Fcode>, describe your tools, and call them from Python. The inference engine is fetched once from Hugging Face and cached; there is nothing else to build, and offline setup for air gapped devices is covered in \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fcactus-compute\u002Fneedle\u002Fblob\u002FHEAD\u002Fdoc\u002Fapis.md\" rel=\"nofollow ugc noopener\">doc\u002Fapis.md\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Self-contained\u003C\u002Fstrong>: weights baked into a single 14MB engine; no separate model files to manage, and inference does no network.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Simple contract\u003C\u002Fstrong>: tool calls come back as structured data, text in, JSON out; a byte-level grammar compiled from your schemas constrains every token.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Confidence-gated\u003C\u002Fstrong>: every response carries a calibrated confidence score from a learned head; set a threshold, act above it, escalate below it.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Tool retrieval\u003C\u002Fstrong>: declare a large catalogue and a built-in retrieval head renders only the top five tools per turn, with the grammar constrained to that subset.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Bounded memory\u003C\u002Fstrong>: a 256-token sliding window with the tools pinned as KV sinks, so total memory stays near 28MB no matter how long the conversation runs.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Weights: \u003Ca href=\"https:\u002F\u002Fhuggingface.co\u002FCactus-Compute\u002Fneedle2\" rel=\"nofollow ugc noopener\">huggingface.co\u002FCactus-Compute\u002Fneedle2\u003C\u002Fa> · source: \u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fcactus-compute\u002Fneedle\" rel=\"nofollow ugc noopener\">github.com\u002Fcactus-compute\u002Fneedle\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fcactus-compute\u002Fneedle\u002FHEAD\u002Fassets\u002Ffrontier.png\" alt=\"Size-quality frontier: mobile-class and below\" \u002F>\u003C\u002Fp>\n\u003Ch2>Simple Attention Network\u003C\u002Fh2>\n\u003Cp>Needle 2 is a Simple Attention Network, our dense small-model recipe: a Hadamard MLP in place of the FFN, GQA attention, engram key-value memory, and multi-lane hyper-connections. See the paper for the design and ablations: \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.18363\" rel=\"nofollow ugc noopener\">arXiv:2607.18363\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fcactus-compute\u002Fneedle\u002FHEAD\u002Fassets\u002Farchitecture.png\" alt=\"Simple Attention Network architecture\" \u002F>\u003C\u002Fp>\n\u003Cp>Each block carries its update rule. Here x̂ is the RMS-normalised flattening of the four residual streams, H the orthonormal Walsh-Hadamard transform (a fixed matrix, applied in n log n time with no weights to read), (kₜ, vₜ) rows gathered from hashed n-gram tables, and P the doubly-stochastic normalisation of the routing logits A, computed by Sinkhorn iteration; a, b, g and all σ-gates are learned and input-dependent. Both attention and MLP residuals are sandwich-normed and gated, the engram sites fire at two layers, and decoding is constrained by a byte-level grammar compiled from the declared schemas.\u003C\u002Fp>\n\u003Ch2>Quickstart\u003C\u002Fh2>\n\u003Cpre>\u003Ccode class=\"language-sh\">pip install cactus-needle\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Needle reads your tool descriptions to decide what to call and how to fill arguments, so describing them well is the whole game.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Simple\u003C\u002Fstrong>: decorate a function. The signature gives the argument types, the docstring is the tool description, and \u003Ccode>run()\u003C\u002Fcode> completes the loop: model picks the call, Needle executes your function, feeds the result back, and returns the final response with the executed tool results attached as \u003Ccode>results\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-python\">import needle\n\n@needle.tool\ndef get_weather(city: str):\n    \"Get the current weather for a city.\"\n    return {\"city\": city, \"temp_c\": 27, \"sky\": \"clear\"}\n\nagent = needle.Needle(tools=[get_weather])\nprint(agent.run(\"what's it like in Lagos right now?\")[\"results\"])\n# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cstrong>Extraction\u003C\u002Fstrong>: to pull structured data out of text, declare the shape and call \u003Ccode>extract()\u003C\u002Fcode>. Pass a Pydantic model and you get a typed object back.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-python\">from pydantic import BaseModel\n\nclass Invoice(BaseModel):\n    vendor: str\n    total: float\n    due_date: str\n\ninvoice = needle.extract(\"Invoice from \n\u003C\u002Fcode>\u003C\u002Fpre>\n",1787581367586]