[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:turbovec":3},"\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FRyanCodrai\u002Fturbovec\u002Fmain\u002Fdocs\u002Fheader.png\" alt=\"turbovec — Google's TurboQuant for vector search\" width=\"100%\" \u002F>\n\u003C\u002Fp>\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FRyanCodrai\u002Fturbovec\u002Fblob\u002Fmain\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fl\u002Fturbovec\" alt=\"License\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fturbovec\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fturbovec?label=pypi&amp;color=blue\" alt=\"PyPI version\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fcrates.io\u002Fcrates\u002Fturbovec\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fcrates\u002Fv\u002Fturbovec?label=crates.io&amp;color=blue\" alt=\"crates.io version\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.19874\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpaper-arXiv-b31b1b.svg\" alt=\"TurboQuant paper\" \u002F>\u003C\u002Fa>\n\u003C\u002Fp>\u003Chr \u002F>\n\u003Cp>\u003Cstrong>A 10 million document corpus takes 31 GB of RAM as float32. turbovec fits it in 4 GB - and searches it faster than FAISS.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>turbovec is a Rust vector index with Python bindings, built on Google Research's \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2504.19874\" rel=\"nofollow ugc noopener\">\u003Cstrong>TurboQuant\u003C\u002Fstrong>\u003C\u002Fa> algorithm — a data-oblivious quantizer with near-optimal distortion and no separate training phase.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Online ingest.\u003C\u002Fstrong> Add vectors, they're indexed — no train step, no parameter tuning, no rebuilds as the corpus grows.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fast SIMD search.\u003C\u002Fstrong> Hand-written kernels — NEON SDOT\u002FSMMLA on ARM, AVX-512 VNNI and \u003Ccode>vpermb\u003C\u002Fcode> on x86, with AVX2 and scalar fallbacks — beat FAISS IndexPQFastScan in every measured config, averaging 3.4× at 4-bit and 23% at 2-bit across the eight cells of each width, on both architectures.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Incremental saves.\u003C\u002Fstrong> \u003Ccode>sync(path)\u003C\u002Fcode> persists just what changed since the last sync — one fsync per call, crash-safe at any byte, and a removal or a small append costs milliseconds however large the index. \u003Ccode>write\u003C\u002Fcode>\u002F\u003Ccode>load\u003C\u002Fcode> stay for whole-file snapshots.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Filter at search time.\u003C\u002Fstrong> Pass an id allowlist (or a slot bitmask) to \u003Ccode>search()\u003C\u002Fcode> and the kernel honours it directly. You always get up to \u003Ccode>k\u003C\u002Fcode> results from the allowed set — no over-fetching, no recall hit on selective filters.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Pure local.\u003C\u002Fstrong> No managed service, no data leaving your machine or VPC. Pair with any open-source embedding model for a fully air-gapped RAG stack.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Building RAG where privacy, memory, or latency matters? \u003Cstrong>You're in the right place.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Ch2>Python\u003C\u002Fh2>\n\u003Cpre>\u003Ccode class=\"language-bash\">pip install turbovec\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cpre>\u003Ccode class=\"language-python\">from turbovec import TurboQuantIndex\n\nindex = TurboQuantIndex(dim=1536, bit_width=4)\nindex.add(vectors)\nindex.add(more_vectors)\n\nscores, indices = index.search(query, k=10)\n\nindex.write(\"my_index.tv\")\nloaded = TurboQuantIndex.load(\"my_index.tv\")\n\nindex.sync(\"my_index.tv\")   # after more changes: durable incremental save\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Ccode>vectors\u003C\u002Fcode> and \u003Ccode>query\u003C\u002Fcode> are 2-D \u003Ccode>float32\u003C\u002Fcode> arrays of shape \u003Ccode>(n, dim)\u003C\u002Fcode> — other dtypes are rejected rather than silently converted, so cast with \u003Ccode>np.asarray(x, dtype=np.float32)\u003C\u002Fcode> first if needed.\u003C\u002Fp>\n\u003Cp>Need stable ids that survive deletes? Use \u003Ccode>IdMapIndex\u003C\u002Fcode>:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-python\">import numpy as np\nfrom turbovec import IdMapIndex\n\nindex = IdMapIndex(dim=1536, bit_width=4)\nindex.add_with_ids(vectors, np.array([1001, 1002, 1003], dtype=np.uint64))\n\nscores, ids = index.search(query, k=10)   # ids are your uint64 external ids\nindex.remove(1002)                         # O(1) by id\n\nindex.write(\"my_index.tvim\")\nloaded = IdMapIndex.load(\"my_index.tvim\")\n\nindex.sync(\"my_index.tvim\")   # durable incremental save, ids included\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch3>Hybrid retrieval (filtered search)\u003C\u002Fh3>\n\u003Cp>Restrict results to a candidate set produced by another system (SQL, BM25, ACL, time window, …):\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-python\">import numpy as np\nfrom turbovec import IdMapIndex\n\nidx = IdMapIndex(dim=1536, bit_width=4)\nidx.add_with_ids(vectors, ids)\n\n# Stage 1: external system narrows to candidate ids.\nallowed = np.array(db.execute(\"SELECT id FROM docs WHERE tenant=?\", (t,)).fetchall(),\n                   dtype=np.uint64)\n\n# Stage 2: dense rerank within the candidate set.\nscores, ids = idx.search(query, k=10, allowlist=allowed)\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Filtering happens inside the SIMD kernel at 32-vector block granula\u003C\u002Fp>\n",1788219551364]