Jiamu Zhang1 Tianze Yang1 Yucheng Shi2 Liang Wu1
1 Nokia, Sunnyvale, CA 2 Tencent Hunyuan
Qwen3-8B on a real BANKING77 item. Every number is a model output.
[!TIP] 🆕 vLLM serves every level, L2 included. An embed server's pooler hands back the hidden state a closed-form head reads, so a decision endpoint is a pooling server plus a few kilobytes of head.
python -m anyjev.pipeline <model>converts, serves and measures in one command. Start here ↓
⚡ Serve it
Three commands take a model off the Hub and put a calibrated decision endpoint in front of it.
pip install "anyjev[hf]"
# 1. keep the blocks a decision needs — usually about two thirds
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18
# 2. serve it. L2 reads a hidden state, so the pooler hands one back untouched
vllm serve ./qwen-b18 --task embed \
--override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'
from anyjev import Decider, Question
from anyjev.backends.vllm import VLLMBackend
d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), level="L2")
route = Question.choice("Which team should handle this?",
["billing", "technical", "sales", "other"], name="route")
d.fit_head(route, states, labels, layers=[-1]) # 100–300 labels, one closed-form solve
d.decide(ticket, [route])["route"].distribution # {"billing": 0.81, "technical": 0.07, ...}
Without labels, turn on the rotation budget — recommended for any K-option choice. L0 asks the
model once per option rotation so that no option is favoured by its position. Most decisions do not need
all K: read them one at a time, stop when the leader is far enough ahead, and the threshold can be
calibrated so the answer matches the full cycle's a stated fraction of the time — measured against
our own full-strength readout, so it needs no labels at all.
d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), adaptive_shifts=True)
d.calibrate_adaptive(route, unlabelled_tickets, target=0.01) # a few hundred states, no labels
d.decide_batch(tickets, route) # diagnostics: shifts_used, stop_threshold
7.2 rotations instead of 18 at a certified 1% disagreement rate, **2.2× the decisions per s
