StackMap
Subscribe
Explore / AnyJev
nokia-applied-research

AnyJev

Nokia research: turn any open LLM into a Jev-style decision model. Typed choice, yes/no or score from one prefill, de-biased with no labels or a fitted head, served on vLLM.

998 130 Python Apache-2.0updated 3 days ago
View on GitHubDispute this mapping →
Curator's take

The cleanest way to get decisions out of a model you already run: ask a typed question, read the next-token distribution from one prefill, nothing generated or parsed. Level L0 removes option-position bias with zero labels; L2 fits a tiny closed-form head from 100 to 300 labels and serves it from a vLLM pooling endpoint, and `anyjev.pipeline` measures accuracy, calibration and latency on your own box instead of asking you to trust their tables. One finding worth stealing: cutting Qwen2.5-7B to 18 of 28 blocks was faster with slightly better accuracy. It is a week-old research repo from two authors, so expect API churn, and it assumes you can run an open model; without a GPU, laya's small checkpoints are the lighter path.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside AnyJev. Ranked by curator confidence.

alternativealternativealternativebuilt withkevlayalocaljevvLLMAnyJev
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md2 min read

Jiamu Zhang1     Tianze Yang1     Yucheng Shi2     Liang Wu1

1 Nokia, Sunnyvale, CA      2 Tencent Hunyuan

Reverse the option order: the raw logit readout flips its answer, AnyJev L0 gives the same answer both ways
Qwen3-8B on a real BANKING77 item. Every number is a model output.

[!TIP] 🆕 vLLM serves every level, L2 included. An embed server's pooler hands back the hidden state a closed-form head reads, so a decision endpoint is a pooling server plus a few kilobytes of head. python -m anyjev.pipeline <model> converts, serves and measures in one command. Start here ↓

⚡ Serve it

Three commands take a model off the Hub and put a calibrated decision endpoint in front of it.

pip install "anyjev[hf]"

# 1. keep the blocks a decision needs — usually about two thirds
python -m anyjev.truncate Qwen/Qwen2.5-7B-Instruct 18 ./qwen-b18

# 2. serve it. L2 reads a hidden state, so the pooler hands one back untouched
vllm serve ./qwen-b18 --task embed \
  --override-pooler-config '{"pooling_type":"LAST","normalize":false,"softmax":false}'
from anyjev import Decider, Question
from anyjev.backends.vllm import VLLMBackend

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), level="L2")
route = Question.choice("Which team should handle this?",
                        ["billing", "technical", "sales", "other"], name="route")

d.fit_head(route, states, labels, layers=[-1])   # 100–300 labels, one closed-form solve
d.decide(ticket, [route])["route"].distribution  # {"billing": 0.81, "technical": 0.07, ...}

Without labels, turn on the rotation budget — recommended for any K-option choice. L0 asks the model once per option rotation so that no option is favoured by its position. Most decisions do not need all K: read them one at a time, stop when the leader is far enough ahead, and the threshold can be calibrated so the answer matches the full cycle's a stated fraction of the time — measured against our own full-strength readout, so it needs no labels at all.

d = Decider(VLLMBackend("http://localhost:8000", "./qwen-b18"), adaptive_shifts=True)
d.calibrate_adaptive(route, unlabelled_tickets, target=0.01)   # a few hundred states, no labels
d.decide_batch(tickets, route)      # diagnostics: shifts_used, stop_threshold

7.2 rotations instead of 18 at a certified 1% disagreement rate, **2.2× the decisions per s