Sieve-9B

Sieve-9B is a decision model from the Sieve family. You give it a piece of text or JSON (the state) and typed questions, and it returns a calibrated probability for every option of every question. It reads the state once, scores the options directly, and never generates text.

It is a LoRA adapter (rank 32) and a 2.1M-parameter pointer head on Qwen/Qwen3.5-9B. The code and training data are at github.com/sthanika-ai/sieve.

How it works

Sieve follows the decision-model design of Kev:

  • Backbone. Only the text part of Qwen3.5-9B is kept: 32 layers and the embeddings, 7.94B parameters.
    • Removed: the language-model head, the multi-token-prediction layer and the vision encoder, 1.72B parameters together.
  • Readout. A pointer head scores the end of each option against the decision point.
  • Isolation. Each question runs as its own row from the state's cache, so questions never see each other.

What Sieve-9B changes:

  • Backbone. The post-trained Qwen3.5-9B instead of the base model.
  • Averaged weights. The adapter and head are the exact average of two rank-16 runs from one initialisation.
  • New data. 1,920 records from 240 rule structures not in the decision-v7 split, plus states padded with unrelated text.
  • Serving:
    • CUDA graphs: 2 options take 18.6 ms instead of 70.2 ms.
    • Long questions run as one row: 255 options take 330.5 ms instead of 371.0 ms.
type options returns
choice the caller's keys, optionally described (up to 255) a probability per key
noul yes / no P(yes)
score ordered levels a probability per level

Use

pip install "sieve-decisions[cuda,serve] @ git+https://github.com/sthanika-ai/sieve"
from sieve import load_sieve, decide

m = load_sieve("sthanika-ai/Sieve-9B")
decide(m, {"subject": "Duplicate charge on invoice 4411", "body": "Billed twice. Refund today or we cancel."},
       {"department": {"type": "choice", "instructions": "Which team handles this?",
                       "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages"}},
        "churn_risk": {"type": "noul", "instructions": "Does the customer threaten to cancel?"}})
  • Loading. Load it with sieve; the adapter alone is not a text-generation model.
  • Server. sieve-serve --model sthanika-ai/Sieve-9B --graphs.
  • GPU memory. About 20 GB; 17.6 GiB was measured with a 4,000-token state.
  • Temperature. head.pt stores T = 1.631, fitted by NLL on the decision-v7 calibration split (1,148 questions).
    • That split is used only to fit T and to choose the checkpoint; no reported number comes from it.
    • T never changes an answer, and temperature=1.0 gives raw probabilities.

Results

All numbers are Sieve-9B's own. They are recomputed from its saved per-question predictions and stored in result.json.

sieve-eval in the GitHub repository computes the same metrics for any labelled JSONL file. With --no-merge it reproduces the public-test rows below exactly.

Decision Index 0.2.1

Sieve-9B scores 41.71 (raw index 56.47, breadth 40.06) on our own run of the full frozen Decision Index suite. The run used the official kit at 87d4650 and its scorer.

  • Coverage. 150,317 requests over 44 benchmarks, all answered, with no truncation or errors. The index averages 38 of the benchmarks.
  • Results. sthanika-ai/Sieve-9B-decision-index-results.
  • Not on the leaderboard yet. The score is self-reported and expected, not confirmed, until the maintainers review it. On the 2026-09-28 board it would place about 18th of 72.
Decision Index Knowledge & Reasoning Language Understanding Retrieval & Classification Tools & Automation Arts & Human Taste
41.71 27.4 46.1 46.1 61.3 22.8

The index and area scores are chance-corrected: 0 is random guessing and 100 is perfect. The per-benchmark scores below are in each benchmark's own metric.

All 44 benchmarks
benchmark metric score requests
ACOS per-review F1 0.241 1,565
ANLI macro-F1 0.562 3,200
API-Bank accuracy 0.726 508
Amazon ESCI macro-F1 0.493 5,000
BANKING77 macro-F1 0.836 3,080
BBH fixed-option tasks accuracy 0.672 5,507
BFCL case exact accuracy 0.954 1,694
BPoMP accuracy 0.665 5,000
BRIGHT nDCG@10 0.420 220
CLINC150+OOS macro-F1 0.760 5,500
CLadder accuracy 0.633 5,000
CRUXEval accuracy 0.537 570
ChessBench accuracy 0.114 5,000
ContractNLI macro-F1 0.615 123
FinEntity macro-F1 0.888 979
ForecastBench Brier (lower is better) 0.181 10,139
GPQA Diamond accuracy 0.424 196
GSM8K accuracy 0.460 2,638
HLE accuracy 0.106 501
Habermas Machine accuracy 0.404 1,676
HellaSwag accuracy 0.830 10,042
HoVer claim verification accuracy 0.635 4,000
Home appliance simulator case exact accuracy 0.307 88
Humicroedit accuracy 0.551 2,628
MMLU-Pro accuracy 0.508 12,032
MuSR accuracy 0.567 752
NLI4CT macro-F1 0.779 5,500
New Yorker caption matching accuracy 0.581 528
POP909-CL accuracy 0.129 2,000
PhishNChips phishing decisions accuracy 0.541 2,000
RAGTruth response-level hallucination F1 on hallucinated class 0.629 2,700
SATA-Bench case exact accuracy 0.296 1,650
ToolRet nDCG@10 0.651 685
VAST macro-F1 0.522 3,006
When2Call MCQ accuracy 0.561 3,652
WinoGrande accuracy 0.751 1,267
cfcolor accuracy 0.580 5,000
iSarcasmEval Sarcasm F1 · track A, English 0.506 4,600
ARC-Challenge (not in the index) accuracy 0.938 1,172
ARC-Easy (not in the index) accuracy 0.978 2,376
MMLU (not in the index) accuracy 0.749 14,033
RouterBench (not in the index) selected quality (quality objective) 0.793 10,000
SGD/SGD-X (not in the index) macro-F1 0.627 2,500
SimpleBench (not in the index) accuracy 0.200 10

Decision suites, held-out benchmark and public test sets

benchmark questions accuracy [95% CI] NLL Brier ECE wrong at p ≥ 0.9
decision-v7 development (in-domain) 1,264 0.872 [0.851, 0.892] 0.330 0.176 0.028 1.7%
transfer-v4 development (out-of-domain) 656 0.829 [0.797, 0.860] 0.447 0.242 0.045 3.7%
transfer-v9 development, strict (out-of-domain) 936 0.753 [0.724, 0.781] 0.689 0.337 0.050 3.8%
held-out decision benchmark ³ 6,224 0.857 [0.848, 0.865] 0.451 0.207 0.055 0.4%
MMLU, 57 subjects, 4-way ² 14,042 0.758 [0.751, 0.765] 0.642 0.333 0.008 1.7%
SciQ 1,000 0.984 [0.976, 0.991] 0.063 0.026 0.024 0.1%
Banking77, 77 intents ¹ 3,079 0.857 [0.844, 0.869] 0.563 0.224 0.043 1.6%
MASSIVE intent, 60 labels 2,974 0.765 [0.749, 0.780] 0.886 0.345 0.093 0.5%
MASSIVE scenario, 18 labels 2,974 0.711 [0.695, 0.728] 0.889 0.401 0.059 0.2%
typed-decisions test 2,000 0.704 [0.681, 0.726] 0.704 0.404 0.040 0.5%
  • Calibration. The probability metrics use the stored temperature (T = 1.631), as served. Raw (T = 1) values are in result.json.
  • Wrong at p ≥ 0.9 is the share of all questions answered wrongly with a probability of at least 0.9.
  • Intervals. They come from a bootstrap with 4,000 resamples over groups of related records.
  • Development partitions only. The decision suites are development partitions; no test partition of them has been scored.
  • Unmerged adapter. These rows were scored with the adapter unmerged. The Decision Index run and the latency use the merged model, as served.

¹ In-distribution: the training data contains 1,000 Banking77 training questions.

² Our conversion. On the Decision Index's own MMLU, 14,033 questions outside the index, Sieve-9B scores 0.749.

³ The held-out benchmark:

  • Size. 6,233 questions over 10 synthetic decision domains, with no family overlap with training. 6,224 were scored; 9 were dropped because their options exceed the evaluation's 12,288-token question limit.
  • Truncation. States over 384 tokens, the training limit, were truncated to 384.
  • Option counts. 77 of the scored questions have 256 or 512 options, more than the served API accepts (255).
held-out benchmark by questions accuracy
state ≤ 384 tokens 5,186 0.889
state > 384 tokens (truncated) 1,038 0.697
2 options 2,452 0.976
3–10 options 2,969 0.823
11–50 options 522 0.728
51–255 options 204 0.431
256–512 options 77 0.377

Tool selection (BFCL v3, decision form)

The model picks the function to call from the offered ones, or "none of these". Argument values are not generated or scored, so these are not BFCL leaderboard numbers.

BFCL v3, decision form questions accuracy
function selection (simple, multiple, live simple, live multiple) 1,909 0.963
irrelevance detection (irrelevance, live irrelevance) 1,118 0.720

Latency

Server-reported p50 on one A100 80GB (bf16, merged, CUDA graphs), for one choice question on a cached state of about 75 tokens, excluding tokenisation and HTTP:

options 2 5 10 25 50 100 255
ms 18.6 19.3 22.8 42.8 74.5 128.4 330.5

Training

  • Recipe:
    • LoRA r=16, α=32 on the attention, MLP and DeltaNet projections, with the pointer head trained from scratch;
    • lr 5e-5, OneCycle, 2 epochs, 8 records per step, bf16;
    • permuted choice options and "none of the above" augmentation, with a cross-entropy loss.
  • Averaging. Two adapters were trained from one initialisation on different data mixes and averaged 50/50 ("model soups", Wortsman et al., 2022).
  • Selection. The checkpoint was chosen by calibration-split accuracy, then NLL. That rule was fixed after both runs were evaluated and before the average was scored.
    • The benchmarks were never used to choose. The second data mix did target weaknesses seen on the development partitions.
    • No outputs of other models were used.
  • Data. The four training files are in data/sieve-9b/, unchanged.
  • Data hygiene. The training files were checked against seven benchmark partitions (listed in provenance.json).
    • No overlap: by exact question, exact state, and an order-insensitive hash of each state.
    • Near-duplicates: the unknowable date-policy file has 4 near-duplicate states, and its families are excluded from every transfer-v9 number.
    • Rule structures: 30 of the 240 are logically equivalent to a decision-v7 structure after pushing negations inward. None equals or contains a held-out structure.

Files

file contents
adapter_model.safetensors, adapter_config.json the LoRA adapter (rank 32, PEFT format)
head.pt the pointer head and the calibrated temperature
tokenizer.json, tokenizer_config.json Qwen3.5-9B's tokenizer, unchanged
training_config.json, training_metrics.json, train.log the recipe, the two runs and how they were averaged
result.json the numbers above, from saved per-question predictions, with the same metrics as sieve-eval
provenance.json SHA-256 of every file and of the training data, and library versions

Credits and license

  • Design. The typed-decision format, the pointer-head design and the stored temperature follow Kev (Apache-2.0).
  • Training data. It comes from Kev's decision-v7 suite, its date-policy files and its rule generator. It also contains text from ten public datasets under their own licences, listed in the GitHub NOTICE.
  • Backbone. Qwen3.5-9B, Apache-2.0.
  • License. The weights are Apache-2.0 (see LICENSE and NOTICE).
Downloads last month
54
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sthanika-ai/Sieve-9B

Finetuned
Qwen/Qwen3.5-9B
Adapter
(785)
this model

Datasets used to train sthanika-ai/Sieve-9B

Spaces using sthanika-ai/Sieve-9B 2

Collection including sthanika-ai/Sieve-9B

Evaluation results

  • Decision Index (chance-corrected; self-reported, not yet on the leaderboard) on Decision Index 0.2.1 (150,317 requests over 44 benchmarks; the index averages 38)
    self-reported
    41.710
  • accuracy on decision-v7 development (1,264 clean questions)
    self-reported
    0.872
  • ECE, at the shipped temperature on decision-v7 development (1,264 clean questions)
    self-reported
    0.028
  • accuracy on transfer-v4 development (656 clean questions)
    self-reported
    0.829
  • Brier, at the shipped temperature on transfer-v4 development (656 clean questions)
    self-reported
    0.242
  • accuracy on transfer-v9 development, strict (936 questions)
    self-reported
    0.753