4. Dedicated, fine-tuned and vanilla models¶
A model's size tells you less about how it will decide than the place its answer is read from. This book sorts every model into one of three families by that place. A dedicated model was trained to answer System One questions and has its own readout: a letter at a trained answer slot, a pointer head, or a decision head on an encoder. A fine-tuned model is a chat model that someone tuned for a task, sometimes this task, sometimes a neighbouring one. A vanilla model is a stock chat model as its publisher released it.
The family decides which engines can run the model. It also decides how much prompt work you have to do before its answers are worth reading.
| family | where the answer comes from | engines in ornotto |
engines and runtimes in the benchmark |
|---|---|---|---|
| dedicated | the model's own trained readout | dohnuts (decider, kev, Dohnuts), pcdServer (decider only, as a chat model) | jev, dohnuts, slot, pcdServer, laya runtimes, PyTorch, ExecuTorch |
| fine-tuned | the first token of each allowed value, under a chat template | pcdServer | pcdServer, slot (Hmm readout) |
| vanilla | the first token of each allowed value, under a chat template | pcdServer | pcdServer |
The best row of every model in the benchmark, grouped by family:
| Family | Model | Engine | GB | ms/query | Translated | Direct | Best method |
|---|---|---|---|---|---|---|---|
| dedicated | jev | jev (hosted) | 466.1 | 65/67 | 65/67 | jev | |
| dedicated | decider-35b-a3b | dohnuts (Metal) | 21.17 | 266.3 | 64/67 | 64/67 | dohnuts-metal@decider-35b-a3b-q4 |
| dedicated | decider-4b | PyTorch | 348.2 | 63/67 | 65/67 | torch@decider-4b-bf16 | |
| dedicated | decider-0.8b | dohnuts (Metal) | 0.81 | 51.4 | 62/67 | 61/67 | dohnuts-metal@decider-0.8b-q8 |
| dedicated | decider-0.8b-dreamblooms | dohnuts (Metal) | 0.81 | 52.2 | 62/67 | 61/67 | dohnuts-metal@decider-0.8b-dreamblooms-q8 |
| dedicated | kev-4b | laya.cpp | 1,103.1 | 62/67 | 55/67 | kev-4b-gguf | |
| dedicated | decider-2b | pcdServer | 3.78 | 42.2 | 60/67 | 60/67 | pcdserver@decider-2b-f16 |
| dedicated | decider-2b-dreamblooms | pcdServer | 2.01 | 40.1 | 59/67 | 60/67 | pcdserver@decider-2b-dreamblooms-q8 |
| dedicated | kev-0.8b | laya.cpp | 231.5 | 57/67 | 45/67 | kev-0.8b-gguf | |
| dedicated | laya-multilingual | MLX | 6.9 | 52/67 | 49/67 | laya-mlx | |
| dedicated | laya-typed-decisions | PyTorch | 17.7 | 48/67 | 43/67 | laya-typed-decisions | |
| fine-tuned | qwen3.5-4b-hmm | pcdServer | 4.48 | 87.6 | 64/67 | 65/67 | pcdserver@qwen3.5-4b-hmm-q8 |
| fine-tuned | qwen3.8-4b-distill | pcdServer | 2.78 | 93.6 | 62/67 | 63/67 | pcdserver@qwen3.8-4b-distill-q4 |
| fine-tuned | qwen3.8-9b-distill | pcdServer | 5.78 | 137.8 | 62/67 | 64/67 | pcdserver@qwen3.8-9b-distill-q4 |
| fine-tuned | qwen3.5-4b-supercoder | pcdServer | 2.61 | 84.2 | 61/67 | 62/67 | pcdserver@qwen3.5-4b-supercoder-q4 |
| fine-tuned | qwen2.5-coder-7b | pcdServer | 4.68 | 108.9 | 61/67 | 62/67 | pcdserver@qwen2.5-coder-7b-q4 |
| fine-tuned | qwen2.5-coder-3b | pcdServer | 2.10 | 60.2 | 60/67 | 59/67 | pcdserver@qwen2.5-coder-3b-q4 |
| fine-tuned | pulse-0.6b | pcdServer | 0.40 | 27.9 | 54/67 | 52/67 | pcdserver@pulse-0.6b-q4 |
| fine-tuned | lq-decide-1.7b | pcdServer | 1.83 | 35.1 | 50/67 | 46/67 | pcdserver@lq-decide-1.7b-q8 |
| fine-tuned | qwen3.8-2b-distill | pcdServer | 1.31 | 42.9 | 49/67 | 51/67 | pcdserver@qwen3.8-2b-distill-q4 |
| fine-tuned | zerocoder-0.8b | pcdServer | 0.38 | 34.7 | 42/67 | 36/67 | pcdserver@zerocoder-0.8b-q2 |
| fine-tuned | qwen3-reranker-0.6b | pcdServer | 0.44 | 24.7 | 40/67 | 44/67 | pcdserver@qwen3-reranker-0.6b-q5 |
| fine-tuned | lyco-router-0.6b | pcdServer | 0.48 | 26.4 | 35/67 | 33/67 | pcdserver@lyco-router-0.6b-q4 |
| fine-tuned | qwen2.5-0.5b-pcb | pcdServer | 0.34 | 18.1 | 27/67 | 23/67 | pcdserver@qwen2.5-0.5b-pcb-q2 |
| fine-tuned | euaiact-0.5b | pcdServer | 0.40 | 23.8 | 20/67 | 20/67 | pcdserver@euaiact-0.5b-q4 |
| fine-tuned | lq-decide-0.6b | pcdServer | 0.64 | 23.1 | 11/67 | 11/67 | pcdserver@lq-decide-0.6b-q8 |
| vanilla | qwen3.5-4b-unsloth | pcdServer | 4.48 | 85.0 | 63/67 | 64/67 | pcdserver@qwen3.5-4b-unsloth-q8 |
| vanilla | qwen2.5-3b | pcdServer | 3.29 | 53.3 | 62/67 | 57/67 | pcdserver@qwen2.5-3b-q8 |
| vanilla | qwen3.5-4b | pcdServer | 3.01 | 89.8 | 62/67 | 63/67 | pcdserver@qwen3.5-4b-q4 |
| vanilla | bonsai-27b | pcdServer | 3.80 | 513.5 | 62/67 | 63/67 | pcdserver@bonsai-27b-q1 |
| vanilla | qwen3.5-2b | pcdServer | 1.22 | 42.7 | 61/67 | 62/67 | pcdserver@qwen3.5-2b-q3 |
| vanilla | qwen3-4b-unsloth | pcdServer | 8.05 | 71.5 | 61/67 | 62/67 | pcdserver@qwen3-4b-unsloth-bf16 |
| vanilla | qwen3-4b | pcdServer | 2.50 | 75.5 | 61/67 | 62/67 | pcdserver@qwen3-4b-q4 |
| vanilla | qwen3-8b | pcdServer | 5.03 | 116.0 | 61/67 | 62/67 | pcdserver@qwen3-8b-q4 |
| vanilla | qwen3.5-9b-unsloth | pcdServer | 5.68 | 136.9 | 61/67 | 63/67 | pcdserver@qwen3.5-9b-unsloth-q4 |
| vanilla | qwen3.5-0.8b | pcdServer | 0.65 | 32.1 | 57/67 | 57/67 | pcdserver@qwen3.5-0.8b-q5 |
| vanilla | qwen3-1.7b-unsloth | pcdServer | 1.83 | 35.1 | 57/67 | 59/67 | pcdserver@qwen3-1.7b-unsloth-q8 |
| vanilla | qwen3-1.7b | pcdServer | 2.17 | 35.2 | 57/67 | 59/67 | pcdserver@qwen3-1.7b-q8 |
| vanilla | qwen3.5-2b-unsloth | pcdServer | 2.01 | 39.5 | 57/67 | 59/67 | pcdserver@qwen3.5-2b-unsloth-q8 |
| vanilla | qwen3.5-0.8b-unsloth | pcdServer | 0.59 | 29.6 | 56/67 | 57/67 | pcdserver@qwen3.5-0.8b-unsloth-q5 |
| vanilla | qwen3.5-0.8b-ggml | pcdServer | 1.56 | 30.3 | 55/67 | 57/67 | pcdserver@qwen3.5-0.8b-ggml-bf16 |
| vanilla | bonsai-1.7b | pcdServer | 0.49 | 34.9 | 54/67 | 55/67 | pcdserver@bonsai-1.7b-q2 |
| vanilla | qwen3-0.6b | pcdServer | 0.64 | 25.5 | 53/67 | 54/67 | pcdserver@qwen3-0.6b-q8 |
| vanilla | qwen1.5-4b | pcdServer | 4.20 | 247.8 | 42/67 | 43/67 | pcdserver@qwen1.5-4b-q8 |
| vanilla | qwen1.5-0.5b | pcdServer | 0.35 | 21.2 | 35/67 | 32/67 | pcdserver@qwen1.5-0.5b-q3 |
| vanilla | qwen1.5-1.8b | pcdServer | 1.02 | 41.8 | 34/67 | 35/67 | pcdserver@qwen1.5-1.8b-q3 |
| vanilla | qwen2.5-0.5b | pcdServer | 0.49 | 20.2 | 28/67 | 26/67 | pcdserver@qwen2.5-0.5b-q4 |
Dedicated models¶
decider¶
Mapika/decider is a family of full fine-tunes of Qwen3.5 trained to answer System One questions. The prompt is fixed by training: Context: and the state, then the question, the options lettered (A), (B), …, and Answer: (. The model is read at that last token, and only the logits of the option letters count. They are divided by a temperature from the checkpoint's decider.json (1.03 for the 0.8B model) before the softmax, so the probabilities are calibrated by the recipe.
The benchmark ran four sizes, from several converters:
| model | GGUF source | licence | best result (translated / direct) |
|---|---|---|---|
| decider-0.8b | DreamBlooms/decider-0.8b-GGUF, mradermacher/decider-0.8b-GGUF | Apache-2.0 | 62/67 and 62/67 on dohnuts (Metal), mradermacher Q6_K, 54 ms; the DreamBlooms Q8_0 file ornotto registers: 62/67 and 61/67, 52 ms |
| decider-2b | DreamBlooms/decider-2b-GGUF, cosetoenor/decider-2b-GGUF | Apache-2.0 | 59/67 and 62/67 on dohnuts (Metal); 60/60 on pcdServer |
| decider-4b | Mapika/decider-4b, in Mapika's PyTorch package | Apache-2.0 | 63/67 and 65/67, 348 ms |
| decider-35b-a3b | mradermacher/decider-35b-a3b-GGUF | Apache-2.0 | 64/67 and 64/67 on dohnuts (Metal), 266 ms, 21 GB |
The DreamBlooms 2B build is a newer checkpoint with calibration-aware reinforcement learning and a temperature of 1.3. On this set it scores the same as the older 2B builds: calibration work shows up in the confidences, which the accuracy column does not see.
decider runs on dohnuts with its own readout, and on pcdServer as an ordinary chat model. The two are not the same measurement. Chapter 6 shows the 0.8B model scoring 62 through its readout and 55 through pcdServer.
In ornotto, decider-0.8b and decider-2b are registered and default to dohnuts:
import ornotto
fast = ornotto.Decider("decider-0.8b") # dohnuts, the trained readout
same_weights = ornotto.Decider("decider-0.8b", engine="pcd") # pcdServer, as a chat model
kev¶
jaredpalmer/kev is a LoRA over Qwen3.5 plus a bilinear pointer head. The prompt marks the decision point and the end of every option with special tokens, and the head scores each option by the dot product of two projected hidden states. The benchmark ran the mys/kev-0.8b-GGUF and kev-4b conversions through laya.cpp's laya serve. kev-4b reached 62/67 translated but 55/67 direct, and took 1,103 ms per query. kev-0.8b dropped to 45/67 on untranslated text.
ornotto registers kev-0.8b from DreamBlooms/kev-0.8b-GGUF (Apache-2.0) on dohnuts, which reads the same pointer head. That build was not part of the benchmark run.
Dohnuts¶
PsiACE/Dohnuts is the model dohnuts.cpp was first written for: Qwen3.5-0.8B with language LoRA adapters and one scalar head read at each candidate marker. It also takes images through an mmproj projector. ornotto registers it as dohnuts-0.8b.
Non-commercial weights
The Dohnuts weights are licensed CC-BY-NC-SA-4.0. The engine is Apache-2.0, the weights are not. Dohnuts was not benchmarked here and is not the ornotto default.
laya-multilingual¶
convaiinnovations/laya-multilingual (Apache-2.0) is an encoder with a decision head, not a chat model; chapter 3 describes its readout and its token budget.
The same checkpoint ran on six engines and runtimes. Every method at 8-bit precision or better with the full prompt scored 52/67 translated and 49/67 direct, so the engine or runtime changed only the speed: 6.9 ms per query on MLX, 57.6 ms on laya.cpp. The compact prompt of the Neural Engine exports cost four to five answers. Of the two 4-bit conversions, one lost an answer and the other lost 16. ornotto does not run laya.
jev¶
jev is TypeSafe's hosted System One model and the best result in the benchmark (chapter 3). Its internals are not public, so the benchmark treats it as the reference, not as a design to copy. dohnuts answers the same request shape, which is why pydantic-ai's TypeSafe model can drive a local engine (chapter 11).
Fine-tuned models¶
These are chat models tuned by third parties. None of them has a trained answer slot, so every one ran on pcdServer, which scores the first token of each allowed value under the model's chat template.
n4ze3m/Qwen3.5-4B-Hmm (Apache-2.0) is the exception worth knowing: it was tuned for this kind of pick-one decision. On pcdServer it scored 64/67 translated and 65/67 direct at every quantization from Q4 to BF16, the best local result in the benchmark, at 88 to 95 ms per query. Read through its own letter readout on llama-server (the slot rows), the same weights scored 62/67 and 63/67 and took three times as long.
The rest spread from close to the top to the bottom of the table:
- The empero-ai Qwen3.8 distills reached 62/67 at 4B and 9B, and 49/67 at 2B.
- The coding models (Qwen2.5-Coder 3B and 7B, jica98's Qwen3.5-4B super-coder) scored 60 to 61.
- Small models tuned for routing, deciding or retrieval scored worse than a vanilla model of the same size: pulse-0.6b 54, lq-decide-1.7b 50, lq-decide-0.6b 11, lyco-router-0.6b 35, Qwen3-Reranker-0.6B 40.
- DavidAU's Qwen3 Zero-Coder 0.8B scored 18 to 42 across seven quantizations, with no order by bit count, well below vanilla Qwen3.5-0.8B's best of 57.
- Qwen2.5-0.5B-PCb reached 27 at best, one answer below vanilla Qwen2.5-0.5B's 28, and EUAIAct-Qwen2.5-0.5B reached 20.
A tune for a nearby task does not carry over to this one. If you pick a fine-tuned model, check it on your own questions first.
ornotto registers qwen3.5-4b-hmm (the Q4_K_M file, 2.7 GB) on pcdServer:
Vanilla models¶
Vanilla chat models need nothing but a chat template that llama.cpp understands, so any GGUF works on pcdServer. The benchmark ran Qwen3.5 (0.8B, 2B, 4B, 9B), Qwen3 (0.6B, 1.7B, 4B, 8B), Qwen2.5 (0.5B, 3B), Qwen1.5 (0.5B, 1.8B, 4B) and PrismML's Bonsai (1.7B, 27B), most of them at several quantizations.
- Qwen3.5-4B reached 62/67 and 63/67 at Q4; unsloth's Q8 build reached 63/67 and 64/67, one answer below the Hmm tune.
- Qwen3.5-2B at Q3_K_M reached 61/67 translated and 62/67 direct at 43 ms per query from a 1.2 GB file.
- Qwen3.5-0.8B reached 57/67 at Q5 and 55/67 at Q8, at about 30 ms.
- Qwen2.5-3B reached 62/67 on translated text but only 57 to 60 on direct text: it handles English well and other languages less well.
- Qwen1.5 scored 42/67 at best, at 4B. An older generation is no substitute for a smaller new one.
- Bonsai-27B at 1 bit (3.8 GB) reached 62/67 and 63/67, but took 514 ms per query.
Licences vary within the Qwen family. Qwen3 and Qwen3.5 are Apache-2.0. Qwen2.5-3B and Qwen2.5-Coder-3B use the Qwen Research licence, and Qwen1.5 the Tongyi Qianwen research licence. Read the model card before you ship one.
ornotto registers qwen3.5-0.8b, qwen3.5-2b and qwen3.5-4b on pcdServer. Any other GGUF works by path or by Hugging Face file:
mine = ornotto.Decider("~/models/Qwen_Qwen3.5-2B-Q3_K_M.gguf") # a local file
hub = ornotto.Decider("hf:bartowski/Qwen_Qwen3.5-2B-GGUF/Qwen_Qwen3.5-2B-Q3_K_M.gguf") # downloaded on first use
An unregistered GGUF runs on pcdServer only. To run one on dohnuts, pass the profile JSON as metadata= (and the scorer head as head= for kev or Dohnuts): dohnuts needs to know which readout the weights were trained for.
Which family for which job¶
- If you want calibrated probabilities you can threshold, use a dedicated model on dohnuts. Only the dedicated readouts divide by a fitted temperature.
- If you want the most accurate local answer and can spend about 90 ms and 2.7 GB, use Qwen3.5-4B-Hmm on pcdServer.
- If you have a model already, or need one that no one has tuned, a vanilla Qwen3.5 on pcdServer works with no training at all. Qwen3.5-2B at Q3 is the smallest vanilla model that stays within one answer of the dedicated 0.8B model.
Chapter 7 shows how far each family can be quantized, and chapter 12 turns these results into a choice.