ornotto¶
ornotto asks a local language model to pick an answer, not to write one. You give it a state (a message, or any JSON value) and a question with a closed set of answers: one of these options, yes or no, or a level on a rubric. It returns the answer and a probability for every alternative. Because nothing is generated, there is no output to parse, repair or retry, and decider-0.8b answers in about 52 ms on an Apple M4 Max.
The package runs two open-source C++ engines on llama.cpp behind one Python API. dohnuts.cpp runs models trained for this job and reads their answer at a trained slot. pcdServer runs any chat GGUF and scores only the tokens of the answers you allow. This book explains how both work, compares them with TypeSafe's hosted jev model and with other local engines and runtimes on 67 real queries, and documents the package.
Install¶
uv add ornotto # or: pip install ornotto
uv add "ornotto[pydantic-ai]" # with the pydantic-ai integration
Decide¶
import ornotto
ornotto.choose("Build a kern feature for A V W T", ["docs", "python", "fea"]).value # 'fea'
answer = ornotto.check("Write me a script that renames glyphs", "Does the user want code?")
answer.value, answer.probability # (True, 0.6)
The first call downloads decider-0.8b (0.81 GB) from Hugging Face and starts dohnuts on a loopback port. Later calls reuse both.
The best 15 of 208¶
Each row is one method, a model file read one way, asked which of five tasks a FontLab user wants, over 67 queries in 30 languages. Click a column header to sort. Chapter 6 has all 208 rows.
| Engine | Model | Family | Quant | GB | ms/query | Translated | Direct |
|---|---|---|---|---|---|---|---|
| jev (hosted) | jev | dedicated | 466.1 | 65/67 | 65/67 | ||
| pcdServer | qwen3.5-4b-hmm | fine-tuned | q8 | 4.48 | 87.6 | 64/67 | 65/67 |
| pcdServer | qwen3.5-4b-hmm | fine-tuned | q4 | 2.71 | 92.6 | 64/67 | 65/67 |
| pcdServer | qwen3.5-4b-hmm | fine-tuned | bf16 | 8.42 | 94.8 | 64/67 | 65/67 |
| dohnuts (Metal) | decider-35b-a3b | dedicated | q4 | 21.17 | 266.3 | 64/67 | 64/67 |
| slot (llama-server) | decider-35b-a3b | dedicated | q4 | 21.17 | 299.1 | 64/67 | 64/67 |
| PyTorch | decider-4b | dedicated | bf16 | 348.2 | 63/67 | 65/67 | |
| pcdServer | qwen3.5-4b-unsloth | vanilla | q8 | 4.48 | 85.0 | 63/67 | 64/67 |
| pcdServer | qwen3.5-4b-unsloth | vanilla | q5 | 3.14 | 94.5 | 63/67 | 64/67 |
| pcdServer | decider-35b-a3b | dedicated | q4 | 21.17 | 296.8 | 63/67 | 63/67 |
| pcdServer | qwen3.8-9b-distill | fine-tuned | q4 | 5.78 | 137.8 | 62/67 | 64/67 |
| pcdServer | qwen3.5-4b | vanilla | q4 | 3.01 | 89.8 | 62/67 | 63/67 |
| pcdServer | qwen3.5-4b-unsloth | vanilla | q6 | 3.53 | 93.6 | 62/67 | 63/67 |
| pcdServer | qwen3.8-4b-distill | fine-tuned | q4 | 2.78 | 93.6 | 62/67 | 63/67 |
| slot (llama-server) | qwen3.5-4b-hmm | fine-tuned | bf16 | 8.42 | 284.2 | 62/67 | 63/67 |
What is in this book¶
- Concepts. Chapter 1 says what a System One decision is. Chapter 2 covers what you ask and what you get back. Chapter 3 shows the five readouts that get an answer out of a model, and defines the terms the rest of the book uses. Chapter 4 covers dedicated, fine-tuned and vanilla models.
- Benchmarks. Chapter 5 is the method. Chapter 6 has the full sortable results. Chapter 7 covers quantization, chapter 8 speed and caching, and chapter 9 confidence and fallback.
- The package. Chapter 10 is the API reference. Chapter 11 covers typed decisions and pydantic-ai agents. Chapter 12 is about choosing an engine and model, building from source and contributing.