Skip to content

5. How we measured

The benchmark asks one question 67 times: which of five things does this FontLab user want the built-in assistant to do? It is the router in front of the FontLab assistant, and it is the decision this book was written to get right. Every method got the same 67 queries, and every answer was scored against a label we assigned before the run.

The router question

The assistant can do five kinds of work, and each query belongs to exactly one:

task what the user wants
docs an explanation of how a FontLab feature, tool, panel or setting works, from the documentation, not code
python a Python script that automates FontLab: batch operations, renaming glyphs, generating files
fea OpenType feature code in FEA syntax: liga, kern, calt, ss01, lookups, classes
vfj a glyph definition in VFJ, FontLab's JSON glyph format: contours, nodes, components, anchors
sample a sample text to type in the Glyph window: a pangram, test strings, glyph sequences

The labels split 17 docs, 14 fea, 14 sample, 13 python and 9 vfj.

System One methods received the question as the instruction "This message was sent to the assistant built into FontLab, the font editor. What is the user asking the assistant to do? Pick the single best match." and each task as an option with a description of one to three sentences. Other methods needed shorter prompts, which the section on prompts lists.

The 67 queries

28 queries are in English and 39 are in 29 other languages, 30 languages in all: Polish (3); German, French, Spanish, Italian, Russian, Japanese, Chinese and Korean (2 each); and Czech, Portuguese, Dutch, Turkish, Hungarian, Finnish, Swedish, Ukrainian, Bulgarian, Greek, Armenian, Georgian, Arabic, Persian, Hebrew, Urdu, Hindi, Bengali, Thai and Vietnamese (1 each). They cover Latin with diacritics, Cyrillic, Greek, Arabic, Hebrew, Devanagari, Bengali, Thai, Georgian, Armenian, Han, Hangul and Kana.

Ten of the queries were written to be hard. They sit on the border between two tasks, and a reasonable person could hesitate over each:

  • "How do I write a liga feature?" asks for an explanation (docs) but names FEA.
  • "Give me the code that makes ff become a ligature" is fea, and "Script the kerning for A V" is fea phrased like python.
  • "What would the JSON for a glyph look like?" is vfj, while "Show me how a glyph is stored in a VFJ file" is docs.
  • "Text with lots of kerning pairs" and "Something to type in the Glyph window that shows off my ligatures" are sample, and so is one Polish variant.
  • "Can FontLab batch rename glyphs?" is docs, while "Rename all glyphs in every open font, I don't want to click" is python.

These ten carry most of the differences between the stronger models. On "How do I write a liga feature?", 9% of all methods answered docs, and jev answered fea. Chapter 9 lists every query with the share of methods that got it right.

Translated and direct

Every model was scored twice.

  • Direct: the original query goes to the classifier, in whatever language it was written.
  • Translated: the 39 non-English queries are first translated into English, and the 28 English ones pass through unchanged.

The translated mode models a pipeline in which a small translator runs before the decision. It costs one more model and about 88 ms per non-English query. It helps weaker multilingual models most and helps the strongest ones not at all: jev scored 65/67 either way.

Language detection

The pipeline has to know whether a query is English before it can decide to translate it. Two detectors were measured over 100 passes of the 67 queries, single-threaded, after a warm-up:

DetectorCorrectMean µsCalls/s
lingua63/67604.31,655
papagan43/6712.182,440
majority63/67616.41,622

lingua was right on 63 of 67 queries and handles 1,650 calls per second. papagan is 50 times faster and was right on 43. The majority vote, which breaks ties in lingua's favour, adds nothing to lingua alone. Its four misses: a Spanish and an English query read as Latin, a Bulgarian one as Russian, and a Hindi query full of English terms as Sotho.

Translation

The translator is Hy-MT2-1.8B at Q4_K_M on llama-server, at temperature 0, with a prompt that names the source language and carries a short glossary. The glossary pins font-editing terms that the model otherwise paraphrased in early tests: without it, the Polish kursywy came back as "curve", and the German kern-Feature as "core feature". Terms such as glyph, kerning, feature, italic, VFJ, FEA and Glyph window map to themselves.

The 39 translations, with the time each took:

Click a column header to sort.
QueryLangEnglishms
Jak zmienić kąt pochylenia kursywy w FontLab?plHow can I change the italic angle in FontLab?149
Napisz skrypt Python, który zmieni nazwy wszystkich glifów z .sc na .smcpplWrite a Python script that changes the names of all glyphs from .sc to .smcp112
Schreib mir ein kern-Feature für die Buchstaben A V W TdeWrite me a kern feature for the letters A V W T90
Wo stelle ich die vertikalen Metriken ein?deWhere do I set the vertical metrics?61
Come faccio a esportare il font in OTF?itHow can I export the font to OTF?77
Scrivimi la definizione VFJ di un glifo quadrato largo 500itWrite me the VFJ definition for a square glyph with a width of 50099
Écris une fonctionnalité liga pour fi et flfrWrite a liga functionality for fi and fl70
Donne-moi un pangramme français pour tester mes accentsfrGive me a French pangram to test my accents71
Dame un texto de muestra con muchas letras redondasesGive me a sample text with many round letters74
Escribe un script en Python que elimine los glifos sin unicodeesWrite a Python script that removes glifos without Unicode.85
Jak vyexportovat font jako OTF?csHow to export a font as OTF?66
Escreva o código da feature smcp para versaletesptWrite the code for the smcp feature for Versaletes85
Hoe voeg ik een nieuwe as toe aan een variabel font?nlHow do I add a new font style to a variable font?91
Tüm glifleri 50 birim sağa kaydıran bir Python betiği yaztrA Python script that shifts all glyphs 50 units to the right95
Írj egy kern feature-t az A és V betűk közéhuWrite a kern feature between the letters A and V81
Anna esimerkkiteksti, jossa on paljon ä- ja ö-kirjaimiafiAn example text with many A and O letters72
Viết định nghĩa VFJ cho glyph dấu gạch ngang rộng 300viWrite the VFJ definition for the wide 300-width horizontal glyph90
Ge mig ett pangram på svenska för att testa typsnittetsvGive me a pangram in Swedish to test the font81
Дай мне текст-образец с кириллицей для окна GlyphruGive me a text sample with Cyrillic characters for the Glyph window98
Напиши скрипт на Python, который удалит все направляющиеruWrite a Python script that removes all leading spaces66
Як змінити сімейну назву шрифту?ukHow do I change the family name of a font?83
Напиши VFJ-описание глифа «точка» с юникодом 002EbgWrite the VFJ description for the glyph “dot” with Unicode 002E104
Πώς αλλάζω τη γωνία των πλάγιων γραμμάτων στο FontLab;elHow do I change the angle of the horizontal lines in FontLab?95
Գրիր kern feature Ա և Վ տառերի համարhyGriir kern feature A and V for spacing81
დამიწერე Python სკრიპტი, რომელიც ყველა გლიფს გაასწორებსkaWrite the Python script, which contains many glyphs76
اكتب كود feature للربط بين الحرفين f و iarWrite a feature code to connect the letters f and i81
چطور فونت را به صورت OTF خروجی بگیرم؟faHow can I export the font as an OTF file?85
תן לי טקסט לדוגמה בעברית כדי לבדוק את הריווחheGive me a Hebrew text example so I can check the kerning89
پائتھن اسکرپٹ لکھو جو تمام گلفس کی چوڑائی 600 کر دےurWrite the font script so that all glyphs have a width of 600.104
FontLab में kerning पैनल कहाँ है?hiWhere is the kerning panel in FontLab?77
একটি VFJ গ্লিফ সংজ্ঞা লিখুন যার নাম period এবং প্রস্থ 250bnWrite a VFJ glyph definition named period with a width of 250.99
เขียนสคริปต์ Python ที่เปลี่ยนชื่อ glyph ทั้งหมดเป็นตัวพิมพ์เล็กthWrite a Python script that changes the name of all glyphs to lowercase.99
FontLabでイタリック角度を変更するには?jaHow to change the italic angle in FontLab?77
fi と fl の liga フィーチャーを書いてくださいjaPlease write the features of fi and fl liga.77
给我一段包含所有拼音字母的示例文本,用来检查字体zhGive me an example text containing all phonetic letters to check the font.88
写一个 Python 脚本,把所有字形的宽度设为 600zhWrite a Python script to set the width of all glyphs to 600.97
A와 V 사이의 커닝 feature 코드를 작성해 줘koWrite the kerning feature code between A and V81
VFJ 형식으로 이름이 hyphen이고 너비가 300인 글리프 정의를 작성해 줘koThe FontLab file has a hyphen name and a glyph definition with a width of 300.113
Napisz coś do okna Glyph, żeby sprawdzić ligatury fi flplWrite something in the Glyph window to check the ligatures fi fl94

Each query was translated once and the English text was cached, so every classifier in translated mode saw the same English.

What each method was asked

Engines and runtimes take different request shapes, so the same decision was phrased in the shape each readout was built for:

engine or readout prompt
jev, dohnuts, laya.cpp (kev), PyTorch, ExecuTorch a System One request: the query as state, one choice question with the long instruction and the five long task descriptions as criteria
slot, decider readout decider's trained layout, tokenized as upstream does it: Context: and the query, then the question, the options lettered (A) to (E) with their long descriptions, then Answer: (
slot, Hmm readout the Qwen chat template with thinking off, options as A: name — description lines, and "Return only the option letter."
pcdServer one enum field task with the choices docs, python, fea, vfj, sample, and a field description made of a short instruction and one short description per task
laya runtimes a System One request with the short instruction and short descriptions, because the encoder shares a 1,024-token budget; the Neural Engine exports hold 96 tokens and got a compact prompt

pcdServer cannot attach a description to each allowed value, which is why its task descriptions live in the field description. Chapter 3 explains what each readout does with its prompt.

One model at a time

Every model was loaded alone, measured, and unloaded before the next one started. The rule comes from an accident. An early driver loaded several GGUF files per batch (up to 9 GB of weights at once) next to the resident translation model, while other experiments ran beside it. The machine, an Apple M4 Max with 48 GB of memory, swapped until its boot disk was full and macOS stopped responding.

Since then the harness checks that no model process is running before it loads one. A watchdog polls every two seconds and aborts a model's run if swap grows by more than 4 GB, free space on the boot disk falls below 30 GB, available memory falls below 4 GB, or the run passes 30 minutes. With those guards in place, swap stayed flat at about 2.7 GB for every model, the 21 GB decider-35b-a3b included. Chapter 8 turns this into rules for your own code.

One model at a time also makes the timings fair: no model competes with another for the GPU or for memory bandwidth.

What the numbers mean

  • Translated and direct count correct answers out of 67. A correct answer is one whose most likely option is the labelled task.
  • The English column counts the 28 English queries in translated mode. Other, translated and other, direct count the 39 non-English queries in each mode.
  • ms/query is the mean wall-clock time from the moment the harness sends a query to the moment it has the probabilities. For server engines it includes the HTTP round trip and JSON parsing on loopback. For jev it includes the network trip to OpenRouter. It excludes model loading and a warm-up call made before timing starts.
  • Load ms is the time to load the model and answer the warm-up call, for runtimes that load in-process or that the harness started itself. It is blank where a server was started outside the timed step.
  • GB is the size of the GGUF file as published on Hugging Face.

In-process runtimes (MLX, Core ML, ONNX Runtime) have no HTTP in their times, so compare their milliseconds with a server's only loosely.

Reproducibility

  • Engines. dohnuts.cpp at upstream commit 9a894b0, with llama.cpp pinned by its submodule at b29c606. The benchmark predates our prefix cache (chapter 8), which leaves single-question requests unchanged. pcdServer at commit 1ce9e55, which fetches llama.cpp tag v0.4.1 when it is configured. slot used a stock llama-server.
  • Translator. Hy-MT2-1.8B at Q4_K_M, temperature 0, with the glossary above.
  • Dates. The runs, the jev calls included, took place on 2026-09-22 and 2026-09-23.
  • Data. The harness itself is private. Every method's scores, timings and file sizes are published in src_docs/data/ in the ornotto repository, and the tables in this book are generated from them.

Limits

  • 67 queries is a small set. One query is 1.5 percentage points of accuracy. Differences of one or two queries between models are within what a different set of 67 would change.
  • The labels are ours. For the ten ambiguous queries another person could label a few differently. jev's two errors are both on ambiguous queries.
  • Thresholds are in-sample. The confidence gates in chapter 9 were tuned on the same 67 queries they are scored on. Expect them to do somewhat worse on new traffic.
  • One machine. All timings come from one Apple M4 Max, on Metal where the engine or runtime supports it. CPU-only machines are several times slower: dohnuts on the CPU with eight threads took 275 ms per query against 52 ms on Metal.
  • One task. A five-way router is one decision. Yes/no questions, rubrics and long option lists behave differently, and chapter 2 shows examples of each.
  • Only the winner is scored. Accuracy looks at the most likely option. It ignores how confident the model was, which matters as soon as you gate on confidence (chapter 9).