kapi steers model output by injecting context: a terms store that mandates renderings, a voice guide, an instruction. This page measures whether each model actually follows that context — not whether it translates well, which is a different question. The sibling batch eval measures structural integrity and cost; this one measures obedience.
How it is measured
The core metric is a differential. An engineered corpus is translated twice through the production pipeline — once bare (no context) and once steered (terms + voice profile + instruction, injected exactly as production injects them) — and both passes are scored with kapi’s own deterministic check tools (term-check, dnt-check, voice-vocab-check, pattern-check). Two numbers fall out per dimension: absolute adherence (did the steered output satisfy the requirement) and lift (steered minus bare — how much the context moved the model). Lift is the decision-relevant one: a model with high absolute adherence but no lift already “knew” it, and our context earns no credit. A model with high lift is genuinely steerable, which is what context injection is buying.
Every fixture is a trap: a naive translation violates the context. A mandated term whose natural rendering differs from the mandate, a product name that reads like a common noun, casual English tempting an informal register, a source that ends in the exclamation mark the instruction forbids. There are distractors (a lowercase “compass” that is a real compass and must be translated) and declared-winner conflicts (the terms store pins a compound containing a forbidden word — the pin wins, and the scorer knows it). Results are reported per dimension — terminology, voice, instruction — never as one collapsed score: a model can be excellent at terminology and poor at voice, and the collapsed number would hide the thing you would act on.
The deterministic checks are the backbone. The genuinely subjective remainder of voice — register, naturalness, restraint — is scored by a cross-family LLM judge under a fixed yes/no rubric, blind to which model and which variant produced a text, and its scores are published only once judge–human agreement has been measured above a stated bar. Until then, judged numbers stay in the record but off this page.
What the measurement found
The most steerable model measured is haiku ( nb ), which the context moved by +73.1pp overall.Every measured model shows positive lift — the context earns its tokens everywhere.The weakest dimension with context applied is instruction at 97.6% steered adherence — the ceiling context injection currently hits.
Target de measured 2026-08-27 · 20 fixtures, 32 checks (17 terminology, 7 voice, 8 instruction) · corpus 66cc1693369f
haiku (claude-code) sonnet (claude-code) open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
haiku
+41.2pp
+42.9pp
+62.5pp
+46.9pp
—
—
—
sonnet
+47.1pp
+42.9pp
+75.0pp
+53.1pp
—
—
—
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
eu.anthropic.claude-sonnet-4-6 (bedrock) haiku (claude-code) opus (claude-code) sonnet (claude-code) gemini-3.1-flash-lite (gemini) gemini-3.1-pro-preview (gemini) gemini-3.5-flash (gemini) open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
eu.anthropic.claude-sonnet-4-6
+100.0pp
+100.0pp
+50.0pp
+70.0pp
150
+466.7pp
$0.014
haiku
+100.0pp
+100.0pp
+41.7pp
+65.0pp
—
—
—
opus
+100.0pp
+100.0pp
+50.0pp
+70.0pp
—
—
—
sonnet
+100.0pp
+100.0pp
+50.0pp
+70.0pp
—
—
—
gemini-3.1-flash-lite
+100.0pp
+100.0pp
+50.0pp
+70.0pp
137
+510.9pp
$0.0011
gemini-3.1-pro-preview
+100.0pp
+100.0pp
+50.0pp
+70.0pp
137
+510.9pp
$0.0076
gemini-3.5-flash
+100.0pp
+100.0pp
+50.0pp
+70.0pp
137
+510.9pp
$0.0047
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
eu.anthropic.claude-sonnet-4-6 bedrock · 1× per variant
haiku (claude-code) sonnet (claude-code) open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
haiku
+52.9pp
+33.3pp
+75.0pp
+54.8pp
—
—
—
sonnet
+52.9pp
+16.7pp
+75.0pp
+51.6pp
—
—
—
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
haiku claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-conflict
1
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / formality
3
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
voice / vocab-conflict
1
100%
100%
+0.0pp
sonnet claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-conflict
1
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / formality
3
100%
100%
+0.0pp
voice / vocab
2
50%
100%
+50.0pp
voice / vocab-conflict
1
100%
100%
+0.0pp
Target nb measured 2026-08-27 · 17 fixtures, 26 checks (16 terminology, 2 voice, 8 instruction) · corpus daf7de0cd29f
haiku (claude-code) sonnet (claude-code) open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
haiku
+68.8pp
+100.0pp
+75.0pp
+73.1pp
—
—
—
sonnet
+50.0pp
+100.0pp
+50.0pp
+53.8pp
—
—
—
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
haiku claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
50%
100%
+50.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
sonnet claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
33%
+33.3pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
Reproducing this
The harness is scripts/contexteval. A run against the built-in demo stub exercises the harness and measures nothing about any model; such runs are marked simulated and are excluded from every chart on this page. Adherence varies by target language, so the published sweep covers more than one.
make context-eval # demo stub: proves the harness, measures nothing
make context-eval-publish # the real sweep → this page's data