kapi steers model output by injecting context: a terms store that mandates renderings, a brand voice guide, an instruction. This page measures whether each model actually follows that context — not whether it translates well, which is a different question. The sibling batch eval measures structural integrity and cost; this one measures obedience.
How it is measured
The core metric is a differential. An engineered corpus is translated twice through the production pipeline — once bare (no context) and once steered (terms + voice profile + instruction, injected exactly as production injects them) — and both passes are scored with kapi’s own deterministic check tools (term-check, dnt-check, brand-vocab-check, pattern-check). Two numbers fall out per dimension: absolute adherence (did the steered output satisfy the requirement) and lift (steered minus bare — how much the context moved the model). Lift is the decision-relevant one: a model with high absolute adherence but no lift already “knew” it, and our context earns no credit. A model with high lift is genuinely steerable, which is what context injection is buying.
Every fixture is a trap: a naive translation violates the context. A mandated term whose natural rendering differs from the mandate, a product name that reads like a common noun, casual English tempting an informal register, a source that ends in the exclamation mark the instruction forbids. There are distractors (a lowercase “compass” that is a real compass and must be translated) and declared-winner conflicts (the terms store pins a compound containing a forbidden word — the pin wins, and the scorer knows it). Results are reported per dimension — terminology, voice, instruction — never as one collapsed score: a model can be excellent at terminology and poor at voice, and the collapsed number would hide the thing you would act on.
The deterministic checks are the backbone. The genuinely subjective remainder of voice — register, naturalness, restraint — is scored by a cross-family LLM judge under a fixed yes/no rubric, blind to which model and which variant produced a text, and its scores are published only once judge–human agreement has been measured above a stated bar. Until then, judged numbers stay in the record but off this page.
What the measurement found
The most steerable model measured is eu.anthropic.claude-sonnet-4-6 (en-GB), which the context moved by +70.0pp overall. Every measured model shows positive lift — the context earns its tokens everywhere. The weakest dimension with context applied is terminology at 98.7% steered adherence — the ceiling context injection currently hits.
Target de measured 2026-07-17 · 20 fixtures, 32 checks (17 terminology, 7 voice, 8 instruction) · corpus 041e6a648e6c
eu.anthropic.claude-sonnet-4-6(bedrock)haiku(claude-code)opus(claude-code)sonnet(claude-code)gemini-3.1-flash-lite(gemini)gemini-3.1-pro-preview(gemini)gemini-3.5-flash(gemini)open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
eu.anthropic.claude-sonnet-4-6
+47.1pp
+42.9pp
+75.0pp
+53.1pp
213
+249.4pp
$0.021
haiku
+47.1pp
+28.6pp
+62.5pp
+46.9pp
—
—
—
opus
+41.2pp
+42.9pp
+75.0pp
+50.0pp
—
—
—
sonnet
+47.1pp
+42.9pp
+62.5pp
+50.0pp
—
—
—
gemini-3.1-flash-lite
+35.3pp
+42.9pp
+75.0pp
+46.9pp
178
+263.3pp
$0.0015
gemini-3.1-pro-preview
+44.1pp
+42.9pp
+75.0pp
+51.6pp
178
+289.7pp
$0.011
gemini-3.5-flash
+41.2pp
+42.9pp
+75.0pp
+50.0pp
178
+280.9pp
$0.0079
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
eu.anthropic.claude-sonnet-4-6 bedrock · 1× per variant
eu.anthropic.claude-sonnet-4-6(bedrock)haiku(claude-code)opus(claude-code)sonnet(claude-code)gemini-3.1-flash-lite(gemini)gemini-3.1-pro-preview(gemini)gemini-3.5-flash(gemini)open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
eu.anthropic.claude-sonnet-4-6
+100.0pp
+100.0pp
+50.0pp
+70.0pp
150
+466.7pp
$0.014
haiku
+100.0pp
+100.0pp
+41.7pp
+65.0pp
—
—
—
opus
+100.0pp
+100.0pp
+50.0pp
+70.0pp
—
—
—
sonnet
+100.0pp
+100.0pp
+50.0pp
+70.0pp
—
—
—
gemini-3.1-flash-lite
+100.0pp
+100.0pp
+50.0pp
+70.0pp
137
+510.9pp
$0.0011
gemini-3.1-pro-preview
+100.0pp
+100.0pp
+50.0pp
+70.0pp
137
+510.9pp
$0.0076
gemini-3.5-flash
+100.0pp
+100.0pp
+50.0pp
+70.0pp
137
+510.9pp
$0.0047
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
eu.anthropic.claude-sonnet-4-6 bedrock · 1× per variant
eu.anthropic.claude-sonnet-4-6(bedrock)haiku(claude-code)opus(claude-code)sonnet(claude-code)gemini-3.1-flash-lite(gemini)gemini-3.1-pro-preview(gemini)gemini-3.5-flash(gemini)open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
eu.anthropic.claude-sonnet-4-6
+52.9pp
+33.3pp
+75.0pp
+54.8pp
197
+278.4pp
$0.019
haiku
+52.9pp
+0.0pp
+75.0pp
+48.4pp
—
—
—
opus
+52.9pp
+33.3pp
+75.0pp
+54.8pp
—
—
—
sonnet
+52.9pp
+33.3pp
+75.0pp
+54.8pp
—
—
—
gemini-3.1-flash-lite
+47.1pp
+33.3pp
+75.0pp
+51.6pp
172
+300.1pp
$0.0014
gemini-3.1-pro-preview
+50.0pp
+33.3pp
+75.0pp
+53.2pp
172
+309.5pp
$0.0083
gemini-3.5-flash
+52.9pp
+33.3pp
+75.0pp
+54.8pp
172
+318.8pp
$0.0061
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
eu.anthropic.claude-sonnet-4-6 bedrock · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-conflict
1
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / formality
3
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
voice / vocab-conflict
1
100%
100%
+0.0pp
haiku claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-conflict
1
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / formality
3
100%
100%
+0.0pp
voice / vocab
2
50%
50%
+0.0pp
voice / vocab-conflict
1
100%
100%
+0.0pp
opus claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-conflict
1
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / formality
3
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
voice / vocab-conflict
1
100%
100%
+0.0pp
sonnet claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-conflict
1
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / formality
3
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
voice / vocab-conflict
1
100%
100%
+0.0pp
gemini-3.1-flash-lite gemini · 2× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
6
0%
100%
+100.0pp
instruction / exclamation
6
0%
100%
+100.0pp
instruction / verbatim
4
100%
100%
+0.0pp
terminology / dnt
12
100%
100%
+0.0pp
terminology / dnt-distractor
2
100%
100%
+0.0pp
terminology / term
16
0%
100%
+100.0pp
terminology / term-conflict
2
0%
0%
+0.0pp
terminology / term-distractor
2
100%
100%
+0.0pp
voice / formality
6
100%
100%
+0.0pp
voice / vocab
4
0%
100%
+100.0pp
voice / vocab-conflict
2
100%
100%
+0.0pp
gemini-3.1-pro-preview gemini · 2× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
6
0%
100%
+100.0pp
instruction / exclamation
6
0%
100%
+100.0pp
instruction / verbatim
4
100%
100%
+0.0pp
terminology / dnt
12
100%
100%
+0.0pp
terminology / dnt-distractor
2
100%
100%
+0.0pp
terminology / term
16
0%
100%
+100.0pp
terminology / term-conflict
2
50%
100%
+50.0pp
terminology / term-distractor
2
100%
100%
+0.0pp
voice / formality
6
100%
100%
+0.0pp
voice / vocab
4
0%
100%
+100.0pp
voice / vocab-conflict
2
100%
100%
+0.0pp
gemini-3.5-flash gemini · 2× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
6
0%
100%
+100.0pp
instruction / exclamation
6
0%
100%
+100.0pp
instruction / verbatim
4
100%
100%
+0.0pp
terminology / dnt
12
100%
100%
+0.0pp
terminology / dnt-distractor
2
100%
100%
+0.0pp
terminology / term
16
0%
100%
+100.0pp
terminology / term-conflict
2
0%
100%
+100.0pp
terminology / term-distractor
2
100%
100%
+0.0pp
voice / formality
6
100%
100%
+0.0pp
voice / vocab
4
0%
100%
+100.0pp
voice / vocab-conflict
2
100%
100%
+0.0pp
Target nb measured 2026-07-17 · 17 fixtures, 26 checks (16 terminology, 2 voice, 8 instruction) · corpus 954c7c2223dd
eu.anthropic.claude-sonnet-4-6(bedrock)haiku(claude-code)opus(claude-code)sonnet(claude-code)gemini-3.1-flash-lite(gemini)gemini-3.1-pro-preview(gemini)gemini-3.5-flash(gemini)open dot = without context · filled dot = with context
Model
terminology lift
voice lift
instruction lift
overall lift
context tok/pass
lift / 1k ctx tok
steered $/pass
eu.anthropic.claude-sonnet-4-6
+50.0pp
+50.0pp
+75.0pp
+57.7pp
190
+303.6pp
$0.016
haiku
+50.0pp
+100.0pp
+75.0pp
+61.5pp
—
—
—
opus
+50.0pp
+100.0pp
+62.5pp
+57.7pp
—
—
—
sonnet
+50.0pp
+100.0pp
+75.0pp
+61.5pp
—
—
—
gemini-3.1-flash-lite
+43.8pp
+100.0pp
+75.0pp
+57.7pp
169
+341.4pp
$0.0013
gemini-3.1-pro-preview
+50.0pp
+100.0pp
+75.0pp
+61.5pp
169
+364.1pp
$0.010
gemini-3.5-flash
+50.0pp
+100.0pp
+75.0pp
+61.5pp
169
+364.1pp
$0.0067
Per-trap breakdown
Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.
eu.anthropic.claude-sonnet-4-6 bedrock · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / vocab
2
50%
100%
+50.0pp
haiku claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
opus claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
67%
+66.7pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
sonnet claude-code · 1× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
3
0%
100%
+100.0pp
instruction / exclamation
3
0%
100%
+100.0pp
instruction / verbatim
2
100%
100%
+0.0pp
terminology / dnt
6
100%
100%
+0.0pp
terminology / dnt-distractor
1
100%
100%
+0.0pp
terminology / term
8
0%
100%
+100.0pp
terminology / term-distractor
1
100%
100%
+0.0pp
voice / vocab
2
0%
100%
+100.0pp
gemini-3.1-flash-lite gemini · 2× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
6
0%
100%
+100.0pp
instruction / exclamation
6
0%
100%
+100.0pp
instruction / verbatim
4
100%
100%
+0.0pp
terminology / dnt
12
100%
100%
+0.0pp
terminology / dnt-distractor
2
100%
0%
−100.0pp
terminology / term
16
0%
100%
+100.0pp
terminology / term-distractor
2
100%
100%
+0.0pp
voice / vocab
4
0%
100%
+100.0pp
gemini-3.1-pro-preview gemini · 2× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
6
0%
100%
+100.0pp
instruction / exclamation
6
0%
100%
+100.0pp
instruction / verbatim
4
100%
100%
+0.0pp
terminology / dnt
12
100%
100%
+0.0pp
terminology / dnt-distractor
2
100%
100%
+0.0pp
terminology / term
16
0%
100%
+100.0pp
terminology / term-distractor
2
100%
100%
+0.0pp
voice / vocab
4
0%
100%
+100.0pp
gemini-3.5-flash gemini · 2× per variant
dimension / kind
checks
bare
steered
lift
instruction / digits
6
0%
100%
+100.0pp
instruction / exclamation
6
0%
100%
+100.0pp
instruction / verbatim
4
100%
100%
+0.0pp
terminology / dnt
12
100%
100%
+0.0pp
terminology / dnt-distractor
2
100%
100%
+0.0pp
terminology / term
16
0%
100%
+100.0pp
terminology / term-distractor
2
100%
100%
+0.0pp
voice / vocab
4
0%
100%
+100.0pp
Reproducing this
The harness is scripts/contexteval. A run against the built-in demo stub exercises the harness and measures nothing about any model; such runs are marked simulated and are excluded from every chart on this page. Adherence varies by target language, so the published sweep covers more than one.
make context-eval # demo stub: proves the harness, measures nothing
make context-eval-publish # the real sweep → this page's data