Gå til hovedinnhold

Context eval

kapi steers model output by injecting context: a terms store that mandates renderings, a voice guide, an instruction. This page measures whether each model actually follows that context — not whether it translates well, which is a different question. The sibling batch eval measures structural integrity and cost; this one measures obedience.

How it is measured

The core metric is a differential. An engineered corpus is translated twice through the production pipeline — once bare (no context) and once steered (terms + voice profile + instruction, injected exactly as production injects them) — and both passes are scored with kapi’s own deterministic check tools (term-check, dnt-check, voice-vocab-check, pattern-check). Two numbers fall out per dimension: absolute adherence (did the steered output satisfy the requirement) and lift (steered minus bare — how much the context moved the model). Lift is the decision-relevant one: a model with high absolute adherence but no lift already “knew” it, and our context earns no credit. A model with high lift is genuinely steerable, which is what context injection is buying.

Every fixture is a trap: a naive translation violates the context. A mandated term whose natural rendering differs from the mandate, a product name that reads like a common noun, casual English tempting an informal register, a source that ends in the exclamation mark the instruction forbids. There are distractors (a lowercase “compass” that is a real compass and must be translated) and declared-winner conflicts (the terms store pins a compound containing a forbidden word — the pin wins, and the scorer knows it). Results are reported per dimension — terminology, voice, instruction — never as one collapsed score: a model can be excellent at terminology and poor at voice, and the collapsed number would hide the thing you would act on.

The deterministic checks are the backbone. The genuinely subjective remainder of voice — register, naturalness, restraint — is scored by a cross-family LLM judge under a fixed yes/no rubric, blind to which model and which variant produced a text, and its scores are published only once judge–human agreement has been measured above a stated bar. Until then, judged numbers stay in the record but off this page.

What the measurement found

The most steerable model measured is haiku ( nb ), which the context moved by +73.1pp overall.Every measured model shows positive lift — the context earns its tokens everywhere.The weakest dimension with context applied is instruction at 97.6% steered adherence — the ceiling context injection currently hits.

Target de measured 2026-08-27 · 20 fixtures, 32 checks (17 terminology, 7 voice, 8 instruction) · corpus 66cc1693369f

0%25%50%75%100%terminologyhaikusonnetvoicehaikusonnetinstructionhaikusonnet
haiku (claude-code) sonnet (claude-code) open dot = without context · filled dot = with context
Modelterminology liftvoice liftinstruction liftoverall liftcontext tok/passlift / 1k ctx toksteered $/pass
haiku+41.2pp+42.9pp+62.5pp+46.9pp
sonnet+47.1pp+42.9pp+75.0pp+53.1pp

Per-trap breakdown

Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.

haiku claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%67%+66.7pp
instruction / exclamation30%100%+100.0pp
instruction / verbatim2100%100%+0.0pp
terminology / dnt6100%100%+0.0pp
terminology / dnt-distractor1100%100%+0.0pp
terminology / term813%100%+87.5pp
terminology / term-conflict1100%100%+0.0pp
terminology / term-distractor1100%100%+0.0pp
voice / formality3100%100%+0.0pp
voice / vocab333%100%+66.7pp
voice / vocab-conflict10%100%+100.0pp

sonnet claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / verbatim2100%100%+0.0pp
terminology / dnt6100%100%+0.0pp
terminology / dnt-distractor1100%100%+0.0pp
terminology / term80%100%+100.0pp
terminology / term-conflict1100%100%+0.0pp
terminology / term-distractor1100%100%+0.0pp
voice / formality3100%100%+0.0pp
voice / vocab333%100%+66.7pp
voice / vocab-conflict10%100%+100.0pp

Target en-GB measured 2026-07-17 · 16 fixtures, 20 checks (3 terminology, 5 voice, 12 instruction) · corpus 600fac47f9dc

0%25%50%75%100%terminologyeu.anthropic.claude-sonnet-4-6haikuopussonnetgemini-3.1-flash-litegemini-3.1-pro-previewgemini-3.5-flashvoiceeu.anthropic.claude-sonnet-4-6haikuopussonnetgemini-3.1-flash-litegemini-3.1-pro-previewgemini-3.5-flashinstructioneu.anthropic.claude-sonnet-4-6haikuopussonnetgemini-3.1-flash-litegemini-3.1-pro-previewgemini-3.5-flash
eu.anthropic.claude-sonnet-4-6 (bedrock) haiku (claude-code) opus (claude-code) sonnet (claude-code) gemini-3.1-flash-lite (gemini) gemini-3.1-pro-preview (gemini) gemini-3.5-flash (gemini) open dot = without context · filled dot = with context
Modelterminology liftvoice liftinstruction liftoverall liftcontext tok/passlift / 1k ctx toksteered $/pass
eu.anthropic.claude-sonnet-4-6+100.0pp+100.0pp+50.0pp+70.0pp150+466.7pp$0.014
haiku+100.0pp+100.0pp+41.7pp+65.0pp
opus+100.0pp+100.0pp+50.0pp+70.0pp
sonnet+100.0pp+100.0pp+50.0pp+70.0pp
gemini-3.1-flash-lite+100.0pp+100.0pp+50.0pp+70.0pp137+510.9pp$0.0011
gemini-3.1-pro-preview+100.0pp+100.0pp+50.0pp+70.0pp137+510.9pp$0.0076
gemini-3.5-flash+100.0pp+100.0pp+50.0pp+70.0pp137+510.9pp$0.0047

Per-trap breakdown

Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.

eu.anthropic.claude-sonnet-4-6 bedrock · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / spelling4100%100%+0.0pp
instruction / verbatim2100%100%+0.0pp
terminology / term30%100%+100.0pp
voice / contractions40%100%+100.0pp
voice / vocab10%100%+100.0pp

haiku claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%67%+66.7pp
instruction / exclamation30%100%+100.0pp
instruction / spelling4100%100%+0.0pp
instruction / verbatim2100%100%+0.0pp
terminology / term30%100%+100.0pp
voice / contractions40%100%+100.0pp
voice / vocab10%100%+100.0pp

opus claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / spelling4100%100%+0.0pp
instruction / verbatim2100%100%+0.0pp
terminology / term30%100%+100.0pp
voice / contractions40%100%+100.0pp
voice / vocab10%100%+100.0pp

sonnet claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / spelling4100%100%+0.0pp
instruction / verbatim2100%100%+0.0pp
terminology / term30%100%+100.0pp
voice / contractions40%100%+100.0pp
voice / vocab10%100%+100.0pp

gemini-3.1-flash-lite gemini · 2× per variant

dimension / kindchecksbaresteeredlift
instruction / digits60%100%+100.0pp
instruction / exclamation60%100%+100.0pp
instruction / spelling8100%100%+0.0pp
instruction / verbatim4100%100%+0.0pp
terminology / term60%100%+100.0pp
voice / contractions80%100%+100.0pp
voice / vocab20%100%+100.0pp

gemini-3.1-pro-preview gemini · 2× per variant

dimension / kindchecksbaresteeredlift
instruction / digits60%100%+100.0pp
instruction / exclamation60%100%+100.0pp
instruction / spelling8100%100%+0.0pp
instruction / verbatim4100%100%+0.0pp
terminology / term60%100%+100.0pp
voice / contractions80%100%+100.0pp
voice / vocab20%100%+100.0pp

gemini-3.5-flash gemini · 2× per variant

dimension / kindchecksbaresteeredlift
instruction / digits60%100%+100.0pp
instruction / exclamation60%100%+100.0pp
instruction / spelling8100%100%+0.0pp
instruction / verbatim4100%100%+0.0pp
terminology / term60%100%+100.0pp
voice / contractions80%100%+100.0pp
voice / vocab20%100%+100.0pp

Target fr measured 2026-08-27 · 19 fixtures, 31 checks (17 terminology, 6 voice, 8 instruction) · corpus 7d42625fd539

0%25%50%75%100%terminologyhaikusonnetvoicehaikusonnetinstructionhaikusonnet
haiku (claude-code) sonnet (claude-code) open dot = without context · filled dot = with context
Modelterminology liftvoice liftinstruction liftoverall liftcontext tok/passlift / 1k ctx toksteered $/pass
haiku+52.9pp+33.3pp+75.0pp+54.8pp
sonnet+52.9pp+16.7pp+75.0pp+51.6pp

Per-trap breakdown

Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.

haiku claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / verbatim2100%100%+0.0pp
terminology / dnt6100%100%+0.0pp
terminology / dnt-distractor1100%100%+0.0pp
terminology / term80%100%+100.0pp
terminology / term-conflict10%100%+100.0pp
terminology / term-distractor1100%100%+0.0pp
voice / formality3100%100%+0.0pp
voice / vocab20%100%+100.0pp
voice / vocab-conflict1100%100%+0.0pp

sonnet claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / verbatim2100%100%+0.0pp
terminology / dnt6100%100%+0.0pp
terminology / dnt-distractor1100%100%+0.0pp
terminology / term80%100%+100.0pp
terminology / term-conflict10%100%+100.0pp
terminology / term-distractor1100%100%+0.0pp
voice / formality3100%100%+0.0pp
voice / vocab250%100%+50.0pp
voice / vocab-conflict1100%100%+0.0pp

Target nb measured 2026-08-27 · 17 fixtures, 26 checks (16 terminology, 2 voice, 8 instruction) · corpus daf7de0cd29f

0%25%50%75%100%terminologyhaikusonnetvoicehaikusonnetinstructionhaikusonnet
haiku (claude-code) sonnet (claude-code) open dot = without context · filled dot = with context
Modelterminology liftvoice liftinstruction liftoverall liftcontext tok/passlift / 1k ctx toksteered $/pass
haiku+68.8pp+100.0pp+75.0pp+73.1pp
sonnet+50.0pp+100.0pp+50.0pp+53.8pp

Per-trap breakdown

Adherence with context, by the kind of trap the fixture set. A model can hold the terms and still translate the product name; this is where that shows.

haiku claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%100%+100.0pp
instruction / exclamation30%100%+100.0pp
instruction / verbatim2100%100%+0.0pp
terminology / dnt650%100%+50.0pp
terminology / dnt-distractor1100%100%+0.0pp
terminology / term80%100%+100.0pp
terminology / term-distractor1100%100%+0.0pp
voice / vocab20%100%+100.0pp

sonnet claude-code · 1× per variant

dimension / kindchecksbaresteeredlift
instruction / digits30%33%+33.3pp
instruction / exclamation30%100%+100.0pp
instruction / verbatim2100%100%+0.0pp
terminology / dnt6100%100%+0.0pp
terminology / dnt-distractor1100%100%+0.0pp
terminology / term80%100%+100.0pp
terminology / term-distractor1100%100%+0.0pp
voice / vocab20%100%+100.0pp

Reproducing this

The harness is scripts/contexteval. A run against the built-in demo stub exercises the harness and measures nothing about any model; such runs are marked simulated and are excluded from every chart on this page. Adherence varies by target language, so the published sweep covers more than one.

make context-eval                    # demo stub: proves the harness, measures nothing
make context-eval-publish            # the real sweep → this page's data