Gå til hovedinnhold

Skill eval

An agent is handed a prompt and a workspace. Every row expands into the exact prompt, the files the agent saw, the tools and commands it reached for, and a diff of what it changed.

21 pass
1 flaky
0 fail
1 false triggers
24d since this run
run 2026-08-27T22:34:43Z · 3 passes per scenario agent 2.1.228 (Claude Code) · the claude CLI's default, not pinned skill 4c682ac83 · last edited 2026-08-15 kapi bin/kapi · kapi v1.2.0-rc29-9-g4b8194cf7 (beta, commit: 4b8194cf7, built: 2026-08-27T22:02:21Z) host Darwin arm64 settings claude -p, bypassPermissions, 3 repeat(s), trigger cap 4 turns (MCP scenarios use their own, since picking a tool can take a step or two). Sampling follows the CLI's defaults and is not pinned.

The agent has to notice kapi is relevant. The only lever is the skill's description, and a false trigger costs more than a miss: missing a positive wastes one kapi user's prompt, firing on a code task wastes everyone's.

What kapi added

Each scenario ran a second time with no skill, no MCP server, and no kapi on PATH. The comparison is deliberately conservative: kapi failing is never counted as a win, the message counts are medians, and a scenario with no gate is unknown rather than assumed.

What it cost. Across 22 scenarios the agent sent 197 messages with kapi and 164 without, and the unaided arm was shorter on 15 of them. That is the honest counterweight to the counts below: on this suite kapi reaches answers the unaided agent cannot, and it is not the cheaper route to the ones it can.

0
kapi enabled it
the unaided agent could not finish and the one with kapi did
0
kapi eased it
both finished, and kapi took materially fewer messages
0
kapi hindered it
the unaided agent finished and the one with kapi did not
17
no difference
both finished, and kapi saved nothing measurable
5
not comparable
no gate, so there is no outcome to compare

Must fire (17)

Must stay quiet (5)