Gå til hovedinnhold

The Kapi format family & the .kpz package

KBF is the native serialization of one content atom — the translatable block. A project owns more than blocks, though: a content memory and a terms store, plus stand-off annotations and media. The Kapi format family gives each of those atoms a native, deterministic, lossless format, and the .kpz package bundles them into one portable artifact.

Two tiers: native vs interchange

For every atom there are two serializations, playing the same two roles KBF and XLIFF play for blocks (see KBF vs XLIFF):

AtomNative (lossless — pack, cache, hash)Interchange (lossy — industry handoff)
blocks + targetsKBF — Kapi Bundle Format, .kbf.jsonXLIFF / PO
stand-off annotationsKBF annotations.overlays.jsonl
content memorycontent-memory bundle, .memory.jsonTMX
terms storeterms bundle, .terms.jsonTBX
whole projectKPZ — Kapi Project Archive, .kpz

Every native form except .kpz is JSON, and the suffix says so. The marker segment ahead of .json keeps the file self-describing; the .json tail means jq, GitHub, and an editor's syntax highlighting all recognise it without configuration. .kpz keeps a dedicated extension because it is a binary zip, not something anyone opens by hand.

The native forms round-trip every field of the internal model. The interchange forms map only the standard, industry-portable subset — which is exactly why they cannot be used to pack a project losslessly:

  • TMX drops the entity mappings (including the conceptId cross-link from a content-memory entity to a terms concept), provenance origins and import sessions, per-entry properties, and notes.
  • TBX drops the term source (terminology vs brand_vocabulary), the competitorTerm flag, and the extensible properties.

The bundles preserve all of these. They share the KBF discipline: a kind magic string, a MAJOR.MINOR schemaVersion with the same forward-compatibility contract as KBF, and a deterministic serializer (sorted records, no HTML escaping, trailing newline) so the bytes are stable for content hashing and git diffing. The content-memory bundle reuses the same Run model as KBF blocks for its variant content, so inline codes, placeholders, and plural/select constructs serialize identically.

One bundle per store

Each store has exactly one native spelling — the bundle — and a project commits it:

kapi memory export -o l10n/tm/cli-nb.memory.json
kapi terms import seeds/terms.json # format detected from the extension
kapi terms export --format bundle -o brand.terms.json # or name it explicitly

The explicit spelling is --format bundle for both stores. It is not called json, because kapi terms --format json already means something else — a separate, lossy JSON export that does not preserve concept identity. Reach for bundle when you want the lossless form.

A bundle diffs line by line in git, which is what makes a reviewed content memory or terms store reviewable at all. There is no separate compressed spelling: a whole project that needs to travel is a .kpz, and a single bundle that needs to travel small is what zip and gzip are for.

The suffix, not the location, is what identifies a bundle — a project puts one wherever it likes and may keep several. Which bundle a given project uses is set by the recipe (defaults.termbase_source / defaults.tm_source) or by an explicit -o path.

Where terms live when nothing is bound

Your terms bundle belongs at the repository root, as terms.json. .kapi/ is gitignored, so a terms bundle kept there would never be committed, never show up in a diff, and never get reviewed, which defeats the point of keeping terminology as source. The bundle is the committed source; .kapi/termbase.db is the throwaway cache it compiles into. The terms page specifies the ladder.

Content memory has no matching convention: a project has one terms bundle but many content-memory bundles, typically one per content surface. This repository's own dogfood commits a bundle per surface under l10n/tm/ against a single l10n/terms.json, so a memory seed is always named rather than guessed.

The .kpz package

A .kpz is a deterministic zip bundling a project's authoritative content, one member per content type — the same content-type set the sync protocol moves over the wire:

project.kpz
├── manifest.json inventory: per-member sha256 + Merkle rootHash
├── blocks/<id>.kbf.json KBF (blocks + targets)
├── annotations/<id>.overlays.jsonl KBF annotations (stand-off overlays)
├── memory.json content memory
├── terms.json terminology
├── media/<name> opaque blobs

│ ── working-state members (AD-025 §5) ──
├── source/<name> ingested source document(s)
├── overlays.json in-progress overlays (targets, annotations, …)
└── history.jsonl advisory provenance log (opt-in; excluded from rootHash)

The memory and terms members use the conventional bare names, not the compound suffix, so unzipping a package by hand gives you the spelling the rest of the tooling teaches. Member names are cosmetic: unpacking routes each member by the manifest's contentType, never by its filename.

(A .kpz also carries the project recipe in manifest.json — flows, plugins, defaults, content — kept out of the rootHash, so the package is a runnable project in a file. Side-effecting recipe like server:/hooks: travels inert.)

The same container serves two profiles, set by the manifest kind:

  • kind: kapi-project — the whole project (all locales, full recipe, content memory, terms, overlays, source identity + skeletons). The snapshot / transport parcel, moved by pack / unpack.
  • kind: kapi-interchange — a task-scoped bilingual slice for one locale pair (blocks, inline codes, segmentation, skeleton, memory and term context). This is neokapi's interchange format, sent to a translator/reviewer by extract and ingested by merge (see KBF vs XLIFF).

Both are parcels, not workspaces: day-to-day work happens in the ambient .kapi project (AD-025 §7).

Because every member is a native lossless format, the whole package is lossless: unpacking can seed a fresh content memory, terms store, and block store, and the project's regenerable caches (the block-store cache, sync hashes) rebuild faithfully from it. The manifest's per-member SHA-256 and Merkle rootHash give the package — and each member — a stable content identity; unpacking verifies both before trusting the contents.

What a package deliberately excludes

A .kpz packs the source of truth, not the regenerable caches. It excludes the block-store cache (blocks.db) and the sync hash cache, which rebuild on demand, and it excludes secrets (the sync claim token). This is what makes the package the at-rest twin of the sync wire format: packing is the sync converters writing files instead of protobuf, and unpacking is the inverse.

Raw source bytes are also excluded by default: a .kpz carries each source's identity (path, format, content hash) + round-trip skeleton — enough to merge — but not the originals, so it doesn't duplicate git-tracked source. Pass --with-source to kapi extract / kapi pack to embed the raw bytes too (needed to re-extract offline). And kapi pack refuses to write a content-less .kpz (a project with nothing extracted/translated yet) — like git bundle refusing an empty bundle; share the kapi.yaml recipe via git instead.

Working state: hand-off and resume

A .kpz is not only an at-rest snapshot of finished content — it can also carry in-progress working state, so work can stop, move between machines, and resume where it left off. There are two routes, and neither needs per-step CLI verbs:

  • .kpz as an ad-hoc workspace (extract / transform / merge). No project required. The .kpz is a portable bundle; the runtime is a persistent shadow cache keyed to the file, so transforms are fast and incremental and the .kpz is rewritten only when you pack (or pass --pack). extract ingests sources (+ a recipe), running a tool/flow on the .kpz transforms it in place, and merge emits the finished files. info shows whether the cache is dirty.

    kapi extract src/*.json -o work.kpz --target-lang fr,qps # ingest
    kapi translate work.kpz --target-lang fr # transform (cache)
    kapi info work.kpz # dirty?
    kapi pack work.kpz # eject for sharing
    kapi merge work.kpz -o l10n/ # emit → l10n/<lang>/<name>

    A small recipe (target locales + output layout) travels with the file, so merge needs no flags. Transforming reuses work already done instead of recomputing it, and each document's work stays isolated.

  • Project snapshot (pack / unpack). Inside a .kapi project the working state also includes the project's content memory and terms store. kapi pack snapshots all of it; kapi unpack rehydrates it into another machine's .kapi/ state dir. Within a project, re-running a flow also reuses the project's persistent cache (.kapi/cache/blocks.db), so resume there is simply running again.

    kapi pack -o snapshot.kpz # snapshot the project's working state
    kapi unpack snapshot.kpz # rehydrate elsewhere, then run to resume

The working-state members are source/<name> (the ingested documents) and overlays.json (the in-progress overlays — targets/<locale>, annotations/<name>, segmentation, … keyed by (kind, blockHash)). Overlays in a multi-document package are scoped per source, since block ids are only unique within one document. A recipe block in the manifest (not a content member, so out of the rootHash) carries the workspace's target locales and output layout.

Progress is derived from content, not a journal

Because the block store is append-only and content-addressed, "has step X run?" is a pure function of the content: does X's overlay exist for the current block hashes? That is what makes cached resume correct — re-running is a no-op where the overlay is already present, and a source change re-hashes its block so only the affected work recomputes. There is no authoritative progress journal — it would be a second source of truth that could drift from the content (the dual-state footgun this codebase avoids).

The one optional log is advisory provenance: kapi pack --log stamps a hash-chained line into history.jsonl recording the pack (tamper-evident custody for hand-off; unpack verifies the chain and warns if broken). It is strictly subordinate to content — excluded from the package rootHash, never read to decide anything, and safe to delete with no loss of work. A default pack (no --log) is byte-deterministic.

When to use it

Use a .kpz to move a whole project losslessly: backup and archival, seeding a fresh server or an offline desktop working copy, transferring a project between machines without a server, or building a deterministic test fixture. Use the interchange formats (XLIFF/PO, TMX, TBX) instead when content crosses into the wider translation industry — a translation vendor, an external CAT tool, or a TMS — and the lossy, standards-based subset is what the other side expects.

See also