E-08: Document structure tiers
Summary
Some formats state their own logical structure; some only imply it. Where a format carries headings, paragraphs, and tables as markup, the reader reads them. Where it does not (PDF is the archetype), structure is recovered through a tiered model, in decreasing order of authority: the document's own tagged structure tree where it exists, geometric inference from positions otherwise, and an ML layout model above both for what geometry cannot reach. Each tier falls through cleanly to the one below, so a document always parses.
PDF is read out-of-core by PDFium, through two backends that share their content-producing logic:
- Native:
kapi-pdfium, a first-party Mode-C plugin (E-05) linking PDFium via cgo and running as an isolated daemon, so a malformed-PDF crash dies with the subprocess rather than withkapi. - Browser: a
WasmReader(build-taggedjs) that bridges Go-WASM to a PDFium WebAssembly module loaded by the web app, giving the in-browser Lab the same extraction without a server.
Both emit the same Part stream: positioned text Blocks carrying a
GeometryAnnotation (bounding box, optionally per-glyph boxes), grouped into
lines by a shared algorithm, with structure recovered through the tiers. PDFium
is the only PDF path on every platform; there is no hand-rolled pure-Go reader
beside it.
Context
A faithful PDF reader needs more than byte extraction. PDFs encode text with font-program glyph indices (CID/Type0, common for CJK) that a naive scanner garbles; correct extraction requires a real font and encoding engine. PDFium is that engine, but it is a large C++ codebase with two consequences the framework must contain:
- Native weight. Linking PDFium via cgo defeats pure-Go cross-compilation and
inflates every
kapiinstall for a capability many invocations never use: the same isolation rationale as the Okapi bridge and the ML segmenter (M-02). PDF therefore lives in a plugin, not in the framework. - Crash surface. Malformed PDFs are a classic parser-crash vector. Running the reader in a subprocess daemon contains a crash to that process.
Beyond text, the visual editor (S-06) and the browser Lab need each text fragment's position on the page, and downstream processing wants document structure (which lines are headings, which blocks form a table), so that a translated document can be reflowed and so that table cells are translated as cells. PDF carries none of this uniformly: some PDFs are tagged with an explicit logical structure tree; most are not, and structure must be inferred from geometry.
Decision
Two backends, one content contract
PDF reading is registered through a build split, so the right backend is wired per target and the framework core stays free of PDFium:
core/formats/register_pdf_js.go(//go:build js) registers the browserWasmReader. It is read-only: there is no writer, so editing tools fail cleanly rather than overwriting the document with extracted text.core/formats/register_pdf_other.go(//go:build !js) is a no-op. Native builds have no in-core PDF reader, so thepdfformat is supplied only by the installedkapi-pdfiumplugin at runtime, or is absent with an actionable "install the plugin" error.
The native plugin is its own Go module,
github.com/neokapi/neokapi/plugins/pdfium, outside the workspace so its cgo and
PDFium dependencies never enter another module's build graph:
plugins/pdfium/
├── go.mod module …/plugins/pdfium
├── manifest.json Mode-C daemon manifest (declares the pdf format + schema)
├── formats/pdf/schema.json config schema: geometry, glyphs, tier3
├── cmd/kapi-pdfium/main.go daemon entry point (gRPC, Mode C)
└── internal/pdfreader/
├── reader.go ReadParts: fast path + geometry path
└── structtree.go tier-1 tagged struct tree (build-tag-free; runtime-gated)
Both backends produce an identical Part stream (a document Layer, per-page
Layers, and Blocks) and share two pieces of framework logic so the native and
browser results match:
core/formats/pdf.GroupRunsmerges PDFium's per-rect fragments (it emits a fresh rect at every mid-line font or style change) into line-level runs, splitting only at large horizontal gaps, which mark a column or cell boundary. Without it a styled or isolated glyph becomes a one-character Block.core/structure.Analyze/core/structure.ToPartsinfer tier-2 structure.
Two extraction modes
The reader has two granularities, selected by the geometry config flag
(formats/pdf/schema.json):
- Fast path (
geometry=false, the default): one plain-text Block per page. Fewest allocations, no positional work; the right choice for batch text operations such as word counts, memory leverage, and the toolbox utilities (S-04). - Geometry path (
geometry=true): one Block per positioned text run, each carrying aGeometryAnnotation. Withglyphs=true(which implies geometry) each Block additionally carries per-character boxes for character-precise highlighting.
Independent of the mode, the reader maps the PDF Info dictionary onto the
content model through the shared core/docmeta helper: Title, Subject, and
Keywords become translatable Blocks on the metadata plane, while Author,
Creator, Producer, CreationDate, and ModDate are recorded as
pdf:-namespaced properties on the document layer, never translated, kept for
inspection. This is the same metadata mechanism the image reader uses
(M-07).
Geometry model and coordinate flip
PDFium reports boxes in PDF user space, where the origin is bottom-left and Y
increases upward. neokapi's content and editor model use a top-left origin (Y
increases downward), matching screen and most document coordinate systems. Every
box is flipped once, at extraction, when the page height is known:
Y_top-left = pageHeight − Y_upper. The flip is implemented identically at all
three sites (the native fast and geometry reader, the native tagged-tree reader,
and the browser bridge), and each stamps Origin: "top-left", or falls back to
"bottom-left" when the page height is unavailable. GeometryAnnotation carries
the page number, the union bounding box, the origin, and the optional per-glyph
boxes (text plus box per character).
Structure tiers
Document structure is recovered through three tiers, in decreasing order of authority:
| Tier | Source | Where it runs | Authority |
|---|---|---|---|
| 1: Tagged struct tree | The document's own logical structure tree (Document › H1 › P › Table › TR › TH/TD …) | Native plugin only | Authoritative: the author's own tags |
| 2: Geometric inference | Block positions: row clustering, column alignment, relative line height | Native and browser | Heuristic |
| 3: ML layout | A vision model over the rendered page | Native plugin + host | Heuristic, highest recall |
Tier 1 (internal/pdfreader/structtree.go) reads a tagged PDF's structure
tree directly: it builds a marked-content-ID → text and bounds map from the
page's text objects, walks the struct tree, and maps elements onto the content
model: Table becomes a table group of rows and cells (TH → header, TD → cell),
H1–H6 become headings, P and its relatives become paragraphs. The
marked-content accessors are PDFium experimental APIs that the Go binding wires
only under the pdfium_experimental build tag, and the bundled library exports
the symbols, so the shipped plugin is built with that tag. The code itself
carries no build tag: it compiles either way and, when the experimental API
is unavailable, returns "no tagged structure" so the reader falls through to tier
2. Untagged documents fall through the same way.
Tier 2 (core/structure.Analyze) is format-agnostic: it consumes Blocks
carrying geometry, clusters them into rows, detects tables where cells align into
stable columns across consecutive rows, and tags remaining prose as heading or
paragraph by relative line height. ToParts emits the result as the same
table/heading/paragraph Part stream a structure-aware reader produces
(F-02), so the existing Markdown and HTML
writers render real tables with no PDF-specific code. Both backends run tier 2 as
their fallback, and because it operates on geometry rather than on PDF, any
format that can produce positioned blocks gets it.
Tier 3 runs a layout model over the rendered page for the cases geometry
cannot reach: borderless tables, multi-column reading order, scanned pages,
figure and caption association. It is not part of the PDF format:
it operates on a page raster plus blocks, and so applies to any format that can
produce them (M-03). Enabled with
the tier3 option, the plugin renders each page to a PNG at 72 DPI (so PDF
points map 1:1 to raster pixels and the text geometry aligns) and emits it as a
Media part marked vision.PageRasterProperty alongside the page's raw positioned
blocks, applying no structure of its own. A host-side reader decorator
(host/pluginhost/tier3_reader.go, wrapping every plugin reader as a pass-through
until tier3 is requested) runs the vision layout model over the raster via
vision.StructureFromLayout: the model's regions and reading order are
authoritative, the blocks' own text fills them
(vision.OCRResultFromBlocks, more accurate than re-running OCR over a vector
PDF), and table regions are reconstructed into row and column cells. The
decorator deletes the rendered raster afterwards and, when the vision plugin is
not installed, strips the request so the plugin falls back to the geometric tier
2, so tier 3 degrades cleanly to tier 2, never worse. Browser builds have no
native vision engine, so tier 3 is native-only.
A consequence of this design is a native/browser asymmetry: the browser
extraction contract exposes only rects, not the struct tree, so the in-browser Lab
always uses tier 2 even for a tagged PDF, while a native kapi run uses tier 1.
The structure of a tagged PDF can therefore differ between the Lab and the CLI;
this is documented at the call site.
Structure from an external converter
A document converter that has already recovered layout can hand its result to
the framework instead of the raw file. Two readers take that route: docling
reads DoclingDocument JSON, the lossless serialization Docling emits, and
doclang reads DocLang XML. Each walks the converter's reading order and maps
its items onto the same content model the tiers above produce: a normalized
role on the structure annotation, the provenance box on a
GeometryAnnotation, page headers and footers on the furniture layout layer,
tables as a group of row groups of cell blocks, captions as caption-role blocks.
Both readers are read-only; the converter owns its file, and the framework
re-emits structure through the DocLang writer or projects it to Markdown or
HTML. This is the file-based way to consume a converter: run it, save its
output, read the output.
Distribution
kapi-pdfium bundles the PDFium shared library at lib/<name> beside the
binary in the release tarball, found via an rpath baked into the binary
(@loader_path/lib on macOS, $ORIGIN/lib on Linux, same directory on Windows):
the same bundling shape as the ML segmenter plugin's ONNX runtime
(M-02) rather than a static link. Tarballs are
built per platform on native runners (cgo, -tags pdfium_experimental, PDFium
pinned to a release build), cosign-signed, and indexed in the registry, so
kapi plugin install pdfium verifies the download like any other plugin
(E-05). The CLI's Homebrew formula depends on the
plugin's, so a brew install brings it along; any other install adds it with
kapi plugin install pdfium; the desktop app installs it on demand the first
time a PDF is opened; and the browser self-hosts PDFium's WASM next to the kapi
WASM.
Consequences
- The portable
kapibinary stays pure-Go, small, and cross-compilable; the PDFium native stack is confined to a separately built, separately installed plugin, and a malformed PDF can crash only the daemon. - Native and browser produce the same positioned-text
Partstream because they shareGroupRunsandstructure.Analyze/ToParts; the only intended divergence is tier-1 structure, which the browser cannot see. - Geometry is available for the visual editor and the Lab without changing the
content model: it rides on the stand-off
GeometryAnnotation, off by default so batch text work pays nothing for it. - Tagged documents get author-fidelity structure for free; untagged ones get a best-effort geometric reconstruction; the ML tier sits above both without reworking either backend, and degrades to tier 2 when its plugin is absent.
- Because tier 2 and tier 3 consume geometry and rasters rather than PDF, the same structure recovery is available to any format that can produce them.
Related
- F-02: The content model: Blocks, stand-off annotations, table groups, and the structure stream
ToPartsemits - E-02: The format system: how the
pdfformat registers and detects - E-05: The plugin system: Mode-C daemon discovery, native-stack isolation, signed registry distribution
- M-02: Segmentation: the precedent for isolating a native stack in a bundled-shared-library plugin
- M-03: Multimodal content: the vision layout model tier 3 calls, and the raster contract it runs on
- S-06: The visual editor: how the editor consumes geometry and structure
plugins/pdfium/: plugin module, readers, and README