Skip to content

PDF pipeline and API

This PDF pipeline page is for contributors changing prodockit.pdf or calling its Python API directly. Document authors should use Generate a PDF.

Follow the pipeline

The PDF build starts only after Zensical has produced a complete website. Prodockit reads those generated articles in navigation order, removes website-only structure, and assembles one HTML document. Pandoc and the Lua filter then prepare that document for WeasyPrint, which lays out the pages; the optional final step extracts the index before the PDF is written.

Figure 49.1 shows this sequence across the top row and then the bottom row. Each box names either the form of the document at that point or the component responsible for the next change.

Pipeline from the Zensical project through generated HTML, Pandoc and WeasyPrint to the final PDF

1. PDF generation pipeline

The public prodockit pdf command validates the completed Zensical site, reads each navigation page's generated article, and constructs Page objects from that output. It inspects those objects for active Mermaid and arithmatex elements before discovering or constructing either optional renderer, then pre-renders only the content that is present and calls the lower-level builder. A generated index adds a second layout pass after term pages are known.

The command never invokes Zensical or cleans the configured site_dir. Passing --markdown-file narrows only the pages assembled into the PDF; the requested article must already exist in the completed site. This keeps a single-page PDF quick without allowing it to conceal an incomplete website build.

Run zensical build --clean --strict before pdk pdf. The completed site is the single supported input; Prodockit no longer contains a second renderer using undocumented Zensical Python interfaces.

Use the public Python surface

Choose the public entry point that matches the caller's available input from Table 49.1.

1. Use the public Python surface

API Purpose
build_pdf_from_built_site() High-level build using navigation, settings, and the completed Zensical site
build_pdf() Lower-level build from prepared Page objects
Page One rendered source page plus its path, appendix, index, and running-header metadata
PdfBuildError Build failure carrying the underlying command output
build_source_bundle_from_zensical_config() High-level Markdown/configuration source-bundle build

Of the entry points in Table 49.1, prefer build_pdf_from_built_site() when a caller already has a Zensical project. Use build_pdf() only when the caller owns page rendering and can supply complete HTML and metadata.

from prodockit.pdf.config import build_pdf_from_built_site

output = build_pdf_from_built_site("zensical.toml")
print(output)

The CLI wraps these functions with progress reporting, captured diagnostics, and non-zero exit status; the functions return paths or raise exceptions.

Know the internal modules

Table 49.2 maps each internal module to the transformation it owns.

2. Know the internal modules

Module Responsibility
prodockit.project_config Direct TOML/YAML reading for the settings Prodockit consumes
prodockit.pdf.site Completed-site validation and generated-page extraction
prodockit.pdf.config Navigation, metadata, optional renderers, and high-level entry points
prodockit.pdf.build Pipeline orchestration and external-command execution
prodockit.pdf.html Page fix-ups, front matter, web/PDF-only content, and heading structure
prodockit.pdf.lua Pandoc Lua filter generation
prodockit.pdf.css Renderer foundations, dynamic page settings, running headers/footers, and duplex layout
docs/stylesheets/pdk-pdf.css Managed PDF presentation defaults that authors can override in print.css
prodockit.pdf.icons Project icon discovery and SVG recovery from built CSS
prodockit.pdf.mermaid Isolated Python Mermaid rendering and diagram assets
prodockit.pdf.source_bundle Markdown/configuration source PDF
prodockit.pdf.index Marker extraction, term-page mapping, and generated index
prodockit.pdf.release Host release lookup for cover markers
prodockit.pdf.runtime_config Strict project-root pdk-pdf.toml policy and supported defaults
prodockit.pdf.runtime_store Project-local locks, archive validation, smoke tests, atomic activation, and fallback state
prodockit.pdf.runtime_prepare Provider-gated pdk pdf --prepare orchestration
prodockit.pdf.weasyprint_runtime Official Windows x64 artifact pin, bounded download, absolute-path probe, and cache executable resolution

Use Table 49.2 to locate the owner of a transformation before changing it. Keep transformation stages narrow. A change to HTML normalisation, CSS, the Lua filter, or an external tool can affect every page, so verify the complete PDF and built-output tests rather than relying on a unit test of the changed module alone.

Preserve source-bundle boundaries

The prodockit source-bundle command includes root README.md, Markdown below docs_dir, and the active Zensical config. Root-level generated Markdown such as changelog, contribution, and licence files stays outside that boundary. It is intentionally separate from the rendered-document pipeline: it discovers files with git, writes a self-contained HTML document, and calls WeasyPrint without Pandoc. The lower-level source-bundle API can discover every non-ignored text file for a specialised caller, but that is not the command-line default.

Copyright text can contain links and line breaks, so it cannot be flattened into a CSS string. The PDF pipeline writes it as an HTML element and places it in the repeated footer with CSS Paged Media's position: running() and content: element(). Check the finished PDF when changing this path; intermediate HTML alone does not prove that links or line breaks survived.

Preserve the runtime trust boundary

pdk-pdf.toml contains committed intent only. Machine-specific resolved state belongs below .prodockit/cache/pdf/, rooted beside the selected Zensical configuration so separate projects and virtual environments never share an active runtime accidentally.

Providers may resolve an approved artifact and copy it into a requested staging path. They do not choose its activation path or extract it themselves. The common store verifies the exact SHA-256, rejects traversal, links, duplicate paths, encrypted ZIP members, special TAR members, excessive entry counts and excessive expansion, checks required content, then runs the provider's fresh-runtime probe. Only that validated staging directory can be atomically activated. A failure retains and reports the current or previous known-good runtime.

The first G1 dependency baseline was measured on the development macOS ARM64 environment on 19 September 2026. These are installed-file sizes rather than download sizes and are evidence for later optionalisation, not release budgets. Table 49.3 records that starting point.

3. G1 base dependency baseline

Base distribution Resolved version Installed files Installed bytes
pypdf 6.18.0 124 3,990,895
mermaidx 0.9.5 48 5,364,347
quickjs-ng 0.16.2.1 9 1,249,527

Importing prodockit.cli did not import any of those three distributions; the measured wall time was 0.36 seconds on that host. Repeat the installed size and cold-import measurement in G5 before removing the PDF-only packages from the base dependency set.

The G2 lazy boundary uses the same HTML parser as the lower-level missing- renderer warnings. On the 60-page contributor guide built on macOS ARM64, a 20-run sample detected the required renderers in a median 84.39 ms (maximum 89.79 ms). The parser is platform-independent Python and is covered on the supported macOS, Ubuntu, and Windows test matrix. Plain-page configuration tests replace Mermaid and MathJax discovery and construction with hard failures, proving that unused components perform no probe, worker, directory, install, or network work; active Mermaid and maths fixtures prove the existing callbacks and output inputs remain selected when required. G2 changed no default renderer and made the external Mermaid fallback independently deletable; that fallback and its configuration were subsequently removed.

Preserve actionable errors

PdfBuildError reports which external command failed and retains its captured output. BuiltSiteError identifies a failed Zensical command, a missing generated article, or a generated HTML layout that no longer exposes the expected article. The legacy wrapper separately names a changed private Zensical render-result shape. Do not replace these with an unlabelled subprocess status or raw selector failure; callers need to know which boundary changed.

MermaidRenderError identifies the failed diagram number and a bounded failure category without echoing diagram source. It stops the supported PDF pipeline before Pandoc can replace the requested PDF or the result can be copied into the built site. MermaidBackendUnavailableError remains separate because a missing or corrupt prepared runtime needs different corrective action.