10  Performance Model

10.1 Reproducible measurements

Run the benchmark harness from the repository root:

python3 tests/benchmark.py --runs 7 --output /tmp/typograph-benchmarks.json

It measures fresh compiler processes with bundled fonts, one warmup per workload, median/min/max wall time, and median peak resident memory on macOS or Linux. Times include startup, compilation, rendering, and PDF export; filesystem caches are not cleared. Use --no-memory when the environment does not permit the platform’s memory probe. A 60-second per-compile timeout prevents runaway measurements; it is configurable with --timeout.

To compare two source trees, copy the same fixture files into both trees and alternate their measurements in one run:

python3 tests/benchmark.py --root before=/path/to/before --root after=. --runs 7

The three existing stress fixtures cover a 400-node/760-edge grid, 256 32-sided polygons with 480 clipped edges, and 100 clipped/labelled cubics. tests/benchmark.typ adds 144 gates with 576 port edges, a 240-node named chain with 160 followers, a 101-node captured chain, and 100 rotated groups with mixed-axis placement. Select a workload with --case captured-chain.

The correctness suite compiles the stress fixtures, without a timing threshold. Successful compilation does not prove the absence of a performance regression. Compare repeated measurements on the same machine and compiler version; small differences within the observed ranges are noise.

10.2 Recorded review measurements

On macOS 26.5.2 ARM64 with Typst 0.15.1, seven fresh-process samples per revision gave these medians for the 2026-08-28 simplification review:

Workload Before (s) After (s) Before RSS (MiB) After RSS (MiB)
grid 0.282 0.281 67.8 67.6
polygons 0.309 0.309 52.4 51.8
curves 0.203 0.201 62.6 62.8
ports 0.205 0.204 67.3 67.4
named-chain 0.272 0.270 77.2 79.3
captured-chain 0.567 0.336 61.9 45.8
grouped-axes 0.497 0.483 98.3 79.8

The captured-chain workload used about 41% less time and 26% less peak memory. Other timing ranges broadly overlapped. Grouped-axis memory fell about 19%; named-chain memory rose about 3% (roughly 2 MiB). These figures compare the pre-review working tree, including relative positioning, with the simplified code—not an older release. Raw samples and the review report are checked in as docs/benchmarks-2026-08-28.json and docs/review-simplification-2026-08-28.md.

10.3 The renderer’s passes correspond to real dependencies

  1. Classify items and collect node endpoints and captured position dependencies.
  2. Deduplicate nodes and prepare their labels/outlines; ports and clipping need those silhouettes before paths can be planned.
  3. Resolve an axis dependency graph when positions are deferred. Numeric-only node/content positions skip this graph.
  4. Resolve paths and styles, accumulating complete bounds.
  5. Emit edges/content first, then nodes on top, once the origin is known.

Combining node and edge preparation would repeat geometry work or prevent forward references. Within those passes:

  • One prepared outline drives drawing, bounds, ports, and clipping.
  • Captures use flat node tables with integer references. A pair projecting the same point imports its table once; normalized nodes are not deduplicated again. Identity buckets never stringify public capture graphs, and exact equality still distinguishes nodes that share a bucket.
  • Unlabelled nodes skip measurement. Coordinate endpoints skip silhouette lookup. Ordinary edges receive only the identity buckets they need, not the whole diagram-wide index in each memoized function call.
  • Polygon/line intersections are analytic. Curved endpoints use bounded sampling and bisection, then split the original Bézier. Distance-based label positions and tangents share their sampled distance table.
  • Regular-polygon factories precompute unit geometry. A group computes trigonometry once, then applies its affine matrix to every item.

10.4 Simplification boundaries

Forward and reverse clipping share one inside-to-outside traversal. Simple and composite node bounds use the same silhouette/label calculation; native shapes share label-overlay handling. One length-scaling helper retains percentage components while scaling physical dimensions. Successive offsets combine their deltas instead of growing an expression stack.

Some separation remains intentional. Numeric layouts skip capture normalization and the dependency solver. Partial-axis group transforms still need their local-frame rebasing, even though whole-point placement has a simpler path. Bounds accumulation keeps four local scalars; this review did not benchmark a replacement and makes no speed claim about hypothetical alternatives.

More detailed identity keys were tested and rejected: serializing normalized coordinate expressions reduced bucket collisions but slowed the grouped-axis workload. Fewer comparisons alone do not establish a faster implementation.

After expression compilation, dependency resolution visits each axis and dependency once. Constructing and interning captured snapshots costs additional work; an entire direct-capture chain is not guaranteed linear-time. For large generated graphs, named references keep each reference compact, provided the referenced nodes are explicitly emitted. Deep arbitrary expression trees, custom shape builders, and enormous diagrams are not resource-bounded by the package.