What each component is responsible for
The medallion mapping, stated explicitly
The warehouse is a textbook medallion architecture, and the folder names are the layers:transform/models/staging/= bronze view onto the raw Parquet (typing, renaming, light cleaning).transform/models/intermediate/= silver: the physics, the skill model, the strategy features, each a tested building block.transform/models/marts/= gold: the wide feature tables and fact tables the app and ML layer consume.
The defining decision: build-time and request-time are separated
Almost every other shape this project could take would put a database and an inference server behind an API. Off The Pace deliberately does not. The heavy compute happens in two places that are not the request path:- Build time (CI): ingestion, the entire dbt warehouse, model training, ONNX export, and every drift gate run in GitHub Actions. The output is a set of static Parquet and ONNX files published to a CDN.
- Request time (client): the browser downloads those files once and runs DuckDB-Wasm and ONNX Runtime Web locally. Every query and every inference happens on the user’s machine.
What it buys
Zero per-user serving cost, no backend to operate or scale, offline-capable analytics, and a build pipeline where correctness is gated before anything ships.
What it costs
A larger one-time download, no row-level authorization, and a data freshness bounded by the publish cadence rather than live streaming.
This is a trade we made on purpose, not a limitation we backed into. For a use case that is read-only, public, and analytical, pushing compute to the client is the cheapest correct answer. A use case needing per-user data or real-time writes would choose differently.
The “why” companions
- Architecture decisions records each choice as an ADR, including ADR-002 (client-side analytics), ADR-003 (CDN Parquet over an active warehouse), and ADR-010 (version-stamped cache-busting so a data change is never masked by a stale cache).
- The project lineage graph is the live, generated DAG of every model and its dependencies, the same view dbt builds the warehouse from.
Where to go next
How it clears a production bar
Contracts, drift gates, lineage, orchestration, and SLOs as standards met.
Engineering highlights
The eight things a reviewer should notice first.
The data pipeline
The staged orchestration DAG, retries, and post-publish verify.