Reference implementation · not client work
Dataset Onboarding reference implementation
One dataset spec goes in; an Airflow DAG, dbt models and tests, a LookML explore with per-customer access and alert rules come out, and two of the four source types are run end to end on an open dataset. A walkthrough is available on request.
23
files, 1,065 lines, generated from one spec for the retail_orders dataset.
8,746
rows from an open retail dataset loaded into DuckDB by the file-source run.
23/23
dbt tests passed across 8 models, with the source freshness check passing.
Figures from a recorded run of the demo. They describe the demo, not a customer system.
What the demo does
Intake
Four datasets are described in one spreadsheet, one row each: a file source, an API source, a database source and an event stream.
Spec
Each row becomes a validated YAML spec: columns, measures, rollups, schedule, owner, freshness limits and the column that identifies the customer.
Generate
From each spec: an Airflow DAG, dbt staging, fact, rollup and serving models with tests, a LookML view and explore with a per-customer access filter, alert rules, Snowflake SQL, and a plan written like a pull-request description.
Run
The file and API datasets are executed end to end on DuckDB: extract, load, dbt run, dbt test, publish. 8,746 rows and 8 models for retail_orders; 3,200 rows and 7 models for web_events.
Per-customer result
Rows are joined to an accounts mapping, and the serving table returns one slice per customer account. A test fails when a row has no account.
Alert
With the clock moved three days ahead, the freshness rule fires. Delivery is a stub: the alert is printed and logged, nothing is sent.
What it does not do
It is a demo-grade starting point, not the production product. It stops where real infrastructure starts.
- DuckDB, an embedded warehouse on one machine, stands in for the warehouse. The Snowflake SQL and dbt Snowflake target are generated and have never been run against Snowflake.
- The Airflow DAGs are syntax-checked and imported by Airflow 3.1.0. They are not scheduled, and no task is executed by Airflow.
- The LookML is parsed for syntax only. It is not validated or deployed in a Looker instance.
- The database and event-stream sources are generated only. No database or broker is started.
- The API source runs against a mock server bundled with the demo, not a public API.
- Loads are full refresh. There is no incremental or merge logic.
- The retail rows are a subset of a real open dataset. Account names and the API events are synthetic.
Source data: a subset of the Online Retail dataset (Chen, D., 2015, UCI Machine Learning Repository), licensed CC BY 4.0. The full table of what runs and what is only generated is on the service page.
Walkthrough on request
There is no public demo link. We run it live in a 30-minute session and open the generated files with you. Ask through the contact form and pick “Dataset onboarding”.
Want the walkthrough?
A 30-minute live session on the reference implementation, then a conversation about what it would take in your stack.