SeenPixel

Service · Dataset Onboarding

Spec to explore

Describe a dataset in a spreadsheet row, a chat, or an API call, and the whole pipeline is generated in your own stack the same day: ingestion, orchestration in Airflow, warehouse tables and dbt models with tests, a Looker explore with per-customer access, and monitoring and alerts.

Dataset onboarding
Spec to explore

Each customer account gets its own slice.

AccountRowsRevenueOrders
Account Alpha2,201148,778.3981
Account Bravo2,20124,733.4899
Account Charlie2,11111,563.4542
Account Delta2,23327,280.7460

Alert · clock moved three days ahead

[ERROR] retail_orders_freshness: last load was 72.0h ago (error threshold 50h)

Reference implementation running on an open dataset (UCI Online Retail, CC BY 4.0). Figures are from a recorded run on DuckDB; account names are invented and the alert is printed, not sent.

What it is

Onboarding a new dataset usually means the same work every time, spread across four or five tools. Here one spec produces all of it at once, as a pull request your engineers review.

Ingestion

An extract task for the source type, and a load into a raw schema in your warehouse.

Orchestration

An Airflow DAG on the schedule in the spec: extract, load, dbt run, dbt test, publish.

Warehouse tables and dbt models

Staging, a fact table, rollups and a serving table, with not-null, unique, relationship and volume tests and a source freshness check.

Looker explore with per-customer access

A LookML view and explore joined to your accounts, with an access filter so each customer sees only its own rows.

Monitoring and alerts

Rules for stale data, volume drops and failed tests, routed to the channel the dataset owner chose.

The plan a reviewer reads

A written summary of the pipeline, the models and the checks, which becomes the description of the pull request.

What gets generated

The file list from one real run of the reference implementation: the retail_orders dataset, 23 files and 1,065 lines from a single spec. Numbers on the right are line counts.

Plan and spec

  • PLAN.md109
  • spec.yaml126

Airflow

  • airflow/onboard_retail_orders.py312

dbt models, tests and config

  • dbt/dbt_project.yml22
  • dbt/macros/generate_schema_name.sql8
  • dbt/models/marts/agg_retail_orders_account_totals.sql15
  • dbt/models/marts/agg_retail_orders_daily_by_country.sql15
  • dbt/models/marts/agg_retail_orders_monthly_by_product.sql13
  • dbt/models/marts/fct_retail_orders.sql19
  • dbt/models/schema.yml112
  • dbt/models/serving/srv_retail_orders.sql7
  • dbt/models/serving/srv_retail_orders_accounts.sql4
  • dbt/models/staging/stg_retail_orders.sql17
  • dbt/models/staging/stg_retail_orders_accounts.sql7
  • dbt/profiles.yml20
  • dbt/tests/assert_retail_orders_min_rows.sql5

LookML

  • lookml/retail_orders.model.lkml23
  • lookml/retail_orders.view.lkml123
  • lookml/retail_orders_accounts.view.lkml16

Alerts

  • alerts/notifier.py26
  • alerts/retail_orders.alerts.yml24

Snowflake SQL

  • snowflake/load_raw.sql27
  • snowflake/row_access_policy.sql15

Paths are relative to generated/retail_orders/. In the reference implementation the Snowflake SQL, the DAG and the LookML are generated and checked, not run in those systems; see the table further down.

How it works

  1. 1 · Spec

    Someone who knows the dataset describes it: where it comes from, its columns and measures, the schedule, the owner, and which column identifies the customer. A spreadsheet row, a chat with an assistant, or an API call all produce the same spec file.

  2. 2 · Plan as a pull request

    The generator turns the spec into a pull request in your repository: the DAG, the models and tests, the LookML, the alert rules, and a plan written for the reviewer.

  3. 3 · Review and merge

    An engineer on your side reads it, asks for changes or approves it, and your existing checks run. Nothing reaches production without that approval.

  4. 4 · Runs on your schedule

    After the merge the pipeline runs in your Airflow, the tables land in your warehouse, the explore is available to the right customers, and the owner hears about stale or broken data.

Sources supported

Four source types share one spec format. Credentials never go in the spec.

Files

CSV or Parquet at a path or URL: exports, partner drops, bucket folders.

Databases

A table read over a standard database connection.

APIs

Paginated REST endpoints.

Event streams

A Kafka topic, consumed in batches.

Who it is for

Companies that onboard new datasets or new customers again and again, where each one has to be tied to the right account.

B2B software with customer-facing analytics

Every new customer or product module is another dataset that has to be scoped to the right account.

Data-product companies

Each new source or supplier is a pipeline, and each customer gets a view of it.

Martech and measurement

New publisher, platform and partner feeds arrive for every customer.

Fintech with partner feeds

Each merchant or bank partner sends data in its own shape, on its own schedule.

If you onboard a dataset once a quarter, this is more machinery than you need, and we will say so.

How we deliver it

Installed in your environment

The generator and its templates are set up in your repository and adapted to the tools you already run: your Airflow, your warehouse, your dbt project, your Looker instance. Connecting it to real sources, secrets, CI checks and alert delivery is the main work of the install.

You own the code

The generator, the specs and everything it produces are plain Airflow, dbt, SQL and LookML in your repositories. If we stop working together, nothing stops running.

Operated and extended by a pod

Once it is in place, a dedicated pod can run it with you: new source types, new checks, upgrades as your tools change, and the datasets that do not fit the template.

See the dedicated pod and how we work.

The reference implementation, stated plainly

The walkthrough at the top of this page is a demo-grade reference implementation, not a production install. It runs on a subset of an open dataset with DuckDB as the warehouse, on one machine with no cloud account. This is what it does and does not do.

PieceIn the demoNotes
Spreadsheet import and spec validationRunsFour datasets are read from one intake spreadsheet and validated.
File source (CSV)Runs8,746 real rows from an open dataset, bundled with the demo.
API source (paginated REST)RunsAgainst a mock server bundled with the demo, not a public API. 3,200 synthetic events.
Database sourceGenerated onlyThe DAG task code is produced and checked as text. No database is started.
Event stream source (Kafka topic)Generated onlyThe DAG task code is produced and checked as text. No broker is started.
Load into the raw schemaRunsOn DuckDB, an embedded warehouse on a single machine. Full refresh on every run.
dbt models, tests and source freshnessRunsOn DuckDB: 8 models and 23 tests for retail_orders, 7 models and 19 tests for web_events. All tests pass.
Snowflake SQL and the dbt Snowflake targetGenerated onlyNever run against Snowflake.
Airflow DAGGenerated and checkedSyntax-checked and imported by Airflow 3.1.0. Not scheduled; no task is executed by Airflow.
LookML view and exploreGenerated and checkedParsed for syntax only. Not validated or deployed in a Looker instance.
Alert rulesEvaluatedDelivery is a stub: the alert is printed and logged, nothing is sent.
Chat interface (MCP server)RunsFour tools: validate a spec, show the plan, generate, run the demo. Exercised by the test suite.

A production install adds what the demo leaves out: a real warehouse with incremental loads, managed connections to your sources, CI checks on every generated pull request, a secret manager, deployment to your Looker instance with per-customer user attributes, and real alert delivery. The file-source demo uses a subset of the Online Retail dataset (Chen, D., 2015, UCI Machine Learning Repository, CC BY 4.0). Account names and the API events are synthetic.

Questions we get

Does it replace our data engineers?

No. It removes the repetitive part: the same DAG, models, tests, explore and alerts written again for each new dataset. Engineers still design the template, review every pull request and handle the datasets that do not fit the pattern.

Do we need Looker?

LookML is what the generator produces today, including the per-customer access filter. Output for another BI tool or semantic layer can be added by arrangement as part of the install.

What about our existing pipelines?

They stay as they are. The generator adds new datasets next to them, following the conventions of your repository. Existing pipelines can be moved onto specs later, one at a time, if that is worth doing.

Who owns the code?

You do. The generator, the specs and all generated code live in your repositories and belong to you. There is no SeenPixel runtime in your production path.

Can a non-engineer onboard a dataset?

A non-engineer can describe one. The spreadsheet row or chat produces the spec and the pull request; an engineer approves it before anything runs.

Is the demo the product?

No. The demo is a reference implementation on an open dataset with an embedded warehouse; the table on this page says what it runs and what it only generates. A production install is built in your stack. See the reference implementation.

Bring one dataset you have to onboard next

We will walk through the reference implementation with you, then talk about what it would take to do the same in your stack.