Custom data and AI engineering · New York
Your team asks a question in plain English. Your data answers, with the right number.
SeenPixel builds the governed AI layer over the warehouse you already have, and the data platform underneath it.
Illustrative example. Names and figures are invented to show the format.
This is for you if
Dataset Onboarding · spec to explore
From a spreadsheet row to a dashboard, the same day
Describe a new dataset once. The ingestion, the Airflow schedule, the warehouse tables and dbt models with tests, the Looker explore with per-customer access, and the alerts are generated together, as a pull request your engineers review.
For companies that onboard new datasets or new customers again and again. It runs on the tools you already have, and you own the code.
Reference implementation running on an open dataset (UCI Online Retail, CC BY 4.0). Figures are from a recorded run on DuckDB; account names are invented and the alert is printed, not sent.
Most AI projects built on company data never go live. The reasons are known.
The AI model is rarely the problem. The basics are: nobody agreed what each number means, who is allowed to see it, how answers get checked, or who watches the bill. In engineering terms, that is undefined metrics, missing access control, no evaluation loop and unowned warehouse costs. We build the parts that are usually skipped.
40%+
of agentic AI projects are expected to be canceled by the end of 2027, over cost, unclear value or inadequate risk controls.
Source: Gartner, 2026
60%
of agentic-analytics projects that rely only on MCP will fail by 2028 without a consistent semantic layer.
Source: Gartner, May 2026
71%
of data practitioners worry about hallucinated numbers reaching stakeholders.
Source: dbt Labs, State of Analytics Engineering 2026
What we build
Five offers, each with a clear scope and a clear end. Two more, data platforms and software products, when the work calls for them.
Flagship
Ask-your-data agents
People across the company ask questions in plain English and get answers they can trust, from the data you already have.
Governed LLM agents over the warehouse you already have: semantic layer, MCP server, evaluations, access control.
Details →Flagship
Dataset Onboarding
Describe a new dataset in a spreadsheet row or a chat, and the whole pipeline for it is generated in your own stack the same day.
Spec to explore: one dataset spec becomes the Airflow DAG, dbt models and tests, Looker explore and alerts, as a pull request in your repo.
Details →Warehouse cost audit
Find out exactly where the data warehouse bill goes, and bring it down.
A fixed-fee review of Snowflake, Databricks or BigQuery spend, with a prioritized savings plan and the fixes.
Details →Semantic layer & BI
One agreed definition of every number, so dashboards, reports and AI all give the same answer.
Looker, Cube and dbt metrics; Looker-to-X migrations; one definition of each metric for dashboards and agents alike.
Details →Dedicated data + AI pod
Senior data and AI engineers for ongoing work, without a long hiring process.
Senior engineers under one accountable lead, scoped quarterly, once an audit or pilot has earned it.
Details →How we work: a ladder with an exit at each step
Engagements start small and earn the next step. At every step you leave with something you own, and you can stop.
Step 0 · About a week
Free readiness scan
A short conversation and a read-only look at usage, with a two-page findings note. No commitment.
Exit: You keep the findings. Nothing else is owed.
Step 1 · 2–3 weeks
Fixed-fee audit
Cost recovery, or readiness for semantic layer and agents. A prioritized plan and the specific fixes, written down.
Exit: Take the plan to your own team or another vendor. It is yours.
Step 2 · 6–8 weeks
Pilot
One governed ask-your-data agent in production, or one semantic-layer migration, with a stop-point at week two.
Exit: A working system, the code and the docs. Stop here if that is all you need.
Step 3 · Scoped quarterly
Dedicated pod
Senior engineers under one accountable lead, working against a quarterly scope you set with us.
Exit: Quarter by quarter. No long-term lock-in.
You own the code, the models and the documentation at every step. Details on how we work and security.
Proof, not slides
Things you can look at today: a product we run, the design we start from, and how we think.
Product we built and operate
Kitvine
A two-sided creator marketplace on Next.js, Prisma and Postgres, with its own automated content and operations routines. We built it and we run it.
Read the case study →Reference architecture
Governed ask-your-data agent
How a question travels from a person to your warehouse and back, with access rules and checks at every step. Agent runtime, MCP server, policy, semantic layer and evaluation loop: the design a pilot starts from.
See the design →Engineering notes
Why 'chat with your data' projects fail
What the failure data says, and what a semantic layer changes about the odds.
Read the post →Stack we build on
Snowflake · BigQuery · Databricks · dbt · Airflow · Kafka · Looker · Cube · Postgres · Next.js · Claude and OpenAI APIs · MCP
Industries we work in most
Media and martech
Warehouse-native marketing data, identity, clean rooms and the reporting that sits on top.
Fintech
Governed access to numbers that have to be right, with an audit trail for every query.
PE-backed SaaS
Operating partners who need working AI and data systems inside a quarter, not a slide deck.
And others. The engineering is the same; the domain is something we learn fast.
Questions we hear first
What does SeenPixel do?
SeenPixel is a data and AI engineering company in New York. We build and run data platforms, semantic layers, governed LLM agents over existing warehouses, and software products. The work is engineering: code, tests, deployments and operations, not strategy decks.
How does an engagement start?
With a free readiness scan: a short conversation and a read-only look at your warehouse usage, followed by a two-page findings note. From there, a fixed-fee audit, then a 6–8-week pilot, then a dedicated pod if it earns it. There is an exit at each step.
How long does a pilot take?
Six to eight weeks for one governed ask-your-data agent in production or one semantic-layer migration, with a stop-point at week two where you can end the work and keep what exists.
Who owns the code and the IP?
The client. Code, models, semantic definitions, documentation and infrastructure configuration are delivered into your repositories and accounts and belong to you.
How do you handle data access and security?
An NDA before any data access, least-privilege roles inside your environment, no production credentials held outside it, and an audit trail for every query. Details are on the security page and available in full under NDA.
Do you work with companies outside New York?
Yes. We work with companies across the United States. Engagements run remotely, with a weekly working session and a written status every week.
Tell us what you're trying to build
A 30-minute technical conversation about your stack and what you are trying to do. If a free readiness scan makes sense, we will say so; if not, we will say that too.