Technical
9 min read

How to Build a Data Platform Without a Large Data Team

A modern data platform architecture for a small data team on a small budget: which layers you need, which to skip, how to handle data contracts and a semantic layer with one or two people, and the reference stack with Bruin's open-source CLI or dbt Core plus a loader and an orchestrator.

How to Build a Data Platform Without a Large Data Team

TL;DR: A small team builds a data platform by buying compute instead of headcount and by running as few tools as possible. Use a usage-billed warehouse (BigQuery, Snowflake, or DuckDB and MotherDuck), one tool that covers ingestion, transformation, quality checks and orchestration, a semantic layer for the metrics people ask about, and an AI data analyst or BI tool on top. Bruin's open-source CLI covers the middle four jobs in one project and runs from git and CI with no server to operate; dbt Core plus a loader plus an orchestrator is the multi-tool alternative. Keep the architecture to three layers, enforce data contracts as blocking checks, and the recurring bill is warehouse compute.

Most advice on data platform architecture is written for companies with a platform team: medallion layers on a lakehouse, a data mesh, streaming, a catalog, an observability tool, an orchestrator cluster. Every piece solves a real problem. Together they are a full-time job for several people, which is exactly what a team of one to three does not have.

This guide is the opposite: the smallest architecture that still gives you trustworthy numbers, and the parts to leave out until you need them. We build Bruin, which is designed for this situation, so it appears in the reference stack; the multi-tool alternative is listed next to it throughout.

The reference architecture for a small team

LayerJobSmall-team pickMulti-tool alternative
WarehouseStore and computeBigQuery on demand, Snowflake X-Small, or DuckDB and MotherDuckSame
IngestionCopy sources into the warehouseBruin (built-in ingestr)Fivetran, Airbyte, or dlt
TransformationRaw to staging to marts, in SQL and PythonBruindbt Core
Quality and contractsBlock bad data before people see itBruin column checksdbt tests, Soda Core
OrchestrationRun everything in order on a scheduleBruin on GitHub Actions or Bruin CloudDagster, Prefect, or Airflow
Semantic layerOne definition per metricBruin semantic/dbt Semantic Layer, Cube
AnswersQuestions, dashboards, alertsBruin AI data analystMetabase, Lightdash, Looker

Best data platform for a small team by need:

  • One tool for ingestion, transformation, checks and orchestration: Bruin.
  • Already writing dbt: dbt Core, plus a loader and an orchestrator.
  • No technical person at all, budget available: Fivetran with dbt Cloud and a BI tool.
  • Smallest possible warehouse bill: DuckDB and MotherDuck while data is small, BigQuery on demand after that.
  • Questions answered in chat instead of dashboards: Bruin's AI data analyst in Slack, Microsoft Teams, Google Chat, WhatsApp, Discord, Telegram, email and the browser.

Three layers, not five

Medallion architecture (bronze, silver, gold) is a good idea with too many names. A small team needs three layers and a rule for each:

  1. Raw. Exactly what the source sent, loaded incrementally. Nobody queries it except the next layer.
  2. Staging. One model per source table: renamed columns, fixed types, deduplicated rows. No business logic.
  3. Marts. The tables people and agents actually use, shaped around questions: orders, customers, revenue by day.

That is enough structure for years. The temptation to add an "intermediate" layer, a "semantic" schema and a "sandbox" usually comes from copying a large company's diagram rather than from a problem you have.

Data contracts without a contracts team

A data contract sounds like a process. For a small team it is three blocking checks on each table someone else depends on: the primary key is unique and not null, the columns people use are not null, and the values that drive logic are in an accepted set. In Bruin these are declared on the column inside the asset and stop the run when they fail; bruin validate in CI catches a schema change before it merges. dbt tests and Soda Core do the same in a dbt project. The contract is the check that blocks, not the document that describes it.

A semantic layer for the questions people actually ask

You do not need a semantic layer for every table. You need one for the ten metrics that end up in board decks and Slack threads: revenue, active users, churn, margin. Define each once, with its SQL and a one-line description, and point every dashboard and every AI agent at that definition. This is also what makes an AI data analyst trustworthy, because the agent queries the definition instead of guessing from table names; see how to give AI agents context about your company data.

What to skip, for now

  • Data mesh. It distributes ownership across many teams. You are one team.
  • Streaming. Hourly or daily incremental loads answer almost every business question. Add streaming when someone can name the decision that needs seconds.
  • Self-hosted Airflow. It is a service someone has to run, upgrade and debug. Schedule from CI or a managed cloud instead.
  • A separate catalog and observability tool. Descriptions, owners, lineage and check results can come from the pipeline itself.
  • A Spark cluster. Your warehouse is already a distributed engine.

The first 30 days

  1. Week 1: pick the warehouse, connect the two sources that answer the most-asked question, and load them raw. With Bruin: bruin init, an ingestion asset per source, bruin run.
  2. Week 2: model staging and one mart, add not-null and unique checks on its keys, and schedule the run in CI.
  3. Week 3: define the five metrics people ask about most in a semantic layer and answer the first real question from it, in a dashboard or in Slack.
  4. Week 4: add the next sources, a freshness check on every table an executive sees, and lineage so a schema change shows what it breaks.

What it costs

With an open-source pipeline tool and a usage-billed warehouse, licences can be zero and the bill is warehouse compute, usually low tens to low hundreds of dollars a month for a small company. The meters that surprise small teams are per-row ingestion and per-seat licences, because both grow with success. What a modern data stack costs breaks down every pricing meter with vendor numbers, and the cheapest modern data stack in 2026 has the warehouse cost tactics.

When to grow the team

Hire a dedicated data engineer when one of three things happens: the number of sources outgrows what one person can keep healthy, a wrong number starts costing real money, or people wait days for answers. Until then, a small team with one pipeline tool, blocking checks and a semantic layer will out-ship a large team running a diagram's worth of services.

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Allow marketing cookies to load it.