Bruin Academy

Guide

End-to-End Pipeline: NYC Taxi

Build a complete data pipeline from scratch using real NYC taxi data - from ingestion to staging to reports, all orchestrated with Bruin and DuckDB.

Skip ahead

The whole project is one command. Run it, point your coding agent at the Bruin MCP, and it can configure and run everything itself.

$ bruin init zoomcamp

Three ways to go from here:

  • Self-service it. Let the agent configure and run the whole thing, and ask it questions as you go.
  • Follow this tutorial. Slower, and it explains why each setting matters - the part an agent will not guess for you.
  • Hand the tutorial to the agent. Point it at this page and have it work through the steps with you.

What

Build a real data pipeline end-to-end using NYC taxi trip data. Go from raw API data to clean, aggregated reports - learning ingestion, transformation, quality checks, and AI-assisted development along the way.

  • End-to-end ELT pipeline: Python ingestion, SQL staging, reporting layers with quality checks
  • Full orchestration with dependency management, execution order, and visual lineage
  • AI integration via bruin ai enhance and Bruin MCP

How

  • Bruin CLI orchestrates the pipeline; DuckDB serves as the local data warehouse
  • Python assets ingest from the NYC TLC API; SQL assets handle transformations
  • Bruin MCP connects an AI agent for pipeline development and data analysis

Sign up to our newsletter

Practical updates on open-source data pipelines, AI analysts, governance, and what we are shipping at Bruin.

The signup form is hosted by Brevo. Accept cookies to load it.