DuckPipe

A serverless-first, DuckDB-native pipeline orchestrator. No scheduler daemon, no central metadata database, no broker — a run is a Python process that starts, does work, records what it did to a .duckdb file, and exits.

CI status PyPI version MIT license
Try it live in your browser → View on GitHub

The link above runs a real DuckDB engine and DuckPipe's own, completely unmodified source — no server, no install — checking a file for sensitive-looking columns without ever uploading it. It's also the fastest way to see what a run actually looks like: the same task-status table and DAG view duckpipe show prints in a terminal, rendered live after a real in-browser run.

Or run it locally in 10 seconds

# pipeline.py
import duckdb
from duckpipe import task, run

TAXI_DATA = "https://d37ci6vzurychx.cloudfront.net/trip-data/yellow_tripdata_2024-01.parquet"

@task
def extract():
    return duckdb.sql(f"SELECT * FROM read_parquet('{TAXI_DATA}')")

@task(cache=True)
def daily_totals(trips=extract):
    return trips.aggregate(
        "date_trunc('day', tpep_pickup_datetime) AS day, sum(fare_amount) AS total"
    ).pl()

if __name__ == "__main__":
    run(__file__)
uv add duckpipe
uv run duckpipe run pipeline.py

That's the whole surface area otherwise — no fixture to find, no scheduler to stand up. Full quickstart and mental model in the README.

Explore

Why DuckPipe

The pain points this design responds to, mapped to the code that answers each one.

Read →

Examples

Nine realistic pipelines over real, bundled open data — one per facet of DuckPipe.

Browse →

Scaling out

Distributed execution, DuckLake observability, a serverless executor, and the browser — each opt-in.

Read →

Design rationale

The full design rationale and prior-art landscape check behind every tenet.

Read →

Docs