DuckPipe

A serverless-first, DuckDB-native pipeline orchestrator.

View the Project on GitHub woozyking/duckpipe

07 — serverless executor, two platforms, one run

The reference serverless executor (DESIGN.md sec 8): not a new mechanism — duckpipe.run(module, only=task, state_uri=...) already is the whole “serverless executor,” unchanged. What was missing was checking, not just claiming, that it’s genuinely not tied to one platform’s invocation model. This example dispatches pipeline.py’s two tasks into one distributed run, each through a different shape:

Both are the same one-line call underneath. Nothing in pipeline.py (or in DuckPipe itself) knows or cares which shape dispatched it.

docker build -f examples/07_serverless_executor/Dockerfile -t duckpipe-worker .
uv run python examples/07_serverless_executor/run_serverless_demo.py

What happens: the script dispatches extract via docker run duckpipe-worker --only extract --state-uri file:///data/duckpipe.db, then dispatches summarize via a direct handler({"task": "summarize", "state_uri": ..., "run_id": ...}) call — the exact shape a FaaS platform’s own runtime would call it with. Both write into the same state_uri-backed bucket, coordinating with nothing but the delta-merge mechanism from ../04_distributed_cluster, then duckpipe compact folds both workers’ deltas into one file to report on.

Making this a real deployment, beyond this local proof:

See ../../docs/remote_execution.md for “beefy node” mode — the other extension of the same primitive, for when the bottleneck is one big task rather than many independent ones.