Welcome to the OpenDuck docs. OpenDuck is an open-source implementation of differential storage, hybrid (dual) execution, and transparent remote databases for DuckDB — a self-hostable, protocol-open take on the architecture pioneered by MotherDuck.
import openduck
con = openduck.connect("mydb")
con.sql("SELECT * FROM users").show() # remote, transparent
con.sql("SELECT * FROM local.t JOIN cloud.t2 ...") # hybrid, one query- Overview — what OpenDuck is, the problems it solves, and how it compares to MotherDuck, Arrow Flight SQL, and DuckLake.
- Architecture — components, data flow, the protocol, and how a query becomes Arrow batches.
- Configuration — every CLI flag, environment variable, TOML key, and DuckDB secret OpenDuck understands.
- Getting started — build the backend and the extension, start the service, run your first query.
- Python client — the
openduckPython package: connections, hybrid queries, pandas/Arrow integration. - DuckDB extension —
LOAD,ATTACH, URI format, secrets, table functions. - Differential storage — append-only layers, snapshots, the three storage modes (Direct, FUSE, In-Process).
- Hybrid execution — how to enable it, how plans split, the
openduck_runhint. - Snapshots and garbage collection — sealing, point-in-time reads,
openduck gc. - Deployment — Docker, multi-worker, DuckLake, S3 tiering, observability.
- Troubleshooting — common errors and how to fix them.
- Protocol definition — the wire format (gRPC + Arrow IPC).
- Examples — runnable Rust and Python examples for every major feature.
- Python client README — full API for the
openduckpackage.
Source layout, build instructions, and CI are in the top-level README.
MIT.