Skip to content

Single-Repo Data Platform (SRDP)

SRDP is a self-hostable data platform assembled from established open-source components and deployed from a single Git repository. It is aimed at teams that want a coherent, governable data stack without operating a large set of separately managed services. Combining a single-repository approach containing components known from homelab stacks and modern data platforms, SRDP is a serendipitous mix of best practices in one data platform. It is designed to be easy to deploy, easy to operate, and easy to extend.

The same logical architecture runs from a laptop (Docker Compose) to a single VM to a Kubernetes cluster, so development and production stay aligned.1

Components

Component Role in SRDP
Zitadel Identity and access management. The single OIDC issuer; handles authentication and coarse project roles.
Traefik Reverse proxy and ingress. Terminates TLS, routes by hostname, and load-balances across services.
OAuth2-Proxy Edge authentication. Runs the OIDC login flow and gates protected services via Traefik forward-auth.
DuckLake / DuckDB Lakehouse storage and the in-process analytical query engine.
Dagster Data orchestration. Asset-based pipelines with scheduling, retries, and lineage.
dlt Data loading. Planned, not yet integrated.
dbt SQL-based data transformation.
Polars DataFrame library for in-process transformation in Python.
marimo Reactive Python notebooks for interactive analysis.

PostgreSQL backs the platform's stateful services (Zitadel and Dagster, and the DuckLake catalog).

Architecture

SRDP architecture

See Architecture & Conventions for the service topology and the edge authentication model, and the Architectural Decision Records for the design decisions behind the platform.


  1. The single-repository approach takes inspiration from the Instant OpenHIE project, which packages an open-source health information exchange the same way. ↩