Troubleshooting
Ingress returns 404 or invalid cert
- Make sure the TLS secret
custom-ingress-certexists in thesrdpnamespace (kubectl get secret custom-ingress-cert -n srdp). - For production, Traefik uses Let's Encrypt via the ACME TLS challenge; verify the
certResolverand ACME email are set invalues-prod.yaml. - Regenerate with
mkcertif the hostnames or IP changed (local dev). - On a cluster other than kind, confirm that DNS (or your hosts file) points the hostnames to the Traefik IP.
*.srdp.localhostneeds no entries.
kind create cluster fails with "port is already allocated"
deploy/kubernetes/kind-config.yamldeliberately maps host ports 18080/18443, not 80/443/8080, becausedeploy/docker/docker-compose.ymlalready publishes exactly those three (http, https, the Traefik dashboard), and both stacks are meant to run side by side. If you still hit this error, something else on your machine is already using 18080/18443, checklsof -i :18443(or the equivalent) and either free it or pick different ports inkind-config.yaml, matching them up withvalues.yaml'straefik.ports.*.nodePort.
Pods stuck in Pending
- Check storage and DB connectivity:
kubectl describe pod <name> -n srdp. - PostgreSQL runs in-cluster; verify the service is up and its PVC is bound:
- Local (standalone): service
db-postgresql, PVCdata-db-postgresql-0 - Production (replication): service
db-postgresql-primary, PVCdata-db-postgresql-primary-0
LoadBalancer stays in Pending
- Traefik needs a LoadBalancer-capable environment. Verify your Scaleway account quotas and that the service type is
LoadBalancer(seevalues-prod.yaml).
OAuth login loops or 401s
- Ensure
global.domain,zitadel.zitadel.configmapConfig.ExternalDomain, and theoauth2-proxy.extraArgscookie/whitelist domains all match the URL you are using. - Re-check the Zitadel client ID/secret and redirect URIs.
Browser rejects the self-signed cert
- Run
just docker-tlsorjust local-tlsagain, which installs mkcert's local CA into your system trust store, and restart the browser. - Firefox keeps its own trust store, so it may need
certutil(NSS tools) installed beforemkcertcan add the CA there.
ACME errors / rate limits
- Let's Encrypt blocks
nip.iofrequently and requires public reachability on ports 80/443. Open those ports on the load balancer/security group, or temporarily point Traefik to the staging CA until production issuance succeeds.
Zitadel login-client missing
- If the Postgres DB already contains Zitadel data, the
login-clientPAT will not be recreated. Use a fresh database (or drop the existing schema) before re-running the chart.
Dagster webserver CrashLoopBackOff with password authentication failed for user "dagster"
- Cause: The
dagsterrole's password doesn't match the Secretsrdp-dagster-postgresql, keypostgresql-password(locallylocalSecrets.dagster.dbPasswordinvalues-local.yaml), or thesrdp-setupJob that creates and re-syncs it didn't run or failed. - Check the Job's logs with
kubectl -n srdp logs job/srdp-setup. The Job stays until the next install or upgrade, so its logs show the last run, failed or successful. A missingSETUP_PASSWORDS__<ROLE>fails the Job before it touches Postgres. - Fix: rerun
helm upgrade(orjust local-deploy). The Job creates any missing role or database and resets every configured role's password, keeping existing data. - If the pods don't recover on their own, run
kubectl -n srdp rollout restart deploy/srdp-dagster-webserver deploy/srdp-dagster-webserver-read-only deploy/srdp-dagster-daemon deploy/srdp-dagster-user-deployments-srdp-etl.
Updated container image not picked up after rebuild
- Cause: The default
imagePullPolicyisIfNotPresent. If you rebuild an image with the same tag (e.g.v1.0), Kubernetes will keep using the cached version. - Fix: Bump the image tag (e.g.
v1.0→v1.1) in both the build command andvalues-prod.yaml, then redeploy. Alternatively, setimagePullPolicy: Alwaysin your values file, but this is slower for routine deployments.
Orphaned Scaleway Load Balancer blocks tofu destroy
- Symptom:
tofu destroyfails withPrivate Network must be empty to be deleted. - Cause: Kapsule creates a Scaleway-managed Load Balancer when a
type: LoadBalancerservice (Traefik) is deployed via Helm. This LB is not tracked by OpenTofu, so it remains attached to the private network after the cluster is deleted. - Fix: The updated
just prod-destroyhandles this automatically. If you already hit this error, delete the LB via the Scaleway API:Wait ~30 seconds, then retrysource ./secrets.sh curl -s -H "X-Auth-Token: $SCW_SECRET_KEY" "https://api.scaleway.com/lb/v1/zones/nl-ams-1/lbs" | python3 -m json.tool curl -X DELETE -H "X-Auth-Token: $SCW_SECRET_KEY" "https://api.scaleway.com/lb/v1/zones/nl-ams-1/lbs/<LB_ID>?release_ip=true"tofu destroy.
Traefik stuck in Init
- The Traefik PVC is ReadWriteOnce; if an old pod still holds it, new pods stay in
Initwith a multi-attach warning. Delete the old Traefik pod (or the PVC if needed) so the new pod can mount/data.
Zitadel project and apps missing after a restart (Docker Compose)
- Cause: you ran
docker compose down -v, which removes all persistent volumes. - Fix: use
docker compose down(without-v) to stop the stack. Only use-vwhen you intend to reset all data.