Creating a PostgreSQL primary and a replica takes more than starting two containers. The platform has to remember the request, reserve the storage, distinguish the primary from its replica, give the application a usable connection and handle an interrupted operation without losing track of what it created.
Hakopod records those decisions in a database resource with its own project, environment and revision. This article follows that resource through the implementation. Self-hosted alpha.47 includes PostgreSQL, Redis and MongoDB, guided creation and the database cockpit. MySQL, ClickHouse, Oracle Database, Neon, Supabase and Vitess remain unavailable while native qualification is incomplete. Their sections below explain the implementation under development. Cloud provisioning still needs operator setup and approved capacity; publishing the engine does not establish a production deployment.
Start with a configuration you can inspect
This example asks for PostgreSQL 17 with a primary and one replica, required TLS and a different node for each member:
schema_version = 1
name = "orders"
engine = "postgresql"
version = "17"
mode = "cluster"
replicas = 1
shards = 1
cpu = "500m"
memory = "1Gi"
storage_gib = 10
[tls]
mode = "required"
[placement]
spread = "nodes"
Save it as orders.toml, then use the shared CLI:
hakopod database create --project demo --environment development --file orders.toml
hakopod database list --project demo --environment development
hakopod database show DATABASE_ID
The two members need separate data volumes and enough eligible capacity. The requested 1Gi applies to each PostgreSQL member. Replacement and recovery headroom add to the reservation. If the two Kubernetes nodes live on one physical VM, the placement rule does not protect the database from losing that VM.
The dashboard sends the same kind of request to the Go API. The handler checks the versioned configuration, scope, permissions, controller prerequisites and allocation. It stores the accepted revision and operation in the control-plane PostgreSQL database. The response means work was accepted; readiness comes from later observations.
The operation survives the request
The API and reconciliation workers run in one Go process by default. PostgreSQL holds their durable work. The lifecycle worker claims an operation with a short lease, makes a bounded reconciliation step, records its result and schedules the next step when needed. Before changing runtime resources, it rechecks the current revision, lease and authority.
A controller rollout can take minutes. Keeping it in durable steps means an HTTP timeout does not become the only record of what happened. Expired claims can return to the queue. A stale worker cannot keep writing after another revision or lease has replaced its authority.
Native observation has separate workers, so a slow database probe does not occupy the lifecycle lane. Observations include a revision and timestamp. The dashboard must not use a healthy result from an earlier revision to approve a new resize.
The implementation is readable in the API and worker code and the operation store.
Controllers have different jobs
Hakopod owns the database identity, accepted configuration, access policy, allocation and operations. Engine-specific controllers manage their native Kubernetes resources. Kubernetes schedules pods and attaches storage; each engine’s replication protocol determines write availability.
| Engine | What runs with the data members |
|---|---|
| PostgreSQL | CloudNativePG, with optional PgBouncer services. |
| Redis | The credential-safe Opstree Redis operator; cluster clients follow slot ownership. |
| MySQL | MySQL Operator, InnoDB Cluster sidecars and MySQL Router. |
| MongoDB | MongoDB Kubernetes Controller and member agents. |
| ClickHouse | Altinity operator and three Keeper members for clustered metadata. |
| Oracle Free | An owned standalone StatefulSet, TCPS listener and separate data/backup volumes. |
| Vitess, under construction | A namespace-scoped operator, MySQL/vttablet, vtgate, vtctld, vtorc and three etcd members. |
Those supporting processes use resources too. For three MySQL members at 500m CPU and 1Gi each, the sidecars and two Routers bring steady requests to 2 CPU cores and 4Gi before operating headroom. ClickHouse adds Keeper volumes and backup staging volumes. A database can fit its normal members and still lack room to replace one safely.
Cloud checks database reservations against the same workspace allocation used by applications. Removing a pod does not refund a retained volume. Deletion must finish reclaiming the exact owned data before releasing its reservation.
Multi-node placement starts with an operator-approved allocation. Cloud checks node names and UIDs, available CPU and memory, other workloads and a separate storage budget. It reserves the full database envelope on each approved worker. That is deliberately conservative: dividing the reservation across workers would leave less room when a member needs to move or be replaced.
The database form shows the approved nodes and their reported architecture, zone and provider. If a selected node disappears, the selection stays visible and progress stops until it is corrected. The CLI exposes the same inventory through hakopod database nodes. A separately operated cluster has its own approval and scoped-key connection flow; its network and storage must already work.
Oracle Free is proprietary software with upstream limits. Its current recovery qualification is incomplete, and creation remains disabled. Earlier standalone checks do not establish Enterprise licensing, Data Guard, RAC or cross-host failover support.
Neon and Supabase need a different recovery boundary
Neon and Supabase are PostgreSQL-based systems with more processes to manage. Their self-hosted implementations are under construction and remain unavailable in Hakopod. New hosted PlanetScale connections have been removed from the creation flow; existing connections retain their inspection, credential refresh and disconnect paths.
Neon separates PostgreSQL compute from durable storage. Safekeepers receive the write-ahead log, pageservers build and serve database pages, and object storage holds durable files. A broker helps those processes discover each other. Starting a PostgreSQL container alone does not deploy that stack. The control plane also needs to prove which tenant, timeline and compute instance it owns after a process restarts, before retrying creation or deletion. The Neon guide explains those boundaries and the work still awaiting native testing.
Supabase includes PostgreSQL, Auth, PostgREST, Realtime, Storage, Edge Runtime, an API gateway, a connection pooler, postgres-meta, an image proxy and Studio. Restoring its SQL rows alone does not restore uploaded objects, Edge Functions or the database encryption key. The candidate recovery manifest binds those separate assets to one resource revision. Its gateway must serve verified HTTPS, while Studio needs administrator access separate from an application’s API key. The Supabase guide describes the pinned stack, storage and access controls.
Open-source code makes these systems inspectable. It does not make a new deployment automatically ready for production. Both candidates still need full-stack startup, credential rotation, recovery and failure testing before the catalog can offer them.
Choose the connection’s behavior
A PostgreSQL primary route, a replica route and a pooled route are different choices. PgBouncer reuses server connections. It does not inspect arbitrary SQL and decide which statements belong on a replica. Transaction pooling releases a server connection at transaction end; session pooling retains it for the session. Applications that rely on session state need a compatible mode.
MySQL Router exposes separate primary and replica ports. Redis Cluster clients follow slot redirections. MongoDB drivers discover members and apply read preference and write concern. ClickHouse needs Distributed tables or explicit query design to combine shards; balancing connections does not make a local table query read every shard.
Vitess routes through an explicit VSchema. The implementation under construction declares integer hash sharding columns and SINGLE transaction mode. It does not promise live resharding or cross-shard transactions. Internal identity verification, replication and recovery still need native acceptance before availability can be claimed.
Every application needs a reconnection policy. Losing a connection after sending COMMIT does not prove that the write failed. Use bounded retries and application idempotency where the transaction outcome is uncertain. The routing guide explains the endpoint choices.
Give the application an account and verified trust
An application specification stores a binding reference:
[services.api.bindings.DATABASE_URL]
managed_database = "DATABASE_ID"
protocol = "postgres"
endpoint = "read_write"
Connection replacement is reviewed against the application and database revisions. Acceptance saves the reference and queues deployment in one transaction. The runtime resolves restricted application credentials and public CA material for that service. Administrative, monitoring and recovery credentials remain separate, and the application never receives Kubernetes credentials.
TLS needs to verify the server the application actually reaches. A private network address does not prove identity. Native checks verify the hostname, issuer, served certificate and authenticated database access. Router backend verification needs its own test: a correctly encrypted client-to-router session says nothing about whether that router verifies its database server.
Renewal has several steps: issue a new identity, load it at the server, preserve supported trust overlap, update bound applications and test new connections. A long-lived session may continue while a fresh client fails. Testing both catches a different class of problem from checking that a Secret changed.
Decide what a backup means before restoring it
Different engines provide different capture boundaries. PostgreSQL uses a logical dump. Redis capture is consistent per shard. MySQL holds a global read lock while its logical dump runs, so writes and DDL wait. MongoDB uses one snapshot read timestamp across supported collections and checks metadata changes. ClickHouse captures native archives sequentially by shard. Oracle Free captures the application schema with Data Pump at a flashback SCN and a DDL guard.
These are not interchangeable promises. A set of valid shard archives does not establish one atomic cross-shard snapshot. A logical archive does not imply continuous log recovery or point-in-time recovery. Vitess’s proposed per-shard archive has the same explicit limit on global consistency and still needs runtime verification.
The managed backup service encrypts archives and verifies stored bytes against recorded evidence. Restore authenticates input before writes and uses a separate compatible empty target. Application ingress closes, existing target sessions are revoked, and recovery progress remains durable. A failed partial target stays isolated.
For PostgreSQL, the final steps are deliberately separate:
hakopod database restore-plan TARGET_ID --artifact-id ARTIFACT_ID
hakopod database restore TARGET_ID --artifact-id ARTIFACT_ID --review-id REVIEW_ID --name recovered-orders
hakopod database inspect TARGET_ID --job-id JOB_ID --revision REVISION --name recovered-orders --inspected
hakopod database connection-plan TARGET_ID --application-id APP_ID --service api --variable DATABASE_URL --endpoint read_write
hakopod database connect TARGET_ID --review-id CONNECTION_REVIEW_ID --name orders-app
Inspect rows, binary values, schema objects and an application path before recording inspection. Connection replacement then queues the application deployment. Writes made to the source after capture are not in the archive; a real cutover needs a planned write boundary or reconciliation procedure. Keep the source until the recovered application is accepted.
Match the evidence to the claim
Observed roles and ready counts describe the sampled database. Resource requests describe reservations. A topology line is not measured replication traffic, and a volume’s requested size is not measured usage. Missing or stale data must remain visible as unavailable.
Likewise, labels describing zones and providers cannot prove independent failure domains. Members must live in one connected Kubernetes cluster with suitable private networking, control-plane access and storage. Two development nodes on one VM can test scheduling, native connections and process recovery. They cannot establish survival after losing an entire zone or provider.
Start a database evaluation with a disposable application and a restore rehearsal. Use the architecture guide to understand process ownership, security guide to verify the connection, and recovery guide to record what the restored application actually passed. Check the managed database guide for the current release and each engine’s remaining limits.