Skip to main content

Database Version Upgrades & Maintenance

Okustera Database-as-a-Service (DBaaS) provides declarative, self-service version upgrades and automated maintenance for tenant database clusters.

Whether applying security patches, upgrading minor versions with zero downtime, or transitioning between major versions, Okustera protects your workloads with automated 5-point pre-flight safety gates, in-place rolling execution, and automatic rollback safety nets.


1. Upgrade Categorization & Strategy Matrix​

Database engine upgrades fall into two operational tiers based on data compatibility and downtime impact:

Upgrade TierDescription & ScopeSupported EnginesApplication ImpactUpgrade Mechanism
Minor / Patch UpgradeBug fixes, security patches, backward-compatible features (e.g., PostgreSQL 16.2 → 16.5, Qdrant 1.9.0 → 1.10.0, Valkey 7.2.4 → 7.2.5).PostgreSQL, Valkey, Qdrant, MySQL, MongoDBZero Downtime ($< 1$s graceful leader switchover for writes, 0s read drop)Automated in-place rolling restart across cluster replicas.
Major Version UpgradeBreaking storage format changes, feature jumps (e.g., PostgreSQL 15 → 16 → 17, MySQL 8.0 → 8.4).PostgreSQL, MySQL, MongoDBNear-Zero Downtime ($< 5$s cutover window)Declarative Blue-Green cluster provisioning with automated bootstrap import.

2. Automated 5-Point Pre-Flight Safety Gates​

Before any upgrade begins, Okustera automatically evaluates 5 mandatory pre-flight safety checks against your cluster. If any check fails, the upgrade is blocked to ensure your data and application availability are never compromised:

Safety Check Verification Criteria​

GateCheck NameVerification RuleTenant Protection
1Quorum & Replica HealthThe primary instance and all standby replicas must report status Running and readiness Ready. Minimum cluster quorum must be intact.Prevents rolling restarts on an already degraded cluster, eliminating the risk of accidental quorum loss.
2Replication Lag & BuffersStreaming replication lag to all standby replicas must be exactly 0 bytes (zero unapplied WAL/binlog).Guarantees that promoting a standby replica during switchover will not drop or delay uncommitted transactions.
3Storage HeadroomPersistent storage volume utilization must be strictly below $80%$ (at least $20%$ unallocated headroom).Prevents out-of-disk failures during switchover caused by WAL log accumulation, temporary sorting tables, or rollback journaling.
4Valid Backup SnapshotA full physical backup and WAL archive cut must have completed and verified within the last 15 minutes.Ensures an instantaneous recovery fallback exists prior to modifying cluster components.
5Concurrency LockA distributed cluster lock is acquired. No concurrent volume expansions, restore operations, or maintenance tasks may be running.Eliminates race conditions between simultaneous infrastructure management actions.

[!TIP] Emergency Override If an urgent security patch must be applied to an instance experiencing non-critical warnings, tenant administrators can pass skip_preflight=true via the API or toggle Emergency Override in the Cloud Portal. All overrides are recorded in the tenant audit log.


3. The 7-Stage Upgrade Lifecycle & Automated Rollback​

Minor and patch updates run through an automated 7-stage state machine that orchestrates sequential container patching, health verification, and rollback safeguards:

Stage Details​

  1. PENDING: Upgrade job is registered in the tenant database catalog.
  2. PREFLIGHT: The 5 automated safety checks are executed against the cluster.
  3. SNAPSHOT: An on-demand Ceph RBD storage snapshot and WAL cut are generated.
  4. ROLLING_RESTART:
    • Standby replica pods are patched and restarted sequentially.
    • Streaming replication synchronization is verified after each pod restarts.
    • A graceful switchover promotes a synchronized, upgraded standby to become the new primary.
    • The former primary pod is restarted with the new image and rejoins as a healthy standby.
  5. POST_CHECK: Synthetic read/write canary probes test connectivity and database engine responsiveness.
  6. COMPLETED: Cluster is verified healthy on the new version; maintenance lock is released.
  7. ROLLING_BACK / ROLLED_BACK: If any error occurs or canary probes fail and Auto-Rollback is enabled, the operator automatically reverts all pods to the prior image tag and re-establishes streaming replication.

4. How to Upgrade Your Database​

Method A: Cloud Portal UI​

  1. Navigate to Databases in the Okustera Cloud Portal.
  2. When a certified update is available for your cluster, an Upgrade Available badge appears next to the current version (e.g. current: v16.2, available: v16.5).
  3. Click the Upgrade button in the cluster table to open the upgrade modal:
    • Target Version: Select from certified versions available for your engine.
    • Upgrade Tier: View whether the upgrade is a Minor Rolling (zero downtime) or Major Migration (blue-green).
    • Pre-Flight Health Inspection: View the real-time status of all 5 safety gates.
    • Auto-Rollback Protection: Checkbox (enabled by default) to automatically revert if post-upgrade health checks fail.
  4. Click Confirm Upgrade. The modal displays real-time progress as the cluster transitions through each stage (PREFLIGHT $\to$ SNAPSHOT $\to$ ROLLING_RESTART $\to$ COMPLETED).

Method B: DBaaS REST API​

Step 1: Discover Available Upgrade Candidates​

Query certified upgrade targets for your running cluster:

curl -X GET \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
https://portal.okustera.com/api/v1/dbaas/clusters/postgres/production/my-db/upgrade-candidates

Response:

{
"current_version": "16.2",
"candidates": [
{
"version": "16.5",
"tier": "minor_rolling",
"is_recommended": true,
"requires_downtime": false,
"breaking_changes": null,
"release_notes_url": "https://www.postgresql.org/docs/release/16.5/"
},
{
"version": "17.0",
"tier": "major_migration",
"is_recommended": false,
"requires_downtime": true,
"breaking_changes": "PostgreSQL 17 requires Blue-Green migration due to on-disk catalog changes.",
"release_notes_url": "https://www.postgresql.org/docs/release/17.0/"
}
]
}

Step 2: Validate Pre-Flight Safety Checks​

Run pre-flight checks before scheduling the upgrade:

curl -X GET \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
"https://portal.okustera.com/api/v1/dbaas/clusters/postgres/production/my-db/preflight?target_version=16.5"

Response:

{
"passed": true,
"cluster_name": "my-db",
"namespace": "production",
"target_version": "16.5",
"checks": [
{"name": "quorum_and_pod_health", "passed": true, "details": "All 3 pods Ready and healthy."},
{"name": "replication_lag", "passed": true, "details": "Lag: 0 bytes."},
{"name": "storage_headroom", "passed": true, "details": "Disk utilization: 42% (58% free headroom)."},
{"name": "valid_backup_snapshot", "passed": true, "details": "Physical snapshot verified 6 minutes ago."},
{"name": "concurrency_lock", "passed": true, "details": "No conflicting jobs active."}
]
}

Step 3: Trigger the Upgrade​

Submit the upgrade request:

curl -X POST \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{
"target_version": "16.5",
"auto_rollback": true,
"skip_preflight": false
}' \
https://portal.okustera.com/api/v1/dbaas/clusters/postgres/production/my-db/upgrade

Response (202 Accepted):

{
"job_id": "job-upgrade-pg165-001",
"cluster_name": "my-db",
"namespace": "production",
"engine": "postgres",
"target_version": "16.5",
"stage": "PENDING",
"created_at": "2026-09-28T14:30:00Z"
}

Step 4: Track Upgrade Progress​

Monitor stage transitions:

curl -X GET \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
https://portal.okustera.com/api/v1/dbaas/jobs/job-upgrade-pg165-001

Response:

{
"job_id": "job-upgrade-pg165-001",
"cluster_name": "my-db",
"target_version": "16.5",
"stage": "COMPLETED",
"auto_rollback": true,
"started_at": "2026-09-28T14:30:02Z",
"completed_at": "2026-09-28T14:32:15Z"
}

Step 5: (Optional) Initiate Manual Rollback​

If application compatibility issues arise post-upgrade:

curl -X POST \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
https://portal.okustera.com/api/v1/dbaas/jobs/job-upgrade-pg165-001/rollback

Method C: Infrastructure as Code (Terraform)​

You can manage database versions declaratively using the Okustera Terraform provider:

resource "okustera_database_postgresql" "production_db" {
name = "production-core-db"
namespace = "production"
instances = 3
version = "16.5" # Target engine version
storage_size = "100Gi"
storage_class = "ceph-rbd"
database_name = "app"
username = "app_user"
}

Applying terraform apply triggers the pre-flight checks and initiates the rolling update without dropping active database connections.


5. Engine-Specific Upgrade Mechanisms​

5.1 Managed PostgreSQL (CloudNativePG)​

  • Minor Version Updates: In-place rolling update. Standby instances restart with the new image tag. Once streaming sync is confirmed, a graceful switchover (pg_ctl promote) switches the leader with sub-second write disruption and zero read drop.
  • Major Version Updates: Blue-Green cluster provisioning. A new cluster running the target major version (e.g., PostgreSQL 17) is provisioned with initdb.import from the source cluster. Once data synchronization is confirmed, traffic is redirected at the service DNS layer.

5.2 Valkey & Redis (Spotahome Sentinel)​

  • Replicas are patched and restarted sequentially.
  • Redis/Valkey Sentinel initiates an automated leader election, promoting an upgraded replica.
  • The former master node is restarted with the new engine image and rejoins the cluster as a replica.

5.3 Qdrant Vector Database (Vector DB)​

  • Qdrant maintains backward compatibility across minor releases. Segment files and Raft logs on NVMe Ceph RBD remain valid.
  • Kubernetes StatefulSet initiates a reverse rolling update (Pod $N-1$ down to Pod $0$).
  • Readiness probes (/readyz) hold client traffic until all vector indices are memory-mapped and verified.

5.4 MySQL / Galera (Percona XtraDB Cluster)​

  • Target node is placed in maintenance drain mode.
  • ProxySQL dynamically isolates incoming SQL traffic, directing transactions to remaining healthy nodes.
  • The drained node restarts on the new version and synchronizes state via Incremental State Transfer (IST).
  • The remaining nodes are patched sequentially.

5.5 MongoDB (Percona Server for MongoDB)​

  • Secondary members are upgraded and restarted sequentially.
  • The replica set primary executes a graceful rs.stepDown(), electing an upgraded secondary as the new primary.
  • The former primary is upgraded and rejoins as a secondary member.

6. Disaster Recovery & Rollback Protocols​

In addition to automated rolling rollback, Okustera provides Point-in-Time Recovery safety nets:

Point-in-Time Recovery (PITR)​

If an application incompatibility is discovered hours after a completed major or minor upgrade:

  • Continuous Write-Ahead Log (WAL) streaming guarantees an RPO $< 1$ minute.
  • Restore a parallel recovery cluster pointing to the exact timestamp prior to the upgrade:
    apiVersion: postgresql.cnpg.io/v1
    kind: Cluster
    metadata:
    name: db-recovered
    namespace: production
    spec:
    instances: 3
    storage:
    size: 100Gi
    storageClass: block-rbd1
    bootstrap:
    recovery:
    source: production-core-db
    recoveryTarget:
    targetTime: "2026-09-28T14:29:00Z"

Instantaneous Storage Snapshots​

Prior to scheduled maintenance, automated Ceph RBD storage snapshots (block-rbd1) are recorded. If severe corruption occurs, volumes can be restored instantly without network data transfer.


7. Service Level Agreements (SLA) & Commitments​

Operational MetricSLA CommitmentHow It Is Enforced
Minor Upgrade Write Disruption$< 1$ secondSub-second graceful leader switchover promoted via Operator API.
Minor Upgrade Read Disruption0 secondsHealthy standby load balancing via -ro service endpoint.
Major Upgrade Cutover Window$< 5$ secondsService endpoint DNS redirect upon verified blue-green bootstrap.
Recovery Point Objective (RPO)$< 1$ minuteContinuous WAL archiving streamed directly to Ceph S3 storage.
Recovery Time Objective (RTO)$< 10$ minutesAutomated Operator bootstrap restore from S3 backup target.
Regulatory Compliance StandardsISO 22301, ISO 27001, NIS2Business continuity, immutable audit trails, and GDPR Article 17 crypto-shredding.