Database Version Upgrades & Maintenance
Okustera Database-as-a-Service (DBaaS) provides declarative, self-service version upgrades and automated maintenance for tenant database clusters.
Whether applying security patches, upgrading minor versions with zero downtime, or transitioning between major versions, Okustera protects your workloads with automated 5-point pre-flight safety gates, in-place rolling execution, and automatic rollback safety nets.
1. Upgrade Categorization & Strategy Matrix
Database engine upgrades fall into two operational tiers based on data compatibility and downtime impact:
| Upgrade Tier | Description & Scope | Supported Engines | Application Impact | Upgrade Mechanism |
|---|---|---|---|---|
| Minor / Patch Upgrade | Bug fixes, security patches, backward-compatible features (e.g., PostgreSQL 16.2 → 16.5, Qdrant 1.9.0 → 1.10.0, Valkey 7.2.4 → 7.2.5). | PostgreSQL, Valkey, Qdrant, MySQL, MongoDB | Zero Downtime ($< 1$s graceful leader switchover for writes, 0s read drop) | Automated in-place rolling restart across cluster replicas. |
| Major Version Upgrade | Breaking storage format changes, feature jumps (e.g., PostgreSQL 15 → 16 → 17, MySQL 8.0 → 8.4). | PostgreSQL, MySQL, MongoDB | Near-Zero Downtime ($< 5$s cutover window) | Declarative Blue-Green cluster provisioning with automated bootstrap import. |
2. Automated 5-Point Pre-Flight Safety Gates
Before any upgrade begins, Okustera automatically evaluates 5 mandatory pre-flight safety checks against your cluster. If any check fails, the upgrade is blocked to ensure your data and application availability are never compromised:
Safety Check Verification Criteria
| Gate | Check Name | Verification Rule | Tenant Protection |
|---|---|---|---|
| 1 | Quorum & Replica Health | The primary instance and all standby replicas must report status Running and readiness Ready. Minimum cluster quorum must be intact. | Prevents rolling restarts on an already degraded cluster, eliminating the risk of accidental quorum loss. |
| 2 | Replication Lag & Buffers | Streaming replication lag to all standby replicas must be exactly 0 bytes (zero unapplied WAL/binlog). | Guarantees that promoting a standby replica during switchover will not drop or delay uncommitted transactions. |
| 3 | Storage Headroom | Persistent storage volume utilization must be strictly below $80%$ (at least $20%$ unallocated headroom). | Prevents out-of-disk failures during switchover caused by WAL log accumulation, temporary sorting tables, or rollback journaling. |
| 4 | Valid Backup Snapshot | A full physical backup and WAL archive cut must have completed and verified within the last 15 minutes. | Ensures an instantaneous recovery fallback exists prior to modifying cluster components. |
| 5 | Concurrency Lock | A distributed cluster lock is acquired. No concurrent volume expansions, restore operations, or maintenance tasks may be running. | Eliminates race conditions between simultaneous infrastructure management actions. |
[!TIP] Emergency Override If an urgent security patch must be applied to an instance experiencing non-critical warnings, tenant administrators can pass
skip_preflight=truevia the API or toggle Emergency Override in the Cloud Portal. All overrides are recorded in the tenant audit log.
3. The 7-Stage Upgrade Lifecycle & Automated Rollback
Minor and patch updates run through an automated 7-stage state machine that orchestrates sequential container patching, health verification, and rollback safeguards:
Stage Details
PENDING: Upgrade job is registered in the tenant database catalog.PREFLIGHT: The 5 automated safety checks are executed against the cluster.SNAPSHOT: An on-demand Ceph RBD storage snapshot and WAL cut are generated.ROLLING_RESTART:- Standby replica pods are patched and restarted sequentially.
- Streaming replication synchronization is verified after each pod restarts.
- A graceful switchover promotes a synchronized, upgraded standby to become the new primary.
- The former primary pod is restarted with the new image and rejoins as a healthy standby.
POST_CHECK: Synthetic read/write canary probes test connectivity and database engine responsiveness.COMPLETED: Cluster is verified healthy on the new version; maintenance lock is released.ROLLING_BACK/ROLLED_BACK: If any error occurs or canary probes fail and Auto-Rollback is enabled, the operator automatically reverts all pods to the prior image tag and re-establishes streaming replication.
4. How to Upgrade Your Database
Method A: Cloud Portal UI
- Navigate to Databases in the Okustera Cloud Portal.
- When a certified update is available for your cluster, an Upgrade Available badge appears next to the current version (e.g. current:
v16.2, available:v16.5). - Click the Upgrade button in the cluster table to open the upgrade modal:
- Target Version: Select from certified versions available for your engine.
- Upgrade Tier: View whether the upgrade is a
Minor Rolling(zero downtime) orMajor Migration(blue-green). - Pre-Flight Health Inspection: View the real-time status of all 5 safety gates.
- Auto-Rollback Protection: Checkbox (enabled by default) to automatically revert if post-upgrade health checks fail.
- Click Confirm Upgrade. The modal displays real-time progress as the cluster transitions through each stage (
PREFLIGHT$\to$SNAPSHOT$\to$ROLLING_RESTART$\to$COMPLETED).
Method B: DBaaS REST API
Step 1: Discover Available Upgrade Candidates
Query certified upgrade targets for your running cluster:
curl -X GET \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
https://portal.okustera.com/api/v1/dbaas/clusters/postgres/production/my-db/upgrade-candidates
Response:
{
"current_version": "16.2",
"candidates": [
{
"version": "16.5",
"tier": "minor_rolling",
"is_recommended": true,
"requires_downtime": false,
"breaking_changes": null,
"release_notes_url": "https://www.postgresql.org/docs/release/16.5/"
},
{
"version": "17.0",
"tier": "major_migration",
"is_recommended": false,
"requires_downtime": true,
"breaking_changes": "PostgreSQL 17 requires Blue-Green migration due to on-disk catalog changes.",
"release_notes_url": "https://www.postgresql.org/docs/release/17.0/"
}
]
}
Step 2: Validate Pre-Flight Safety Checks
Run pre-flight checks before scheduling the upgrade:
curl -X GET \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
"https://portal.okustera.com/api/v1/dbaas/clusters/postgres/production/my-db/preflight?target_version=16.5"
Response:
{
"passed": true,
"cluster_name": "my-db",
"namespace": "production",
"target_version": "16.5",
"checks": [
{"name": "quorum_and_pod_health", "passed": true, "details": "All 3 pods Ready and healthy."},
{"name": "replication_lag", "passed": true, "details": "Lag: 0 bytes."},
{"name": "storage_headroom", "passed": true, "details": "Disk utilization: 42% (58% free headroom)."},
{"name": "valid_backup_snapshot", "passed": true, "details": "Physical snapshot verified 6 minutes ago."},
{"name": "concurrency_lock", "passed": true, "details": "No conflicting jobs active."}
]
}
Step 3: Trigger the Upgrade
Submit the upgrade request:
curl -X POST \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
-H "Content-Type: application/json" \
-d '{
"target_version": "16.5",
"auto_rollback": true,
"skip_preflight": false
}' \
https://portal.okustera.com/api/v1/dbaas/clusters/postgres/production/my-db/upgrade
Response (202 Accepted):
{
"job_id": "job-upgrade-pg165-001",
"cluster_name": "my-db",
"namespace": "production",
"engine": "postgres",
"target_version": "16.5",
"stage": "PENDING",
"created_at": "2026-09-28T14:30:00Z"
}
Step 4: Track Upgrade Progress
Monitor stage transitions:
curl -X GET \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
https://portal.okustera.com/api/v1/dbaas/jobs/job-upgrade-pg165-001
Response:
{
"job_id": "job-upgrade-pg165-001",
"cluster_name": "my-db",
"target_version": "16.5",
"stage": "COMPLETED",
"auto_rollback": true,
"started_at": "2026-09-28T14:30:02Z",
"completed_at": "2026-09-28T14:32:15Z"
}
Step 5: (Optional) Initiate Manual Rollback
If application compatibility issues arise post-upgrade:
curl -X POST \
-H "Authorization: Bearer <YOUR_API_TOKEN>" \
https://portal.okustera.com/api/v1/dbaas/jobs/job-upgrade-pg165-001/rollback
Method C: Infrastructure as Code (Terraform)
You can manage database versions declaratively using the Okustera Terraform provider:
resource "okustera_database_postgresql" "production_db" {
name = "production-core-db"
namespace = "production"
instances = 3
version = "16.5" # Target engine version
storage_size = "100Gi"
storage_class = "ceph-rbd"
database_name = "app"
username = "app_user"
}
Applying terraform apply triggers the pre-flight checks and initiates the rolling update without dropping active database connections.
5. Engine-Specific Upgrade Mechanisms
5.1 Managed PostgreSQL (CloudNativePG)
- Minor Version Updates: In-place rolling update. Standby instances restart with the new image tag. Once streaming sync is confirmed, a graceful switchover (
pg_ctl promote) switches the leader with sub-second write disruption and zero read drop. - Major Version Updates: Blue-Green cluster provisioning. A new cluster running the target major version (e.g., PostgreSQL 17) is provisioned with
initdb.importfrom the source cluster. Once data synchronization is confirmed, traffic is redirected at the service DNS layer.
5.2 Valkey & Redis (Spotahome Sentinel)
- Replicas are patched and restarted sequentially.
- Redis/Valkey Sentinel initiates an automated leader election, promoting an upgraded replica.
- The former master node is restarted with the new engine image and rejoins the cluster as a replica.
5.3 Qdrant Vector Database (Vector DB)
- Qdrant maintains backward compatibility across minor releases. Segment files and Raft logs on NVMe Ceph RBD remain valid.
- Kubernetes StatefulSet initiates a reverse rolling update (Pod $N-1$ down to Pod $0$).
- Readiness probes (
/readyz) hold client traffic until all vector indices are memory-mapped and verified.
5.4 MySQL / Galera (Percona XtraDB Cluster)
- Target node is placed in maintenance drain mode.
- ProxySQL dynamically isolates incoming SQL traffic, directing transactions to remaining healthy nodes.
- The drained node restarts on the new version and synchronizes state via Incremental State Transfer (IST).
- The remaining nodes are patched sequentially.
5.5 MongoDB (Percona Server for MongoDB)
- Secondary members are upgraded and restarted sequentially.
- The replica set primary executes a graceful
rs.stepDown(), electing an upgraded secondary as the new primary. - The former primary is upgraded and rejoins as a secondary member.
6. Disaster Recovery & Rollback Protocols
In addition to automated rolling rollback, Okustera provides Point-in-Time Recovery safety nets:
Point-in-Time Recovery (PITR)
If an application incompatibility is discovered hours after a completed major or minor upgrade:
- Continuous Write-Ahead Log (WAL) streaming guarantees an RPO $< 1$ minute.
- Restore a parallel recovery cluster pointing to the exact timestamp prior to the upgrade:
apiVersion: postgresql.cnpg.io/v1kind: Clustermetadata:name: db-recoverednamespace: productionspec:instances: 3storage:size: 100GistorageClass: block-rbd1bootstrap:recovery:source: production-core-dbrecoveryTarget:targetTime: "2026-09-28T14:29:00Z"
Instantaneous Storage Snapshots
Prior to scheduled maintenance, automated Ceph RBD storage snapshots (block-rbd1) are recorded. If severe corruption occurs, volumes can be restored instantly without network data transfer.
7. Service Level Agreements (SLA) & Commitments
| Operational Metric | SLA Commitment | How It Is Enforced |
|---|---|---|
| Minor Upgrade Write Disruption | $< 1$ second | Sub-second graceful leader switchover promoted via Operator API. |
| Minor Upgrade Read Disruption | 0 seconds | Healthy standby load balancing via -ro service endpoint. |
| Major Upgrade Cutover Window | $< 5$ seconds | Service endpoint DNS redirect upon verified blue-green bootstrap. |
| Recovery Point Objective (RPO) | $< 1$ minute | Continuous WAL archiving streamed directly to Ceph S3 storage. |
| Recovery Time Objective (RTO) | $< 10$ minutes | Automated Operator bootstrap restore from S3 backup target. |
| Regulatory Compliance Standards | ISO 22301, ISO 27001, NIS2 | Business continuity, immutable audit trails, and GDPR Article 17 crypto-shredding. |