Replication, Backups & Disaster Recovery
Okustera Database-as-a-Service (DBaaS) enforces enterprise-grade data durability, zero data loss replication, and automated Point-in-Time Recovery (PITR) across all tenant database clusters.
All backup policies and replication topologies comply with ISO/IEC 22301 (Business Continuity), ISO/IEC 27001/27017 (Data Protection & Cloud Security), and ISO/IEC 27018 / GDPR Chapter V (Data Sovereignty).
1. Architectural Overview
Every database cluster provisioned in Okustera is backed by an automated high-availability replication mesh and continuous backup streaming to Ceph S3 object storage:
2. Replication Mechanisms & High Availability
| Database Engine | Replication Architecture | Failover Time | Read/Write Separation |
|---|---|---|---|
| PostgreSQL 16 (CloudNativePG) | Physical streaming replication (sync/async standbys) | $< 5$ seconds | Built-in -rw (primary) and -ro (read replicas) services |
| MySQL 8.0 (Percona XtraDB) | Multi-master synchronous Galera replication | Instantaneous ($0$ lag) | Fronted by internal connection pooler |
| Valkey (Sentinel) | Master-replica replication with 3-node Sentinel quorum | $< 1$ second | Native Sentinel client routing |
| MongoDB (PSMDB) | Distributed replica sets with majority voting election | $< 3$ seconds | Primary writes; secondary read preference |
| Qdrant (Vector DB) | Distributed collection sharding with replication factor $\ge 2$ | $< 2$ seconds | Cluster-aware peer consensus |
3. Automated Backups & Continuous Archiving
3.1 Continuous WAL Streaming (PostgreSQL / Barman)
For PostgreSQL workloads, every committed transaction is captured in the Write-Ahead Log (pg_wal).
The operator continuously streams WAL files to your tenant S3 bucket via Barman Object Store:
- Base Backups: Scheduled daily physical snapshots of the data directory.
- WAL Streaming: Uploaded in near real-time (every 16MB WAL segment or every 60 seconds).
- RPO Guarantee: Recovery Point Objective is $< 1$ minute, allowing restoration up to the exact minute before an outage.
3.2 Scheduled Physical Backups (MySQL & MongoDB)
- Non-Locking Snapshots: Uses Percona XtraBackup to take hot physical backups without locking tables or interrupting live transactions.
- MongoDB Oplog Archiving: Percona Backup for MongoDB (PBM) captures continuous oplog slices for point-in-time document recovery.
4. Triggering Backups & Restores in the Cloud Portal
The Okustera Cloud Portal (/databases) provides graphical backup management:
1. Triggering an On-Demand Backup
- Navigate to Databases in the left sidebar.
- Select your active cluster (e.g.
production-postgres). - Click Backups tab $\to$ Create Backup.
- Enter a descriptive label (e.g.
pre-v2-schema-migration). - Click Start Backup. The operator captures a fresh base backup and syncs WAL files to S3 immediately.
2. Restoring to a Point-in-Time (PITR)
Scenario: An application bug executed a malformed query at 2026-09-15 14:35:00 UTC, corrupting production data.
- In the database details view, click Restore / Clone Cluster.
- Cluster Name: Enter a name for the new restored cluster (e.g.
production-db-restored). - Recovery Target: Select Point-in-Time (PITR).
- Specify target timestamp:
2026-09-15 14:34:59 UTC(the exact second before corruption). - Click Launch Restored Cluster.
- CloudNativePG automatically provisions fresh pods, fetches the base backup from S3, replays the WAL segments up to
14:34:59 UTC, and marks the new cluster ready.
5. Declarative Point-in-Time Recovery via YAML
You can also trigger a point-in-time restore declaratively using Kubernetes manifests:
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: production-db-restored
namespace: default
spec:
instances: 3
storage:
size: 100Gi
storageClass: ceph-nvme
bootstrap:
recovery:
source: production-db
recoveryTarget:
targetTime: "2026-09-15 14:34:59.000000+00"
backup:
name: production-db-latest
6. RPO & RTO Service Level Agreement (SLA) Matrix
- Recovery Point Objective (RPO): Maximum acceptable data loss duration.
- Recovery Time Objective (RTO): Maximum duration to restore services after a disaster.
| Database Engine | Protection Mechanism | Backup Target | Target RPO | Target RTO |
|---|---|---|---|---|
| PostgreSQL 16 | Barman S3 WAL streaming + daily base | Encrypted S3 Bucket | $< 1$ minute (PITR) | $< 10$ minutes |
| MySQL 8.0 (PXC) | Percona XtraBackup scheduled snapshots | Encrypted S3 Bucket | $1$ hour | $< 15$ minutes |
| MongoDB (PSMDB) | Percona Backup for MongoDB (PBM) oplog | Encrypted S3 Bucket | $< 5$ minutes | $< 15$ minutes |
| Valkey Sentinel | Ceph NVMe RDB snapshots + AOF | Persistent Volume + S3 | $< 15$ minutes | $< 5$ minutes |
| Qdrant Vector DB | Point-in-time snapshot API | Encrypted S3 Bucket | $1$ hour | $< 10$ minutes |
7. Security, Immutability & ISO Standards Alignment
7.1 S3 Object Lock (WORM - Write Once, Read Many)
Backup storage buckets have S3 Object Lock enabled in Compliance Mode with a mandatory retention window (e.g., 30–90 days). In Compliance Mode:
- No user, including tenant admins and cloud root administrators, can delete or overwrite a backup object before its retention expiration date.
- Protected against compromised credentials or malicious ransomware actions.
7.2 Envelope Encryption via Barbican KMS
- All database backups are encrypted at rest using unique, per-tenant AES-256 keys managed by OpenStack Barbican KMS.
- Master key escrow is kept out-of-band to ensure host destruction does not destroy decryption keys.
7.3 GDPR Right to Erasure via Crypto-Shredding
Under GDPR Article 17 and ISO/IEC 27018, tenants possess the right to be forgotten. Because WORM-locked backups cannot be selectively modified, tenant offboarding triggers the deletion of that tenant's unique encryption key in Barbican (crypto-shredding). All historical backup archives become cryptographically unreadable, fulfilling the legal erasure obligation without compromising backup chain integrity.
7.4 Zero Cross-Border Leakage (ISO/IEC 27018 & GDPR Chapter V)
Automated backup replication targets default strictly to data centers within the same national sovereign jurisdiction. Cross-border replication requires explicit tenant configuration and cryptographic verification.
8. Can a Tenant Opt Out or Refuse Backups?
Yes, tenants can refuse or opt out of backups for their own database clusters and workloads.
8.1 Workload Opt-Out (Dev / Ephemeral Clusters)
For development, CI/CD runners, transient data processing, or cost minimization:
- In the Cloud Portal: Toggle off Enable Automated Backups when provisioning the cluster, or set the backup retention window to
0days. - Declarative Manifests: Omit the
backup.barmanObjectStoreblock from yourpostgresql.cnpg.io/v1cluster manifest or delete the correspondingScheduleCRD.
8.2 Strict Zero-Retention & Privacy Mandates (GDPR Article 17 / HIPAA)
For workloads where regulatory rules forbid long-term data copies:
- Tenants can disable continuous WAL archiving entirely.
- For immutable WORM-locked backups that have already been created, Okustera supports Crypto-Shredding: revoking the tenant's dedicated Barbican KMS key mathematically destroys all past backup blocks, making them permanently unrecoverable.
8.3 SLA & Recovery Waiver
[!WARNING] Opting out of backups waives Okustera's Disaster Recovery SLA (RPO $< 1$ min, RTO $< 10$ mins) for that database. If an operator accidentally executes
DROP DATABASEor an application bug corrupts state, catastrophic data loss cannot be reversed.Note: While tenant database backups can be disabled, platform-level control plane resilience (Undercloud
etcdnode health) remains active to maintain cluster network and hypervisor routing. Platform backups never store tenant application payload data.
9. Backup Pricing & Cost Transparency
Backup storage in Okustera is metered transparently with no hidden egress or restoration fees.
9.1 Where Pricing is Displayed
- Cloud Portal (
/billing):- Navigate to Billing in the sidebar $\to$ Rate Card & Cloud Pricing tab.
- View the active rate card for Ceph S3 Object Storage and Managed Databases.
- Cost Breakdown & Invoices:
- Cost Breakdown tab shows current Month-to-Date spend categorized under Storage & DBaaS.
- Invoices tab provides downloadable PDF receipts itemizing exact backup storage usage (e.g.,
Ceph RGW S3 Standard Storage (15 GB) - $0.38).
- API & Terraform:
- Query the live rate sheet via
GET /api/v1/billing/pricing. - Inspect rates in Terraform via
data "omc_pricing_rates" "current" {}.
- Query the live rate sheet via
9.2 Unit Rates Reference
As documented in the FinOps & Billing Guide:
- Backup Storage (Ceph RGW S3): $0.015 – $0.025 / GB-month ($0.000034 / GB-hour) for compressed base dumps and continuous WAL archives.
- Persistent Volume Snapshots (Ceph RBD NVMe): $0.08 – $0.10 / GB-month ($0.00014 / GB-hour).
- Managed Database Instances (CloudNativePG): Includes automated failover, Barman continuous WAL streaming, and backup orchestration.