Skip to main content

Replication, Backups & Disaster Recovery

Okustera Database-as-a-Service (DBaaS) enforces enterprise-grade data durability, zero data loss replication, and automated Point-in-Time Recovery (PITR) across all tenant database clusters.

All backup policies and replication topologies comply with ISO/IEC 22301 (Business Continuity), ISO/IEC 27001/27017 (Data Protection & Cloud Security), and ISO/IEC 27018 / GDPR Chapter V (Data Sovereignty).


1. Architectural Overview​

Every database cluster provisioned in Okustera is backed by an automated high-availability replication mesh and continuous backup streaming to Ceph S3 object storage:


2. Replication Mechanisms & High Availability​

Database EngineReplication ArchitectureFailover TimeRead/Write Separation
PostgreSQL 16 (CloudNativePG)Physical streaming replication (sync/async standbys)$< 5$ secondsBuilt-in -rw (primary) and -ro (read replicas) services
MySQL 8.0 (Percona XtraDB)Multi-master synchronous Galera replicationInstantaneous ($0$ lag)Fronted by internal connection pooler
Valkey (Sentinel)Master-replica replication with 3-node Sentinel quorum$< 1$ secondNative Sentinel client routing
MongoDB (PSMDB)Distributed replica sets with majority voting election$< 3$ secondsPrimary writes; secondary read preference
Qdrant (Vector DB)Distributed collection sharding with replication factor $\ge 2$$< 2$ secondsCluster-aware peer consensus

3. Automated Backups & Continuous Archiving​

3.1 Continuous WAL Streaming (PostgreSQL / Barman)​

For PostgreSQL workloads, every committed transaction is captured in the Write-Ahead Log (pg_wal).

The operator continuously streams WAL files to your tenant S3 bucket via Barman Object Store:

  • Base Backups: Scheduled daily physical snapshots of the data directory.
  • WAL Streaming: Uploaded in near real-time (every 16MB WAL segment or every 60 seconds).
  • RPO Guarantee: Recovery Point Objective is $< 1$ minute, allowing restoration up to the exact minute before an outage.

3.2 Scheduled Physical Backups (MySQL & MongoDB)​

  • Non-Locking Snapshots: Uses Percona XtraBackup to take hot physical backups without locking tables or interrupting live transactions.
  • MongoDB Oplog Archiving: Percona Backup for MongoDB (PBM) captures continuous oplog slices for point-in-time document recovery.

4. Triggering Backups & Restores in the Cloud Portal​

The Okustera Cloud Portal (/databases) provides graphical backup management:

1. Triggering an On-Demand Backup​

  1. Navigate to Databases in the left sidebar.
  2. Select your active cluster (e.g. production-postgres).
  3. Click Backups tab $\to$ Create Backup.
  4. Enter a descriptive label (e.g. pre-v2-schema-migration).
  5. Click Start Backup. The operator captures a fresh base backup and syncs WAL files to S3 immediately.

2. Restoring to a Point-in-Time (PITR)​

Scenario: An application bug executed a malformed query at 2026-09-15 14:35:00 UTC, corrupting production data.

  1. In the database details view, click Restore / Clone Cluster.
  2. Cluster Name: Enter a name for the new restored cluster (e.g. production-db-restored).
  3. Recovery Target: Select Point-in-Time (PITR).
  4. Specify target timestamp: 2026-09-15 14:34:59 UTC (the exact second before corruption).
  5. Click Launch Restored Cluster.
  6. CloudNativePG automatically provisions fresh pods, fetches the base backup from S3, replays the WAL segments up to 14:34:59 UTC, and marks the new cluster ready.

5. Declarative Point-in-Time Recovery via YAML​

You can also trigger a point-in-time restore declaratively using Kubernetes manifests:

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: production-db-restored
namespace: default
spec:
instances: 3
storage:
size: 100Gi
storageClass: ceph-nvme
bootstrap:
recovery:
source: production-db
recoveryTarget:
targetTime: "2026-09-15 14:34:59.000000+00"
backup:
name: production-db-latest

6. RPO & RTO Service Level Agreement (SLA) Matrix​

  • Recovery Point Objective (RPO): Maximum acceptable data loss duration.
  • Recovery Time Objective (RTO): Maximum duration to restore services after a disaster.
Database EngineProtection MechanismBackup TargetTarget RPOTarget RTO
PostgreSQL 16Barman S3 WAL streaming + daily baseEncrypted S3 Bucket$< 1$ minute (PITR)$< 10$ minutes
MySQL 8.0 (PXC)Percona XtraBackup scheduled snapshotsEncrypted S3 Bucket$1$ hour$< 15$ minutes
MongoDB (PSMDB)Percona Backup for MongoDB (PBM) oplogEncrypted S3 Bucket$< 5$ minutes$< 15$ minutes
Valkey SentinelCeph NVMe RDB snapshots + AOFPersistent Volume + S3$< 15$ minutes$< 5$ minutes
Qdrant Vector DBPoint-in-time snapshot APIEncrypted S3 Bucket$1$ hour$< 10$ minutes

7. Security, Immutability & ISO Standards Alignment​

7.1 S3 Object Lock (WORM - Write Once, Read Many)​

Backup storage buckets have S3 Object Lock enabled in Compliance Mode with a mandatory retention window (e.g., 30–90 days). In Compliance Mode:

  • No user, including tenant admins and cloud root administrators, can delete or overwrite a backup object before its retention expiration date.
  • Protected against compromised credentials or malicious ransomware actions.

7.2 Envelope Encryption via Barbican KMS​

  • All database backups are encrypted at rest using unique, per-tenant AES-256 keys managed by OpenStack Barbican KMS.
  • Master key escrow is kept out-of-band to ensure host destruction does not destroy decryption keys.

7.3 GDPR Right to Erasure via Crypto-Shredding​

Under GDPR Article 17 and ISO/IEC 27018, tenants possess the right to be forgotten. Because WORM-locked backups cannot be selectively modified, tenant offboarding triggers the deletion of that tenant's unique encryption key in Barbican (crypto-shredding). All historical backup archives become cryptographically unreadable, fulfilling the legal erasure obligation without compromising backup chain integrity.

7.4 Zero Cross-Border Leakage (ISO/IEC 27018 & GDPR Chapter V)​

Automated backup replication targets default strictly to data centers within the same national sovereign jurisdiction. Cross-border replication requires explicit tenant configuration and cryptographic verification.


8. Can a Tenant Opt Out or Refuse Backups?​

Yes, tenants can refuse or opt out of backups for their own database clusters and workloads.

8.1 Workload Opt-Out (Dev / Ephemeral Clusters)​

For development, CI/CD runners, transient data processing, or cost minimization:

  • In the Cloud Portal: Toggle off Enable Automated Backups when provisioning the cluster, or set the backup retention window to 0 days.
  • Declarative Manifests: Omit the backup.barmanObjectStore block from your postgresql.cnpg.io/v1 cluster manifest or delete the corresponding Schedule CRD.

8.2 Strict Zero-Retention & Privacy Mandates (GDPR Article 17 / HIPAA)​

For workloads where regulatory rules forbid long-term data copies:

  • Tenants can disable continuous WAL archiving entirely.
  • For immutable WORM-locked backups that have already been created, Okustera supports Crypto-Shredding: revoking the tenant's dedicated Barbican KMS key mathematically destroys all past backup blocks, making them permanently unrecoverable.

8.3 SLA & Recovery Waiver​

[!WARNING] Opting out of backups waives Okustera's Disaster Recovery SLA (RPO $< 1$ min, RTO $< 10$ mins) for that database. If an operator accidentally executes DROP DATABASE or an application bug corrupts state, catastrophic data loss cannot be reversed.

Note: While tenant database backups can be disabled, platform-level control plane resilience (Undercloud etcd node health) remains active to maintain cluster network and hypervisor routing. Platform backups never store tenant application payload data.


9. Backup Pricing & Cost Transparency​

Backup storage in Okustera is metered transparently with no hidden egress or restoration fees.

9.1 Where Pricing is Displayed​

  1. Cloud Portal (/billing):
    • Navigate to Billing in the sidebar $\to$ Rate Card & Cloud Pricing tab.
    • View the active rate card for Ceph S3 Object Storage and Managed Databases.
  2. Cost Breakdown & Invoices:
    • Cost Breakdown tab shows current Month-to-Date spend categorized under Storage & DBaaS.
    • Invoices tab provides downloadable PDF receipts itemizing exact backup storage usage (e.g., Ceph RGW S3 Standard Storage (15 GB) - $0.38).
  3. API & Terraform:
    • Query the live rate sheet via GET /api/v1/billing/pricing.
    • Inspect rates in Terraform via data "omc_pricing_rates" "current" {}.

9.2 Unit Rates Reference​

As documented in the FinOps & Billing Guide:

  • Backup Storage (Ceph RGW S3): $0.015 – $0.025 / GB-month ($0.000034 / GB-hour) for compressed base dumps and continuous WAL archives.
  • Persistent Volume Snapshots (Ceph RBD NVMe): $0.08 – $0.10 / GB-month ($0.00014 / GB-hour).
  • Managed Database Instances (CloudNativePG): Includes automated failover, Barman continuous WAL streaming, and backup orchestration.