Predictive Cluster Sentinel Daemon (omc-sentinel)
omc-sentinel is Okustera's autonomous, predictive cluster health engine. Unlike reactive monitoring tools that alert only after services crash, Sentinel continuously analyzes telemetry metrics, Loki log streams, Linux kernel resources, and Kubernetes states to detect failure precursors and resource depletion velocity before workloads are disrupted.
omc-sentinel is an internal diagnostic and failure-prediction daemon executed by cloud infrastructure operators across physical host hypervisors. For tenant applications, service metrics, and container log streams, please refer to the Tenant Observability & Grafana Guide.
1. Architectural Overview
Sentinel correlates multi-layer signals across physical hypervisors, storage clusters, the OpenStack control plane, and Kubernetes tenant workloads:
2. Predictive Failure Signatures Detected
Sentinel tracks specific mathematical and semantic precursors that herald upcoming outages:
A. Storage Depletion Velocity (Burn Rate)
- Precursor Metric: Filesystem allocation rate (change in volume over time) calculated across consecutive inspection cycles.
- Predictive Rule: If disk usage exceeds 80% and the calculated velocity projects capacity exhaustion in under 24 hours (or under 4 hours for critical), Sentinel flags
PREDICTIVE_RISKbefore Kubelet hits its 85% eviction threshold.
B. Ceph Placement Group (PG) Inactive & Peering Delays
- Precursor Metric: Ceph
inactive_pgs_ratiowhile write operations (write_ops > 0) are active. - Forecast: Unpeered placement groups block incoming write I/O. Sentinel warns operators before virtual machines or databases experience I/O hangs.
C. Hypervisor Storage & Keyring Mismatches
- Precursor Log Signature:
StorageError: Could not determine disk usage/[errno 13] RADOS permission denied. - Forecast: Nova's Resource Tracker cannot determine hypervisor storage capacity, which causes Nova scheduler to reject new instance allocations or miscalculate cluster density.
D. Pod Flapping & Probe Acceleration
- Precursor Metric: Container restart acceleration rate (> 1.0 restart per hour) combined with transient liveness probe failures in the Kubernetes events stream.
- Forecast: Predicts impending pod crash loops caused by microservice memory leaks or message bus disconnections.
E. Database Connection Pool & Deadlock Saturation
- Precursor Log Signature:
too many connections,Lock wait timeout,lost connection to mysql,wsrep: failed to replicate. - Forecast: ProxySQL client pool starvation, predicting incoming HTTP 500/503 Service Unavailable errors across OpenStack APIs.
F. Messaging Broker Heartbeat Drops
- Precursor Log Signature:
connection.blocked,resource_limit_alarm,[Errno -2] Name or service not known. - Forecast: RabbitMQ backpressure or socket disconnection, predicting message queue timeouts in Nova, Cinder, and Neutron RPC calls.
3. Command Line Interface (CLI) Usage
The omc-sentinel command is installed globally at /usr/local/bin/omc-sentinel.
Run Instant Diagnostic Scan (--check)
Executes a full diagnostic and predictive analysis across all layers and prints a color-coded operational report:
omc-sentinel --check
Example Output
================================================================================
OpenCloud (OMC) Sentinel - Autonomous Health & Predictive Daemon
================================================================================
Timestamp: 2026-09-15 11:10:20
Health Score: 0% [CRITICAL FAILURE ACTIVE]
Diagnostics: 7 Total Findings Detected
--------------------------------------------------------------------------------
🚨 ACTIVE CRITICAL OUTAGES / FAILURES (1)
[1] 1 Pods In Failed/Pending State (K8s/Undercloud)
Evidence: kube-system/cilium-operator-677b54b8b-tm7vw (Pending: non-running)
Impact: Associated microservices are offline or failing health probes.
Remediation: Examine pod logs: `kubectl describe pod <pod-name> -n <namespace>`.
⚡ PREDICTIVE RISKS - SOMETHING WILL GO WRONG SOON (6)
[1] Active Pod Flapping / CrashLoop Velocity Detected (K8s/Undercloud)
Precursor: openstack/nova-conductor-6f68b68998-js7dp (262 lifetime restarts)
Forecast: Pod is intermittently crashing; underlying service has unhandled exceptions or failing probes.
Remediation: Check liveness/readiness probe logs and OOM kill events.
[2] High Error Storm in nova-compute (450 errors/10m) (Logs/openstack)
Precursor: Stream openstack/nova-compute emitted 450 error logs in last 10 minutes.
Forecast: Microservice nova-compute is encountering repeated execution exceptions or backend disconnections.
Remediation: Inspect recent logs: `kubectl logs -n openstack -l app.kubernetes.io/name=nova-compute --tail=50`.
[3] Nova Hypervisor RADOS Ceph Permission Failure (Storage/Nova)
Precursor: Stderr: '[errno 13] RADOS permission denied (error connecting to the cluster)'
Forecast: Nova Resource Tracker cannot determine hypervisor storage capacity; instance provisioning will fail.
Remediation: Check Ceph client.openstack or client.cinder keyring permissions in `nova-compute` container.
[4] Ceph Data Availability Risk (Inactive PGs) (Storage/Ceph)
Precursor: 12.9% of Ceph PGs are not active (undersized/peered).
Forecast: Writes to unpeered placement groups will block until OSD peering completes.
Remediation: Check OSD pod status and quorum: `kubectl get pods -n rook-ceph -l app=rook-ceph-osd`.
================================================================================
Real-Time Interactive Watch Dashboard (--watch)
Refreshes live system health every N seconds in a terminal dashboard (similar to top or watch):
omc-sentinel --watch --interval 10
Machine-Readable JSON Output (--json)
Emits pure structured JSON suitable for CI/CD pipelines, Prometheus Pushgateway, or external SIEM ingestion:
omc-sentinel --json | jq .
Example Output
{
"timestamp": "2026-09-15T11:12:54.015191+00:00",
"health_score": 75,
"status": "WARNING",
"findings": [
{
"timestamp": "2026-09-15T11:12:52.025978+00:00",
"category": "K8s/Undercloud",
"title": "Active Pod Flapping / CrashLoop Velocity Detected",
"severity": "PREDICTIVE_RISK",
"evidence": "openstack/nova-conductor-6f68b68998-js7dp (262 lifetime restarts)",
"impact": "Pod is intermittently crashing; underlying service has unhandled exceptions or failing probes.",
"remediation": "Check liveness/readiness probe logs and OOM kill events.",
"metric_data": {}
}
]
}
One-Line Summary (--summary)
Ideal for integration into shell prompt statuses, tmux status bars, or MOTD banners:
omc-sentinel --summary
OMC Cluster Health: 85% | Critical: 0 | Predictive Risks: 1 | Warnings: 2
4. Background Systemd Daemon
Sentinel is packaged with a dedicated systemd service unit to run continuously in the background.
Managing the Service
# Check running service status
sudo systemctl status omc-sentinel
# View real-time daemon logs via journald
sudo journalctl -u omc-sentinel -f -n 50
# Restart daemon after updating configuration
sudo systemctl restart omc-sentinel
Configuration File (omc_sentinel.conf)
The configuration file is located at infrastructure/scripts/omc_sentinel.conf:
{
"interval_seconds": 60,
"state_dir": "/var/log/omc-sentinel",
"prometheus_url": "http://prometheus-k8s.monitoring.svc.cluster.local:9090",
"loki_url": "http://loki-gateway.monitoring.svc.cluster.local:3100",
"webhook_url": "https://hooks.slack.com/services/T00/B00/XXXXX",
"auto_heal": false,
"thresholds": {
"cpu_load_warning_ratio": 1.2,
"cpu_load_critical_ratio": 2.0,
"mem_warning_pct": 85.0,
"mem_critical_pct": 92.0,
"disk_warning_pct": 80.0,
"disk_critical_pct": 90.0,
"disk_hours_to_exhaustion_warning": 24.0,
"disk_hours_to_exhaustion_critical": 4.0,
"log_error_storm_count_10m": 150,
"ceph_inactive_pg_warning_ratio": 0.10
}
}
Setting Up Webhook Alerts (Slack / Discord / Teams)
Set webhook_url in the configuration file or export the environment variable:
export ALERT_WEBHOOK_URL="https://discord.com/api/webhooks/YOUR_WEBHOOK_ID"
sudo systemctl restart omc-sentinel
When an issue with CRITICAL or PREDICTIVE_RISK severity is identified, Sentinel posts an immediate notification with the root cause and remediation command.
5. Automated Self-Healing (--auto-heal)
Sentinel supports safe, non-destructive automated remediation hooks:
# Execute check and automatically repair known safe network and socket anomalies
omc-sentinel --check --auto-heal
When enabled, Sentinel can automatically:
- Re-establish missing
veth-host/api0interfaces and OVS bridge mappings by invokingsetup_veth_bridge.sh. - Fix unprivileged container access permissions on
/var/run/openvswitch/db.sock(chmod 0666). - Invoke
omc_doctor.sh --fixto re-sync missing libvirt Ceph RBD secrets.
6. Integration Cheatsheet
| Task | Command |
|---|---|
| Run Diagnostic Check | omc-sentinel --check |
| Terminal Watch Dashboard | omc-sentinel --watch |
| Inspect JSON Output | omc-sentinel --json | jq . |
| View Past Alert History | tail -n 20 /var/log/omc-sentinel/alerts.jsonl | jq . |
| Check Systemd Service | sudo systemctl status omc-sentinel |
| Inspect Daemon Logs | sudo journalctl -u omc-sentinel -n 50 --no-pager |
| Trigger Auto-Healing | omc-sentinel --check --auto-heal |