Observability & Telemetry (Grafana & Loki)
Okustera provides full-stack operational observability powered by Prometheus, Grafana Loki, and dynamic Grafana telemetry dashboards.
The platform automatically collects metrics, distributed traces, and container log streams across infrastructure hypervisors, serverless microservices, managed databases, and cloud workflows without requiring manual telemetry pipeline maintenance.
Observability Architecture
Dynamic Service-Aware Telemetry Dashboards
Okustera features an adaptive, service-aware telemetry synthesis engine. When you deploy applications or managed services via the Cloud Console, CLI, or Terraform, the platform backend automatically provisions dedicated, tenant-isolated Grafana dashboards with zero-trust metric scoping:
Supported Dynamic Dashboards
| Service Type | Key Metrics Captured | Visualization Panels |
|---|---|---|
| Serverless Functions | Invocations/sec, HTTP status code breakdown, execution duration percentiles (p50/p95/p99), replica scale counts. | Rate charts, error rate indicators, latency histograms, cold-start timelines. |
| Managed Databases | Active connection pool utilization, commit/rollback rate, query latency, replication lag, disk IOPS. | Pool saturation gauges, transaction throughput, storage capacity trends. |
| Compute & Virtual Machines | KVM hypervisor vCPU allocation, memory consumption, private SDN network RX/TX, disk read/write throughput. | Core utilization gauges, RAM consumption graphs, bandwidth meters. |
| Managed Workflows (Airflow) | DAG execution states, task success vs. failure rates, scheduler heartbeat, worker slot consumption. | DAG run Gantt charts, task duration trends, queue backlog trackers. |
Subtractive Pruning & Zero-Resource State
- Subtractive Pruning: When a workload is decommissioned (such as deleting a database cluster or deleting a serverless function), the synthesis engine automatically deletes the corresponding dashboard from the tenant folder to keep the workspace uncluttered.
- Zero-Resource State: For newly initialized workspaces with no active workloads, the telemetry portal displays interactive guidance and one-click deployment shortcuts to launch serverless functions, databases, compute instances, or workflow pipelines.
Platform Administration Telemetry
For platform administrators (omc_role == "admin"), Okustera provides pre-installed, cluster-wide infrastructure dashboards grouped in the platform-admin folder with administrative RBAC enforcement:
| Dashboard | Source & Scrape Endpoint | Telemetry Scope |
|---|---|---|
| OpenStack Overview | http://openstack-exporter.monitoring.svc.cluster.local:9103/metrics | Health, API responsiveness, hypervisors, and storage pools for Keystone, Nova, Neutron, Cinder, Glance, and Placement. |
| Apache APISIX API Gateway | http://apisix.ingress-apisix.svc.cluster.local:9091/apisix/prometheus/metrics | Ingress request rate (RPS), HTTP 2xx/4xx/5xx status ratios, upstream round-trip latency percentiles, and route analytics. |
| OpenFaaS Serverless Platform | http://gateway.openfaas.svc.cluster.local:8080/metrics | System-wide serverless throughput, function autoscaling down to zero, queue depth, and gVisor sandbox invocation overhead. |
| Apache Airflow Cluster | http://airflow-statsd.omc-system.svc.cluster.local:9102/metrics | Global scheduler heartbeats, executor process slots, DAG processing duration, and operator error patterns. |
| OpenStack Clouds | Prometheus OpenStack Node Exporter | Physical compute nodes, hypervisor CPU overcommit ratios, RAM allocation, and local SSD storage. |
| Loki Centralized Logs | http://loki-stack.monitoring.svc.cluster.local:3100 | Real-time multi-container log viewer across all Kubernetes control plane components and worker nodes. |
Okustera Cloud Console: Dedicated Dashboard Launchers
The Okustera Cloud Console provides a centralized Platform Grafana Telemetry (for administrators) and Service-Aware Telemetry Dashboards (for tenant members) interface:
- No Embedded Iframes: Rather than confining telemetry to cramped iframe containers that suffer from cross-origin cookie restrictions and header framing conflicts, each workload card acts as a dedicated launcher opening the Grafana dashboard in a new browser tab.
- Direct Path Routing: Dashboards open with pre-authenticated session tokens directly routed through the platform proxy (
/grafana/d/{dashboard-uid}). - Datasource Introspection: The portal automatically audits active Prometheus and Loki datasource health before rendering links.
Log Ingestion with Grafana Loki
All pods running across Okustera clusters stream container stdout and stderr logs to Loki via Promtail daemonsets.
LogQL Query Patterns
Query logs directly from the Okustera Cloud Console or Grafana Explore using LogQL:
# Filter error logs across backend services
{namespace="omc-system", app="omc-backend"} |= "ERROR"
# Search for APISIX ingress 5xx status codes
{namespace="ingress-apisix", app="apisix"} | json | status >= 500
# Stream logs for a specific serverless function
{namespace="openfaas-fn", fn="document-processor"} |= "Exception"
Custom Metrics with Prometheus
Workloads can expose custom Prometheus metrics by adding standard annotations to Kubernetes Pod or Deployment manifests:
apiVersion: apps/v1
kind: Deployment
metadata:
name: billing-worker
namespace: tenant-production
spec:
template:
metadata:
annotations:
prometheus.io/scrape: "true"
prometheus.io/path: "/metrics"
prometheus.io/port: "8080"
spec:
containers:
- name: app
image: artifact.example.com/apps/billing-worker:v1.2.0
ports:
- containerPort: 8080
Prometheus automatically discovers and scrapes marked endpoints within 30 seconds of deployment.
Generative AI Observability & Tracing (Langfuse)
For generative AI and LLM workloads deployed via the Okustera AI Inference PaaS, the platform integrates self-hosted Langfuse:
- End-to-End Distributed Tracing: Captures prompt requests, reasoning chains, tool calls, vector database retrievals (Qdrant), and streaming token chunks via OpenTelemetry.
- Latency & Performance Profiling: Real-time tracking of Time-to-First-Token (TTFT), inter-token arrival latencies, and queue wait times across vLLM worker replicas.
- Token Usage & Spend Correlation: Automatically correlates prompt and completion token counts with tenant billing accounts and API keys.
- Prompt Versioning & Evaluation: Production prompt playground with A/B scoring and zero-downtime rollback controls.
Tenant Observability in the Cloud Portal
The Okustera Cloud Portal provides built-in telemetry tools designed for developers and DevOps teams:
1. Embedded Grafana with Single Sign-On (SSO)
Tenants can inspect their dedicated telemetry dashboards without logging into a separate service:
- In the Cloud Portal, navigate to Observability $\to$ Grafana (
/grafana). - The portal securely authenticates your user session via header-based SSO.
- Access your tenant workspace folder with pre-provisioned dashboards for your compute instances, PostgreSQL databases, serverless functions, and Airflow workflows.
2. Live Log Viewer & LogQL Search
Debug applications and microservices in real time without accessing container shells:
- Navigate to Observability $\to$ Logs (
/observability). - Use LogQL expressions to filter logs:
- By application:
{app="order-service"} - By level:
{namespace="tenant-prod"} |= "ERROR" - By regex:
{job="openfaas-fn"} |~ "timeout|exception"
- By application:
- Live logs stream directly into the browser with timestamping and log severity badges.
3. Real-Time Resource Telemetry
Inspect live performance metrics across your infrastructure:
- Compute Utilization: Real-time vCPU and memory consumption gauges.
- Network Throughput: RX/TX bandwidth meters for private and floating IP interfaces.
- Storage IOPS: Disk read/write rates and volume saturation.
- API Traffic: Ingress request volume, latency percentiles (p50/p95/p99), and HTTP status codes (2xx/4xx/5xx).
REST API Reference
GET /api/v1/observability/logs— Query Grafana Loki log streams with LogQL expressions and time ranges.GET /api/v1/observability/timeseries/{metric_type}— Retrieve historical time-series data for CPU, RAM, Network, or Disk.GET /api/v1/observability/dashboard— Retrieve tenant-scoped telemetry summary.GET /api/v1/grafana/info— Retrieve tenant Grafana workspace URL and SSO endpoint.