Skip to main content

AI Inference PaaS & Model Foundry (vLLM, KubeRay, Qdrant & Kueue)

Operational Status (Phase 8 Baseline & Node 2 Multi-Model Topology Live)
  • Node 2 Dedicated AI Compute: High-throughput vLLM (vllm/vllm-openai-cpu:latest-x86_64) running with hardware AMD Zen 5 AVX-512 vector extensions and PagedAttention, offloading heavy token generation from the control plane.
  • Kev Admission Router: Sub-15ms System 1 classification and prompt injection guardrails (HTTP 403 Forbidden) routing requests across iac_devops, deep_reasoning, openstack_api, and malicious_prompt.
  • Bounded RAM Footprint: Strict allocation caps (--gpu-memory-utilization 0.35 / $\le$ 28 GB total budget) reserving 25+ GB host RAM for OpenStack Nova KVM hypervisors and Ceph storage.
  • Direct OpenStack SDK Route: Deterministic execution bypassing generative LLM hallucination for administrative cloud tasks.
  • Central Qdrant Vector RAG: Live on private ClusterIP network with unbuffered SSE streaming at https://ai.okustera.com.
Sovereign AI Security Gateway & Zero-Trust Perimeter

Looking for enterprise PII sanitization (Microsoft Presidio) and Model Context Protocol (MCP) tool governance behind Fortinet FortiGate firewalls? Check out the Sovereign AI Security Gateway Quickstart and live interactive console at https://shield.okustera.com.

Okustera AI Inference PaaS & Model Foundry transforms cloud infrastructure into an enterprise-grade, zero-license-tax AI Inference Cloud and Sovereign Model Garden.

Rather than bundling heavy, multi-purpose training frameworks, Phase 8 is hyper-focused on high-throughput production model serving, sovereign model provenance, dynamic multi-LoRA serving, enterprise vector RAG, asynchronous frontier MoE batch execution, agentic code sandboxing, and token-level FinOps governance.

Okustera AI Inference PaaS & Sovereign Model Foundry Console


High-Level Architecture​

The platform delivers an OpenAI-compatible cloud tier on top of the tenant Kubernetes workload plane, fully decoupled into an intelligent edge routing plane, real-time serving tier, vector retrieval engine, and zero-copy shared storage:


1. The Okustera Model Foundry: Sovereign Model Garden​

The Okustera Model Foundry is a sovereign, zero-license-tax alternative to hyperscaler model repositories (AWS Bedrock, Azure AI Foundry, Vertex AI Model Garden). It eliminates vendor lock-in and high API markups by hosting certified open weights directly on the tenant's private infrastructure.

Curated Foundation Catalog & Certified Hardware Blueprints​

Every model blueprint pairs vetted open weights with certified quantization formats, minimum GPU footprints, and tensor-parallel parameters:

CategoryRecommended Foundry ModelsQuantization FormatsHardware Blueprint (Min Spec)Tensor Parallel SplitContext Window
Admission & GuardrailKev 0.8B (jaredpalmer/kev)INT4 (ONNX Runtime)CPU (2 Threads, ~600MB RAM)None (Sub-15ms Fast Classifier)4k tokens
IaC & AutomationNimbus 4B v2.1 (Nimbus-Labs/Nimbus-4B-v2.1)W4A16 / AWQCPU AVX-512 (4 Threads, ~2.2GB RAM)TP=14k tokens
Deep Reasoning & DiagnosticsGemma 4 26B-A4B MoE (google/gemma-4-26b-a4b-it)W4A16 / AWQCPU AVX-512 (8 Threads, ~12.5GB RAM)TP=1 (MoE 4B Active)8k tokens
Reasoning & MathDeepSeek-R1-Distill-Qwen (8B / 32B), DeepSeek-R1-Distill-Llama (70B)FP8, AWQ, BF161x A10G (8B) / 4x A100 (70B)TP=1 (8B) / TP=4 (70B)32k – 64k tokens
General ChatLlama 3.1 / 3.2 (3B, 8B, 70B), Qwen 2.5 (7B, 14B, 32B)AWQ, GPTQ, BF16CPU AVX-512 (3B/7B) or 1x A10GTP=1 (7B/14B) / TP=4 (70B)8k – 128k tokens
Code GenerationQwen 2.5 Coder (7B, 14B, 32B), StarCoder2AWQ, BF161x A10G / 2x A10GTP=1 (7B) / TP=2 (14B)32k – 128k tokens
Multimodal (VLM)Qwen2-VL (7B), Llama 3.2 Vision (11B)AWQ, BF161x A10G (24GB VRAM)TP=18k – 32k tokens
Embeddings & RAGBGE Small/Large EN v1.5, Nomic Embed Text v1.5FP16, INT8, ONNXCPU (TEI / FastEmbed) or 1x T4TP=18k tokens
Frontier MoE BatchDeepSeek-V4 Flash (284B), GLM-5.2 (744B)INT4, FP8Colibrì NVMe Tier + 16GB RAMMoE Expert Streaming8k – 32k tokens

Multi-Source Model Ingestion Pipeline​

  1. Pre-Warmed CephFS Shared Cache (/models/foundry/): Curated foundation models are pre-loaded onto the cluster's high-speed CephFS RWX volume. vLLM worker pods mount the volume read-only, reducing cold-start spin-up time from 20 minutes down to under 15 seconds via zero-copy Linux page cache memory mapping.
  2. Hugging Face Hub Pull-Through Mirror: Tenants specify any public repository ID (e.g., Qwen/Qwen2.5-7B-Instruct). The Foundry pulls weights directly to CephFS, caching layers locally to prevent egress fees and avoid rate limits.
  3. OCI Model Packages via Artifact Keeper (Phase 7 Integration): Fine-tuned weights and LoRA adapters can be packaged as OCI Artifacts using ORAS or ModelKit and pushed directly to Okustera's Universal Artifact Registry (artifacts.okustera.com).
  4. Ceph RGW S3 Custom Model Buckets: Tenants can upload custom weights directly via S3 API keys into isolated buckets (s3://omc-models/<tenant-id>/).

Automated Security Gatekeeper (omc-security-guardian)​

Before any model enters the catalog or is scheduled for serving, it passes through automated security validation:

  • SafeTensors Mandatory Rule: Untrusted PyTorch pickle files (.bin, .pt, pickle) can execute arbitrary code upon deserialization (CVE-2024-34359). The Foundry rejects pickled weights and strictly enforces .safetensors.
  • Cryptographic Hash Verification: Validates file hashes against upstream commit manifests to prevent supply-chain tampering.
  • License Governance: Automatically tags the model license (Apache 2.0, MIT, Llama Community License) to prevent unauthorized commercial deployment.

2. Multi-Model Admission Routing & Guardrails (Kev & System 1)​

To protect large generative models from prompt injection, minimize inference latency, and achieve deterministic reliability for cloud operations, Okustera implements a two-tier System 1 / System 2 admission dispatch topology running on Node 2:

The 4 Execution Pathways​

  1. Prompt Injection & Security Guardrails (malicious_prompt): Kev inspects incoming prompts at the perimeter for jailbreak attempts, system prompt overrides, and destructive commands. Malicious inputs are dropped immediately with HTTP 403 Forbidden before consuming backend compute cycles or entering logging tables.
  2. Deterministic OpenStack SDK Bypass (openstack_api): Administrative queries (checking tenant quotas, listing instances, inspecting flavors, rebooting servers) are routed directly to OpenStack APIs rather than generative LLMs. This guarantees 100% deterministic accuracy with zero generative hallucination.
  3. IaC & Automation Specialist (iac_devops): Dispatched to Nimbus-4B-v2.1, a 4B parameter transformer specialized in Terraform HCL, Ansible playbooks, and Kubernetes CRDs.
  4. Deep Reasoning & Diagnostics (deep_reasoning): Dispatched to Gemma-4-26B-A4B, a Mixture-of-Experts engine activating 4B parameters out of 26B total for complex multi-turn root-cause analysis, log postmortems, and architecture design.

Host Memory Bounding Strategy​

On bare-metal compute nodes where AI inference co-exists with OpenStack Nova KVM hypervisors, AI memory must operate within a bounded ceiling:

  • AVX-512 VNNI 4-Bit Quantization: W4A16 / AWQ quantization shrinks Gemma-4-26B from ~52 GB down to ~12.5 GB and Nimbus-4B to ~2.2 GB.
  • FP8 Dynamic KV Cache: --kv-cache-dtype fp8 cuts attention cache consumption by 50%.
  • Host Allocation Cap: Configured with --gpu-memory-utilization 0.35 in vLLM, bounding total AI host RAM to $\le$ 28 GB and permanently reserving 25+ GB host RAM for tenant virtual machines and Ceph OSDs.

3. Dynamic Multi-LoRA Serving: 50+ Custom Models on 1 Base GPU​

To eliminate GPU sprawl and prohibitive infrastructure costs, Okustera implements Dynamic Multi-LoRA Serving:

  • The Architecture: vLLM hosts a single underlying base model in GPU VRAM (e.g., Meta-Llama-3.1-70B-Instruct).
  • On-the-Fly Switching: Tenants upload lightweight domain-specific LoRA adapters (10MB–200MB) to Ceph S3 or Artifact Keeper.
  • Sub-10ms Scheduling: When a request specifies "model": "llama-70b-legal" or "model": "llama-70b-finance", vLLM's multi-LoRA scheduler (--enable-lora --max-loras 64) dynamically loads and applies the adapter weights in GPU memory in under 10 milliseconds.
  • Cost Impact: Up to 64 distinct customer fine-tunes run concurrently on a single shared GPU instance, slashing infrastructure spend by over 90%.

4. Asynchronous Frontier MoE Batch Tier (Kubernetes Kueue & Colibrì)​

Frontier models exceeding 200B+ parameters (such as DeepSeek-V4 Flash at 284B and GLM-5.2 at 744B) typically require multi-hundred-thousand-dollar GPU clusters (8x H100) when run in standard full-residency VRAM engines. To enable sovereign access to frontier reasoning on commodity or mixed hardware, Okustera establishes an Asynchronous Frontier Batch Tier:

  1. Job Queuing & Fair Sharing with Kubernetes Kueue: Manages multi-tenant batch inference queues (ResourceFlavor, ClusterQueue, LocalQueue), dynamically allocating GPU/CPU quota and preventing cluster overload.
  2. Colibrì NVMe Expert Streaming: Containerized Colibrì workers load base weights into host memory while streaming sparse Mixture-of-Experts (MoE) layers directly from high-speed NVMe storage on demand.
  3. Execution Profile: Ideal for synthetic data generation, document processing, code analysis, and model evaluations where immediate conversational latency is not required.

5. Vector DBaaS & Enterprise RAG Infrastructure (Qdrant)​

Okustera incorporates dedicated vector similarity search powered by Qdrant (written in Rust):

  • On-Disk NVMe Vector Storage: Vectors and payloads are stored on disk with memory-mapped files (mmap), keeping only the HNSW graph index in RAM.
  • Scalar Quantization: Compresses 32-bit floating-point embeddings to 8-bit integers, reducing memory consumption by up to 90% with negligible recall degradation (less than 1%).
  • Hybrid Search (Dense + Sparse BM25): Combines dense vector semantics with keyword BM25 sparse vectors in a single query for maximum precision.
  • Payload Filtering: High-performance pre-filtering based on tenant ID, document classification, or access control tags before vector similarity distance calculations.
  • Zero-Trust Private ClusterIP Access: Vector databases run strictly on private internal ClusterIP networks (:6333 HTTP, :6334 gRPC). Insecure public subdomains (qdrant.okustera.com) are eliminated, preventing unauthorized external access or vector data exfiltration.
  • Tenant-Isolated Vector DBaaS: In addition to the platform's central RAG knowledge base in ai-inference, tenants can provision dedicated, isolated Qdrant clusters with Ceph RBD storage and Barbican KMS API keys directly via Managed Databases: Qdrant Vector DB.
  • Interactive Web Dashboard: Built-in visual collection manager, vector distance explorer, and index telemetry.

6. Agentic Code Execution Sandboxes (Google gVisor)​

AI agents executing tool-calls, data analysis, or code generation require zero-trust execution boundaries:

  • Ephemeral Micro-Sandboxes: Spawns isolated Python and Bash execution pods inside Google gVisor (runsc) virtualized kernels in under 50 milliseconds.
  • Strict Multi-Tenant Isolation: Filters host syscalls via the gVisor Sentry process, preventing kernel escape, host file system traversal, or local network snooping.
  • Network & Quota Boundaries: Egress is strictly blocked from internal OpenStack control planes and cloud metadata services (169.254.169.254), with hard CPU/memory cgroups and execution timeouts.

7. LLM Observability & Tracing (Langfuse)​

Full operational observability for generative AI applications powered by self-hosted Langfuse:

  • OpenLLMetry & OpenTelemetry Integration: Trace every user prompt, reasoning chain, tool execution, and vector retrieval across distributed microservices.
  • Time-to-First-Token (TTFT) & Latency Heatmaps: Granular telemetry tracking token generation latency, queue wait time, and model execution time.
  • Prompt Versioning & Management: Centralized repository for prompt templates with A/B testing and production rollback capabilities.
  • Token Usage & Cost Attribution: Precise token count attribution per tenant, user, and API key, cross-referenced with Okustera's FinOps Billing Engine.

8. AI Gateway Ingress & Token FinOps (Apache APISIX)​

Extending Okustera's centralized Phase 2 API Gateway, Apache APISIX provides specialized AI Gateway capabilities:

  • Unbuffered SSE Streaming: Configured with proxy_buffering: "off" to ensure token streams reach client applications with sub-50ms Time-to-First-Token.
  • Model Payload Routing: Inspects incoming JSON request bodies ("model": "...") and routes traffic dynamically across autoscaled vLLM RayService instances.
  • Semantic Prompt Caching: Integrates with Valkey DBaaS (omc-valkey) to return identical or semantically equivalent prompt completions in under 5 milliseconds with zero GPU computation.
  • Token-Level FinOps: Automatically extracts usage.prompt_tokens and usage.completion_tokens from response streams and sends accounting events to the FinOps Billing Engine to debit tenant quotas.

9. AI Compute Hardware & Multi-Node Scaling Strategy​

To scale beyond a single bare-metal prototype, Okustera supports decoupled multi-node compute architectures:

Hardware Sizing Profiles​

ProfileTarget Hardware SpecVRAM / CapacityTarget Workloads & PerformanceCost Profile
Profile A: Consumer GPU PowerhouseAMD Ryzen 9 7950X, 64-128GB DDR5, 1-2x RTX 3090/409024 GB – 48 GB8B models at 60–90 tok/s (BF16); 70B models at 35–55 tok/s (AWQ/FP8)$1,500 – $2,800 one-off (~€120/mo rental)
Profile B: Enterprise Sovereign AI NodeDual AMD EPYC / Xeon Gold, 256GB ECC RAM, 1-2x NVIDIA L40S or A100/H10048 GB – 80 GB (NVIDIA MIG support)70B models unquantized; NVIDIA MIG partitions card into 7 isolated hardware slices$8,000 – $25,000 one-off (~€600/mo rental)
Profile C: High-RAM CPU ServerDual Xeon Gold / AMD EPYC (32–64 cores), 128–256GB ECC RAM, AVX-512Zero GPU (CPU Vector Extensions)Quantized 70B models (Q4_K_M) fit 100% in RAM at 8–14 tok/s; zero GPU power draw$1,000 – $2,000 refurbished enterprise server

Performance Matrix: Single-Node vs. Multi-Node Expansion​

Capability / MetricSingle Host Prototype (AIO)Multi-Node (+ Profile A: RTX 3090/4090)Multi-Node (+ Profile B: L40S / A100)
Max Real-Time Model7B (Quantized Q4)70B (Quantized AWQ/FP8)70B (Unquantized) / 405B (FP8)
Generation Speed (7B)12–18 tokens/s (CPU)70–90 tokens/s (GPU)100+ tokens/s (GPU)
Generation Speed (70B)0.1–1.5 tok/s (NVMe MoE)35–55 tokens/s (GPU)50–80 tokens/s (GPU)
Time-To-First-Token (TTFT)500ms – 2,000msunder 150msunder 80ms
Concurrent Tenant Streams2 – 4 streams20 – 50 streams100+ streams (via MIG)
Control Plane ContentionElevated during heavy promptsZero contention (Decoupled)Zero contention (Decoupled)

10. Example Manifest: vLLM on KubeRay (RayService)​

The following Kubernetes custom resource illustrates deploying an autoscaled vLLM model serving cluster using KubeRay:

apiVersion: ray.io/v1
kind: RayService
metadata:
name: vllm-llama-70b
namespace: default
spec:
serviceUnhealthyThreshold: 300
rayClusterConfig:
rayVersion: '2.35.0'
headGroupSpec:
rayStartParams:
dashboard-host: '0.0.0.0'
template:
spec:
containers:
- name: ray-head
image: rayproject/ray-ml:2.35.0-py310-gpu
resources:
limits:
cpu: '4'
memory: '16Gi'
workerGroupSpecs:
- groupName: vllm-workers
replicas: 1
minReplicas: 1
maxReplicas: 4
rayStartParams: {}
template:
spec:
tolerations:
- key: "dedicated"
operator: "Equal"
value: "ai-inference"
effect: "NoSchedule"
containers:
- name: vllm-worker
image: vllm/vllm-openai:v0.6.2
args:
- "--model"
- "/models/foundry/Meta-Llama-3.1-70B-Instruct"
- "--tensor-parallel-size"
- "4"
- "--max-model-len"
- "32768"
- "--enable-lora"
- "--max-loras"
- "32"
volumeMounts:
- name: model-weights
mountPath: /models/foundry
readOnly: true
resources:
limits:
nvidia.com/gpu: "4"
memory: "64Gi"
volumes:
- name: model-weights
persistentVolumeClaim:
claimName: cephfs-foundry-models-pvc

11. Querying via Python OpenAI SDK​

Applications query Okustera AI endpoints using the standard, drop-in OpenAI client library:

import os
from openai import OpenAI

# Initialize client pointing to Okustera AI Gateway
client = OpenAI(
base_url="https://ai.okustera.com/v1",
api_key=os.environ.get("OKUSTERA_API_KEY"),
)

# 1. Real-Time Streaming Chat Completion (With Kev Admission Routing)
stream = client.chat.completions.create(
model="qwen2.5:1.5b",
messages=[
{"role": "system", "content": "You are a sovereign cloud architect."},
{"role": "user", "content": "Explain zero-copy model weight memory mapping on CephFS."}
],
# Optional: explicitly request specific execution route, or let Kev classify automatically
extra_headers={"X-AI-Route": "deep_reasoning"},
stream=True,
temperature=0.7,
)

print("Streaming response:")
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print("\n")

# 2. Generating Vector Embeddings
embedding_response = client.embeddings.create(
model="bge-large-en-v1.5",
input=["Sovereign cloud infrastructure with European data residency."],
)

vector = embedding_response.data[0].embedding
print(f"Generated embedding vector of dimension {len(vector)}")