Gaps & capacity

How the environments differ, what is missing, and how much headroom is left.

Environment comparison

Differences between columns are where deployments break.

Propertydevstageproddatacenter
Runtimesdiffersdocker + systemd + nginxkubernetes + systemd + nginxdocker + kubernetes + systemd + nginxdocker + systemd
Kubernetes—v1.28.15v1.28.15—
Container runtimediffers—containerd://2.2.1containerd://2.2.4—
vCPUdiffers4288
Memory availablediffers13.7 / 15.6 GB5.8 / 7.8 GB26.3 / 31.3 GB25.6 / 31.3 GB
Disk availablediffers154.8 GB / 192.7 GB72.4 GB / 95.8 GB306.1 GB / 386.4 GB304.6 GB / 386.4 GB
Workloadsdiffers7133331
Public hostnamesdiffers54150
Datastores reachablediffers899n/a
Datastores blockeddiffers322n/a

7 architectural gaps

critical

Datastore connectivity

prod

prod cannot open connections to 2 shared datastore(s) on ogm-datacenter: milvus, postgres-n8n.

Why it matters — Any application deployed to prod that needs one of these will fail at startup or on its first query — and it will look like an application bug rather than a firewall rule.

What to do

ogm-datacenter publishes these ports through Docker, so plain 'ufw allow' does not apply. Add this environment's public IP to ALLOWED_SOURCES in the ogm-datacenter repository's scripts/setup-firewall.sh and re-run it, so the rule lands in the DOCKER-USER chain.

warning

Runtime parity

dev, stage, prod

Dev runs applications as plain Docker containers behind nginx, while Stage and Prod both run kubeadm Kubernetes. A deployment that works on Dev exercises none of the Kubernetes path it will hit next.

Why it matters — Anything Kubernetes-specific — manifests, probes, ingress rules, service accounts, image pull secrets, resource limits — is first tested in Stage. Failures that could have been caught locally surface one environment later, on shared infrastructure.

What to do

Install a single-node kubeadm cluster on Dev KVM4 at the same v1.28.x as Stage, or run k3s if the 4 vCPU / 16 GB budget is tight. Dev has 172 GB free and 14.8 GB of available memory, so it has the headroom. The Blueprint designer will then be able to generate identical manifests for all three environments instead of Compose for one and Kubernetes for the others.

warning

Runtime sprawl on Prod

prod

Prod KVM8 serves applications through three independent mechanisms at once: Kubernetes workloads, Docker containers, and bare-metal systemd units — all fronted by a single nginx.

Why it matters — There is no one place to ask "what is running here". Each mechanism has its own failure mode, its own restart policy and its own logs, and only nginx knows they are related. A service that dies leaves a route pointing at a dead port, and nothing notices.

What to do

Converge on Kubernetes for anything that serves traffic, in line with the Stage/Prod target. Migrate the systemd units and Docker containers into Deployments one at a time, pointing nginx at a NodePort or the ingress controller as each moves. Inspector's topology view lists exactly which routes still terminate outside Kubernetes.

warning

Datastore connectivity

dev

dev cannot open connections to 3 shared datastore(s) on ogm-datacenter: milvus, nfs, postgres-n8n.

Why it matters — Any application deployed to dev that needs one of these will fail at startup or on its first query — and it will look like an application bug rather than a firewall rule.

What to do

ogm-datacenter publishes these ports through Docker, so plain 'ufw allow' does not apply. Add this environment's public IP to ALLOWED_SOURCES in the ogm-datacenter repository's scripts/setup-firewall.sh and re-run it, so the rule lands in the DOCKER-USER chain.

warning

Datastore connectivity

stage

stage cannot open connections to 2 shared datastore(s) on ogm-datacenter: milvus, postgres-n8n.

Why it matters — Any application deployed to stage that needs one of these will fail at startup or on its first query — and it will look like an application bug rather than a firewall rule.

What to do

ogm-datacenter publishes these ports through Docker, so plain 'ufw allow' does not apply. Add this environment's public IP to ALLOWED_SOURCES in the ogm-datacenter repository's scripts/setup-firewall.sh and re-run it, so the rule lands in the DOCKER-USER chain.

warning

Single-node cluster

prod

prod runs Kubernetes on a single node, which is both the control plane and the only worker.

Why it matters — There is nowhere to reschedule a workload. Losing the node loses the cluster, and control-plane maintenance is indistinguishable from an outage.

What to do

Add at least one worker node to KVM8's cluster so workloads survive control-plane trouble, and keep etcd backed up to ogm-datacenter.

info

Single-node cluster

stage

stage runs Kubernetes on a single node, which is both the control plane and the only worker.

Why it matters — There is nowhere to reschedule a workload. Losing the node loses the cluster, and control-plane maintenance is indistinguishable from an outage.

What to do

Acceptable for a non-production environment; note it as a known limitation.