dograh/deploy/helm/dograh/examples/values-k3s-prod.yaml
prabhatlepton 6d1051757c
feat(helm): add HPA for arq-worker + ui, ship a lean k3s prod example (#516)
* feat(helm): add HPA for arq-worker + ui, ship a lean k3s prod example

## Problem

The chart's autoscaling story only covers the `web` tier — one
`web-hpa.yaml` template gated by `autoscaling.web.enabled`. Operators
scaling the `arq-worker` (background jobs) or `ui` (Next.js SSR) tiers
have to write their own HPA manifests out-of-band or fork the chart.

Turning the existing memory-utilization target on for freshly-installed
workloads also silently breaks: idle Python at the chart's default
`128Mi` (workers) / `256Mi` (ui) memory request already sits above
`80%`, so HPA scales every tier to `maxReplicas` on cold start with no
traffic. On a tight node this cascades into "insufficient CPU" and
blocks new-workload scheduling.

## Fix

**New HPA templates** — `templates/arq-worker-hpa.yaml` and
`templates/ui-hpa.yaml`, both mirroring the existing
`templates/web-hpa.yaml` shape (autoscaling/v2, resource metrics,
gated on `.Values.autoscaling.<tier>.enabled`).

**Extended `values.yaml`**:
- `autoscaling.workers` and `autoscaling.ui` blocks with sane defaults
  (`enabled: true`, `minReplicas: 1`, `maxReplicas: 5`,
  `targetCPUUtilizationPercentage: 70`).
- `targetMemoryUtilizationPercentage: null` on both tiers by default,
  with an inline comment explaining why memory-utilization HPA is a
  broken signal at the chart's default request sizes.
- Header comment reworked to (a) document the `metrics-server`
  requirement, (b) note that HPA takes ownership of Deployment
  `replicas` after first sync, (c) call out that CPU is a poor signal
  for the web tier (long-lived WebSockets), and (d) note that CPU is
  a fine signal for workers and ui.

**Example**: `examples/values-k3s-prod.yaml` — a single-node k3s
production override that exercises the new HPA blocks and demonstrates
the paired safety changes (memory targets nulled, sized resource
requests, migration job CPU sized for a tight node). Ship-ready
starting point for the operator flow: hosted-AI only (no local
models), all state on the node's local-path StorageClass, invite-only
signup, TLS terminated at a shared Cloudflare Origin cert.

## Behavior

Fresh install with defaults:
- Workers scale 1 → 5 on CPU 70% target only. No memory-based
  scale-up storm on cold start.
- UI scales 1 → 5 on CPU 70% target only.
- Web autoscaling stays `enabled: false` by default (unchanged) —
  operators opt in per the existing README warning.

Operators who want memory-based HPA back can:
1. Bump `workers.resources.requests.memory` (~256Mi) or
   `ui.resources.requests.memory` (~384Mi).
2. Set `autoscaling.<tier>.targetMemoryUtilizationPercentage: 80`.

* address review: omit replicas when HPA on, suppress empty-metrics HPA, docs

Fixes raised on #516:

- **Worker/UI Replicas Reset On Upgrade** — arq-worker-deployment.yaml and
  ui-deployment.yaml now wrap `replicas:` in `{{- if not .Values.autoscaling.<tier>.enabled }}`,
  mirroring the existing web-deployment guard. With HPA on, Helm no longer
  reapplies the static replicaCount on upgrade and briefly shrink an
  HPA-scaled pool.

- **Empty Metrics Render Invalid HPA** — arq-worker-hpa.yaml and ui-hpa.yaml
  now short-circuit the whole HPA object when both CPU and memory targets
  are null. Previously the template emitted `spec.metrics:` with no items
  (rejected by the k8s API server).

- **`enableSignup: false` removed from examples/values-k3s-prod.yaml** — that
  knob depends on #514 which hasn't landed; unwiring it here avoids
  suggesting a lockdown that isn't in effect until the sibling PR merges.

- **Header comment mismatch** — `# HPA: 1 → 5 on CPU 70% / memory 80%` claimed
  memory was on while every tier had `targetMemoryUtilizationPercentage: null`.
  Updated to "CPU 70% only (memory HPA opt-in)".

- **Wrong default in comment** — `values.yaml` said workers default is `128Mi`;
  actual is `256Mi`. Fixed.

- **UI comment said "idle Python"** — UI is Next.js/Node.js. Corrected on the
  UI HPA memory comment and the per-tier comments in values-k3s-prod.yaml
  (web: FastAPI, workers: Python/ARQ, ui: Node.js).

All lints pass; verified with `helm template`:
- Defaults render both HPAs and Deployments without static `replicas:`.
- `--set autoscaling.workers.targetCPUUtilizationPercentage=null --set autoscaling.workers.targetMemoryUtilizationPercentage=null`
  renders only the Deployment (HPA suppressed).
- `--set autoscaling.workers.enabled=false` renders the Deployment with
  static `replicas:` restored.

* address review: align Deployment replicas gate with HPA render gate

Follow-up on #516: my earlier fix guarded `spec.replicas` on only
`autoscaling.<tier>.enabled`, but the HPA-empty-metrics guard I added
suppresses the HPA object when both metric targets are null while
`enabled: true`. That combination produced a Deployment with neither
a `spec.replicas` value nor an HPA owner — a k8s Deployment defaults
to `replicas: 1` in that case, but the chart no longer expresses intent.

Fix: the Deployment `replicas` gate now mirrors the HPA render gate
exactly. Rendered outcomes verified with `helm template`:

| autoscaling.<tier>            | HPA rendered? | Deployment replicas? |
|-------------------------------|---------------|----------------------|
| enabled: true, target set     | yes           | omitted (HPA owns)   |
| enabled: true, both null      | no            | static (kept)        |
| enabled: false                | no            | static (kept)        |

* fix(helm): default worker/ui autoscaling off; ui HPA floor of 2

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(helm): align web replicas/HPA gate with worker/ui pattern

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(helm): document worker/ui HPAs in README; polish k3s example

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: prabhat pankaj <prabhatiitbhu@gmail.com>
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 14:48:48 +05:30

117 lines
4.3 KiB
YAML

# Dograh — production values for k3s
# Hosted-AI only (no local models). All state on the node's local disk via
# k3s local-path StorageClass. TLS terminated at Cloudflare edge; re-encrypted
# to origin using the shared *.yourdomain.com Cloudflare Origin CA cert
# (secret `cloudflare-origin-tls`, copied from a neighboring namespace).
#
# HPA is enabled for web / workers / ui with min=1, max=5. Requires
# metrics-server in the cluster (k3s ships it by default).
image:
tag: latest # pin to a released tag once stable
# --- Bundled stateful deps (in-cluster, on local-path PVCs) -----------------
database: { mode: internal }
redis: { mode: internal }
storage: { mode: internalMinio }
# --- HTTP exposure ----------------------------------------------------------
exposure:
mode: ingress
ingress:
className: traefik
host: aicalling.yourdomain.com
tls:
enabled: true
secretName: cloudflare-origin-tls
# --- Runtime config (rendered into the ConfigMap) --------------------------
config:
environment: production
logLevel: INFO
backendApiEndpoint: https://aicalling.yourdomain.com # kills the CF-tunnel fallback
minioPublicEndpoint: https://aicalling.yourdomain.com # browser fetches /voice-audio/ via Traefik
enableSignup: false # invite-only: 403 the public signup endpoint
# --- Workloads not needed for hosted-AI web-only calls ---------------------
ariManager: { enabled: false } # Asterisk SIP singleton
campaignOrchestrator: { enabled: false } # scheduled outbound singleton
coturn: { enabled: false } # TURN relay
# --- Web tier (FastAPI + WebSocket signaling) ------------------------------
web:
replicaCount: 1
resources:
requests: { cpu: 100m, memory: 384Mi }
limits: { cpu: "1", memory: 1Gi }
pdb: { enabled: false }
# --- ARQ background workers ------------------------------------------------
workers:
replicaCount: 1
resources:
requests: { cpu: 50m, memory: 128Mi }
limits: { cpu: 500m, memory: 512Mi }
# --- Next.js UI ------------------------------------------------------------
ui:
replicaCount: 1
resources:
requests: { cpu: 50m, memory: 256Mi }
limits: { cpu: 500m, memory: 512Mi }
pdb: { enabled: false }
# --- Alembic migration Job (post-install / pre-upgrade hook) ---------------
# CPU request kept low so it schedules on a tight node; migrations are
# short-lived and don't need much headroom.
migrate:
resources:
requests: { cpu: 20m, memory: 256Mi }
limits: { cpu: 500m, memory: 512Mi }
# --- HPA: 1 → 5 on CPU 70% only (memory HPA opt-in — see comments) ---------
# The chart ships all autoscaling blocks disabled; this file opts in for all
# three tiers. NOTE: CPU is a poor scale signal for the web tier (long-lived
# WebSockets don't move CPU much) and a coarse one for the IO-bound ARQ
# workers — see the autoscaling notes in values.yaml. Fine starting point for
# a single node; minReplicas=1 is why the web/ui PDBs are disabled above.
autoscaling:
web:
enabled: true
minReplicas: 1
maxReplicas: 5
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: null # idle FastAPI already close to request; CPU HPA is enough
workers:
enabled: true
minReplicas: 1
maxReplicas: 5
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: null # idle Python (ARQ) sits near the memory request; CPU HPA is enough
ui:
enabled: true
minReplicas: 1
maxReplicas: 5
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: null # idle Node.js (Next.js SSR) sits near the memory request; CPU HPA is enough
# --- Internal Postgres (pgvector/pg17) -------------------------------------
postgresql:
persistence: { size: 5Gi, storageClass: local-path }
resources:
requests: { cpu: 100m, memory: 256Mi }
limits: { cpu: "1", memory: 1Gi }
# --- Internal Redis --------------------------------------------------------
redisinternal:
persistence: { size: 1Gi, storageClass: local-path }
resources:
requests: { cpu: 50m, memory: 64Mi }
limits: { cpu: 300m, memory: 256Mi }
# --- Internal MinIO (audio + artifacts on node disk) -----------------------
minio:
persistence: { size: 10Gi, storageClass: local-path }
resources:
requests: { cpu: 50m, memory: 128Mi }
limits: { cpu: 500m, memory: 512Mi }