Peter Wiggers
Clay: Kubernetes & platform

Your requests and limits are where the cloud bill hides

How right-sizing Kubernetes workloads cut public cloud spend by roughly 40% across thirty-odd clusters.

1 min read

Kubernetes schedules on requests, not on what your pods actually use. So when every team copies cpu: 1, memory: 2Gi from the last service they wrote, the cluster autoscaler dutifully buys nodes for capacity nobody touches.

The pattern

  1. Measure real usage over a representative window (p95, not average).
  2. Set requests close to that, and leave headroom where it matters.
  3. Be careful with CPU limits: throttling is often worse than the noisy neighbour you were afraid of.
  4. Automate it, because hand-tuned values drift within a quarter.
resources:
  requests:
    cpu: 150m
    memory: 320Mi
  limits:
    memory: 512Mi

Tooling beats policy

We wrote a small Go tool that compared requests against observed usage and opened pull requests with suggestions. Engineers accepted most of them, because a PR with numbers is easier to say yes to than a wiki page with rules.