Skip to content
Cost Optimization DevOps Platform Engineering

Cloud Cost Governance v2: Pipelines, Admission, and SLO‑Aware Controls

Ian David Rossi
Ian David Rossi November 15, 2023 · 5 min read

TL;DR

Cloud Cost Governance v2: Pipelines, Admission, and SLO‑Aware Controls without the fluff: focus on outcomes, measure them, and stop pretending slides are progress.

Policy as Code for Costs

  • CI checks: block oversized requests, disallowed instance types, missing labels
  • Admission: enforce quotas, request ceilings, and region policies
  • Exceptions: time‑boxed with owner + reason

SLO‑Aware Budgets

  • Budgets informed by SLOs to avoid harmful cuts
  • Alerts on burn rate of budgets with actionable guidance

Ownership and Process

  • Cost champions per team; monthly reviews
  • Tickets with remediation steps and measured outcomes

Tooling

  • Kubecost data into CI comments; admission policies with Kyverno/OPA
  • Dashboards per team; PR annotations with expected cost delta

Conclusion

Governance works when cost feedback meets delivery workflows—early, actionable, and SLO‑aware.

“Make Cloud Cost Governance v2: Pipelines, Admission, and SLO‑Aware Controls boring: repeatable, measurable, and rehearsed.”

Long‑Form Addendum: Making Cloud Cost Governance v2: Pipelines, Admission, and SLO‑Aware Controls Feel “Obvious” to Teams

“Good platforms remove decisions; they don’t add dashboards.”

1) Start With a Promise

Write a one-paragraph promise to engineers: what the platform guarantees (templates, security defaults, observability, rollbacks) and what it requires (standards, ownership, on-call discipline). Publish SLOs for the platform itself so teams can trust it.

2) The Golden Path Mechanics

  • Scaffold: a template creates repo, pipeline, service skeleton, and dashboards.
  • Guardrails: policies are enforced in CI/CD and at deploy time.
  • Self-service: teams can create a service, namespace, secrets, and alerts without opening a ticket.
  • Upgrades: templates carry upgrades forward; deprecations are announced with timelines.

3) What to Measure (Outcomes)

  • Time to first deploy for a new service
  • Lead time for change and change failure rate
  • Platform ticket volume per team and time-to-close
  • % of services on the default path and number of exceptions
  • OKR progress tied to business outcomes (availability, onboarding speed)

4) Common Failure Modes

  • Too many options: teams choose differently and you can’t support it. Fix with fewer, better defaults.
  • Optional policies: “recommended” guardrails become ignored. Fix by enforcing with clear remediation.
  • No ownership: nobody owns migrations or breaking changes. Fix with product ownership and a deprecation policy.

5) Checklist

  • One default template exists per service type (API, worker, UI).
  • Exceptions are time-bound and visible.
  • Platform publishes SLOs and a changelog.
  • Monthly review: adoption, tickets, and incident learnings feed back into templates.

Glossary (Tooltips)

  • IDP: Where templates, guardrails, and paved roads live.
  • DX: What improves when you remove toil and ambiguity.
  • CI/CD: The path from commit to production.
  • SLO: How you keep the platform honest.
  • OKR: Helps align platform work to outcomes.

Appendix 1: Making the Standard Path the Fast Path

A Practical Product Loop

Platforms win by compounding: every week you make the default path easier and the exception path more expensive.

  1. Pick 1–2 templates as “blessed defaults.”
  2. Add guardrails in CI/CD and at deploy time, with actionable errors.
  3. Publish a changelog and deprecation policy so teams trust upgrades.
  4. Review adoption and exceptions monthly; prioritize work that reduces repeated tickets.

30/60/90 (Adoption and Reliability)

  • 30 days: define platform SLOs; ship one golden path with docs + runbook; measure baseline tickets and lead time.
  • 60 days: add self-service for the top ticket categories; introduce an exception registry (owner + expiry).
  • 90 days: enforce defaults for new services; migrate the highest-risk legacy patterns; tie outcomes to OKRs.

Checklist

  • One obvious default path exists (and it works reliably).
  • Exceptions are time-bound and visible.
  • Feedback loops exist: office hours, a roadmap, and real metrics.

Appendix 2: Making the Standard Path the Fast Path

A Practical Product Loop

Platforms win by compounding: every week you make the default path easier and the exception path more expensive.

  1. Pick 1–2 templates as “blessed defaults.”
  2. Add guardrails in CI/CD and at deploy time, with actionable errors.
  3. Publish a changelog and deprecation policy so teams trust upgrades.
  4. Review adoption and exceptions monthly; prioritize work that reduces repeated tickets.

30/60/90 (Adoption and Reliability)

  • 30 days: define platform SLOs; ship one golden path with docs + runbook; measure baseline tickets and lead time.
  • 60 days: add self-service for the top ticket categories; introduce an exception registry (owner + expiry).
  • 90 days: enforce defaults for new services; migrate the highest-risk legacy patterns; tie outcomes to OKRs.

Checklist

  • One obvious default path exists (and it works reliably).
  • Exceptions are time-bound and visible.
  • Feedback loops exist: office hours, a roadmap, and real metrics.

Appendix 3: Making the Standard Path the Fast Path

A Practical Product Loop

Platforms win by compounding: every week you make the default path easier and the exception path more expensive.

  1. Pick 1–2 templates as “blessed defaults.”
  2. Add guardrails in CI/CD and at deploy time, with actionable errors.
  3. Publish a changelog and deprecation policy so teams trust upgrades.
  4. Review adoption and exceptions monthly; prioritize work that reduces repeated tickets.

30/60/90 (Adoption and Reliability)

  • 30 days: define platform SLOs; ship one golden path with docs + runbook; measure baseline tickets and lead time.
  • 60 days: add self-service for the top ticket categories; introduce an exception registry (owner + expiry).
  • 90 days: enforce defaults for new services; migrate the highest-risk legacy patterns; tie outcomes to OKRs.

Checklist

  • One obvious default path exists (and it works reliably).
  • Exceptions are time-bound and visible.
  • Feedback loops exist: office hours, a roadmap, and real metrics.

Appendix 4: Making the Standard Path the Fast Path

A Practical Product Loop

Platforms win by compounding: every week you make the default path easier and the exception path more expensive.

  1. Pick 1–2 templates as “blessed defaults.”
  2. Add guardrails in CI/CD and at deploy time, with actionable errors.
  3. Publish a changelog and deprecation policy so teams trust upgrades.
  4. Review adoption and exceptions monthly; prioritize work that reduces repeated tickets.

30/60/90 (Adoption and Reliability)

  • 30 days: define platform SLOs; ship one golden path with docs + runbook; measure baseline tickets and lead time.
  • 60 days: add self-service for the top ticket categories; introduce an exception registry (owner + expiry).
  • 90 days: enforce defaults for new services; migrate the highest-risk legacy patterns; tie outcomes to OKRs.

Checklist

  • One obvious default path exists (and it works reliably).
  • Exceptions are time-bound and visible.
  • Feedback loops exist: office hours, a roadmap, and real metrics.