Skip to content
back to the quest log
legendary⎈ infrastructureinternal

Kaizin Platform

GKE, but it defends itself

A GKE platform wired end to end: observability, runtime security with automatic response, FinOps governance and an AI on-call that phones a human.

role
Site Reliability Engineer
org
Kaizin Digital
when
2023 — 2024

What it is

The brief was reliability, but reliability without security is theatre. This platform runs on GKE with Anthos Service Mesh, Config Sync and Policy Controller — everything reaches the cluster through Cloud Build, Cloud Deploy and Terraform pull requests, nothing by hand.

Telemetry funnels through an OpenTelemetry collector into Prometheus, Jaeger, Elastic and Grafana. Tetragon watches syscalls at runtime and feeds a SIEM, which drives SOAR playbooks that can push a WAF block rule, revoke tokens, rotate Vault secrets, quarantine a namespace, cordon nodes or trigger a CI/CD rollback — before anyone has opened a laptop.

Alongside it: FinOps with billing exported to BigQuery and quarterly reviews, CIS Kubernetes controls mapped to ISO 27001 and SOC 2, and an AI agent that reaches a human on Discord or by phone with text-to-speech when a playbook needs a decision.

where the effort went

  • Reliability95
  • Security96
  • Automation90
  • FinOps78

the numbers

Control plane
GKE + Anthos
MTTR
↓

blameless postmortems, error budgets

Autoscaling
HPA + VPA
Compliance
CIS → ISO 27001 / SOC 2

built with

  • GKE
  • Anthos Service Mesh
  • Config Sync
  • Terraform
  • Cloud Build
  • OpenTelemetry
  • Prometheus
  • Jaeger
  • Grafana
  • Elasticsearch
  • Tetragon
  • Vault
  • Cloudflare
  • BigQuery

delivery record

What I made

The platform needed reliability, security response, cost governance, and auditability to reinforce each other instead of living in separate operational silos.

  • 01

    GitOps delivery to GKE through Config Sync and Policy Controller

  • 02

    Anthos service mesh and OpenTelemetry observability pipeline

  • 03

    Tetragon runtime detection connected to SIEM and SOAR playbooks

  • 04

    Automated containment, credential rotation, quarantine, and rollback actions

  • 05

    BigQuery billing export and recurring FinOps review process

  • 06

    Human escalation through Discord and text-to-speech phone calls

Hard problems

The constraints mattered as much as the finished interface.

01field note

Automating response without automating damage

problem
Security playbooks can reduce response time but a false positive can block users, rotate secrets, or quarantine healthy workloads.
response
Actions were split by reversibility and confidence, with human escalation for decisions that cross the safe automation boundary.
02field note

Keeping the cluster reproducible

problem
Manual emergency changes undermine both recovery and compliance evidence.
response
Terraform and pull-request delivery remain the source of truth, and rollback is itself an automated pipeline action.

Architecture

after shipping

What stayed with me

  • Fast response is useful only when the response path has explicit limits and a human decision point.

  • Reliability, security, and cost signals become more actionable when they share one operational timeline.