Kaizin Platform
GKE, but it defends itself
A GKE platform wired end to end: observability, runtime security with automatic response, FinOps governance and an AI on-call that phones a human.
- role
- Site Reliability Engineer
- org
- Kaizin Digital
- when
- 2023 — 2024
What it is
The brief was reliability, but reliability without security is theatre. This platform runs on GKE with Anthos Service Mesh, Config Sync and Policy Controller — everything reaches the cluster through Cloud Build, Cloud Deploy and Terraform pull requests, nothing by hand.
Telemetry funnels through an OpenTelemetry collector into Prometheus, Jaeger, Elastic and Grafana. Tetragon watches syscalls at runtime and feeds a SIEM, which drives SOAR playbooks that can push a WAF block rule, revoke tokens, rotate Vault secrets, quarantine a namespace, cordon nodes or trigger a CI/CD rollback — before anyone has opened a laptop.
Alongside it: FinOps with billing exported to BigQuery and quarterly reviews, CIS Kubernetes controls mapped to ISO 27001 and SOC 2, and an AI agent that reaches a human on Discord or by phone with text-to-speech when a playbook needs a decision.
where the effort went
- Reliability95
- Security96
- Automation90
- FinOps78
the numbers
- Control plane
- GKE + Anthos
- MTTR
- ↓
- Autoscaling
- HPA + VPA
- Compliance
- CIS → ISO 27001 / SOC 2
blameless postmortems, error budgets
built with
- GKE
- Anthos Service Mesh
- Config Sync
- Terraform
- Cloud Build
- OpenTelemetry
- Prometheus
- Jaeger
- Grafana
- Elasticsearch
- Tetragon
- Vault
- Cloudflare
- BigQuery