Skip to content
back to the quest log
A Thousand Rallies infrastructure diagram — multi-region load balancing, on-prem and EKS Kubernetes clusters, Unity DevOps pipeline
legendary⎈ infrastructureshipped

A Thousand Rallies

A game platform on three continents

Multi-region delivery for a Unity title — on-prem Kubernetes first, AWS EKS as the fallback, three exit nodes across ID, DE and US.

role
Infrastructure architect
org
Estella Studio
when
2025

What it is

A game needs to feel local everywhere. Rallies runs behind Cloudflare with three regional load-balancer/exit nodes — Sora in Jakarta, Shion in Frankfurt, Suisei in Ohio — fronting the web game, the marketing site and the back-end services.

The compute is deliberately hybrid: two on-prem Kubernetes clusters (Acheron and Mirai, the latter 48C/96T) carry normal load, and two AWS EKS clusters (Kawa and Mori) stand by to take over if the on-prem side fails. Assets live in Cloudflare R2.

The Unity side is wired in as a first-class citizen — Unity DevOps, VCS, Accelerator, Analytics and Cloud Diagnostics — so artists, designers and playtesters all land in the same pipeline as the engineers. Tetragon watches runtime; Terraform and GitHub Actions own everything else.

where the effort went

  • Scale90
  • Resilience95
  • Security72
  • Automation84

the numbers

Regions
3

ID · DE · US

K8s clusters
4

2 on-prem, 2 EKS

Biggest node
48C / 96T

Mirai

Asset store
Cloudflare R2

built with

  • Kubernetes
  • AWS EKS
  • Cloudflare
  • Cloudflare R2
  • Terraform
  • GitHub Actions
  • Tetragon
  • Unity
  • Docker
  • Prometheus

delivery record

What I made

The title needed low-latency access across three continents without making one cloud region or the on-premises cluster a single point of failure.

  • 01

    Three regional edge and exit nodes in Indonesia, Germany, and the United States

  • 02

    Two primary on-premises Kubernetes clusters

  • 03

    Two AWS EKS fallback clusters

  • 04

    Cloudflare edge and R2 asset delivery

  • 05

    Unity DevOps, version control, acceleration, analytics, and diagnostics pipeline

  • 06

    Terraform delivery, monitoring, and Tetragon runtime visibility

Hard problems

The constraints mattered as much as the finished interface.

01field note

Failover across unlike compute

problem
On-premises Kubernetes and managed EKS have different capacity, networking, and failure characteristics.
response
I kept the workload container contract portable and placed regional exits in front of both compute tiers so failover does not change the player-facing route.
02field note

One pipeline for artists and operators

problem
Game assets, Unity builds, services, and infrastructure move at different speeds and involve different teams.
response
Unity tooling joins the same delivery and observability path as containers and infrastructure, with assets separated into R2.

The rack

Named machines, as they appear in the diagram. Yes, they are all named after something.

  • MiraiOn-prem Kubernetes48C / 96T · 64 GiB
  • AcheronOn-prem Kubernetes2C / 4T · 12 GiB
  • KawaAWS EKSa1.metal
  • MoriAWS EKSa1.metal
  • SoraLoad balancer · exit nodeJakarta, ID
  • ShionLoad balancer · exit nodeFrankfurt, DE
  • SuiseiLoad balancer · exit nodeOhio, US
  • YumeCloud utility2 vCPU · 1 GiB
  • Asa · Hiru · YoruCloud utility2 vCPU · 1 GiB each

Architecture

Drawn as it was built. Open it full size — the labels are the interesting part.

after shipping

What stayed with me

  • Hybrid resilience depends on a portable workload and a stable edge, not on making every cluster identical.

  • A game platform includes the artist build path and diagnostic feedback loop, not only the runtime servers.