Skip to content
back to the quest log
High availability Chatwoot architecture across Cloudflare, two AWS availability zones, EKS, RDS, Redis, S3, and SES
epic⎈ infrastructureshipped

Chatwoot on AWS

One command, multi-AZ, recoverable

An idempotent Terraform and Ansible deployment for highly available Chatwoot on EKS, RDS, Redis, S3, SES, and Cloudflare.

role
Infrastructure author
org
Independent open source
when
2026

What it is

The repository turns a fresh AWS account into a multi-availability-zone Chatwoot installation through one deploy command and nine gated phases.

Terraform provisions the network and managed services. Ansible coordinates cluster bootstrapping and application rollout. EKS runs web and Sidekiq workloads with autoscaling, topology spread, and disruption budgets.

A read-only discovery phase protects shared accounts before creation. Secrets move from Secrets Manager through External Secrets and IRSA, so static AWS keys do not enter pods or Terraform state.

where the effort went

  • Reliability96
  • Security94
  • Automation92
  • Recovery86

the numbers

Phases
9
Availability zones
2
App autoscale
2-8
Terraform files
23

built with

  • Terraform
  • Ansible
  • Amazon EKS
  • Amazon RDS
  • ElastiCache
  • Amazon S3
  • Amazon SES
  • Cloudflare
  • External Secrets
  • Kubernetes

delivery record

What I made

I wanted a reference deployment that treats teardown, secret boundaries, shared-account safety, and verification as first-class parts of high availability.

  • 01

    Two-AZ VPC with public, private, and database subnet tiers

  • 02

    EKS application platform with autoscaling and disruption controls

  • 03

    RDS PostgreSQL 16 with pgvector and Multi-AZ failover

  • 04

    Redis high availability, S3 media, SES email, and Cloudflare edge

  • 05

    Secrets Manager to External Secrets flow secured by IRSA

  • 06

    Nine idempotent deployment phases with gates, verification, and teardown

  • 07

    Read-only CIDR discovery and collision checks for shared AWS accounts

Hard problems

The constraints mattered as much as the finished interface.

01field note

Building safely in a shared account

problem
A deployment script can collide with existing networks, names, and resources before Terraform has enough context to protect them.
response
Phase zero performs read-only discovery, selects an unused CIDR, checks names, and stops before any create operation when a collision exists.
02field note

Keeping credentials out of state and pods

problem
Passing application secrets through Terraform variables or static Kubernetes secrets leaves long-lived copies in sensitive control planes.
response
Secrets live in AWS Secrets Manager, sync through External Secrets, and use IRSA for workload access without static AWS credentials.
03field note

High availability that survives maintenance

problem
Multiple replicas are not enough if the scheduler places them together or an upgrade evicts all workers at once.
response
Topology spread, pod disruption budgets, horizontal autoscaling, two NAT gateways, Multi-AZ RDS, and Redis failover cover separate failure domains.

Architecture

Drawn as it was built. Open it full size — the labels are the interesting part.

after shipping

What stayed with me

  • A reliable deployment begins with discovery and ends with teardown; resource creation is only the middle.

  • Availability claims should map to named failure domains and a verification command, not only to replica counts.