Amazon ECS

Deploy Stategraph on Amazon ECS Fargate with the Stategraph ECS Terraform module.

Before you begin

  • Terraform 1.0 or later
  • AWS CLI 2 with credentials for ECS, RDS, ALB, IAM, Secrets Manager, and CloudWatch
  • A VPC with private and public subnets, or the complete example that creates one
  • An ACM certificate for HTTPS
  • A domain name for the instance
  • The ECS Terraform module, available on request with onboarding support: contact Stategraph for the source

Architecture

Internet
Application Load BalancerHTTPS:443
ECS serviceFargate tasks in private subnets
RDS PostgreSQLMulti-AZ
Secrets Managercredentials
CloudWatch Logscontainer logs
HTTPS requests from the internet reach the Application Load Balancer on port 443.
The load balancer sends them to the Fargate tasks of the ECS service, in private subnets.
The tasks use Multi-AZ RDS PostgreSQL, Secrets Manager for the credentials, and CloudWatch Logs for the container logs.

The module creates the resources in the diagram, and also:

  • ECS cluster
  • Security groups for the ALB, tasks, and RDS
  • IAM roles for the tasks
  • Auto scaling policies

Quick start

1. Request an ACM certificate

aws acm request-certificate \
  --domain-name stategraph.example.com \
  --validation-method DNS \
  --region us-east-1

Validate it through DNS as the AWS console instructs, and note its ARN.

2. Write the Terraform configuration

mkdir stategraph-ecs && cd stategraph-ecs
# main.tf
module "stategraph" {
  source = "<stategraph-ecs-module>?ref=v1.0.0"

  # Network (your existing VPC)
  vpc_id             = "vpc-xxxxx"
  private_subnet_ids = ["subnet-xxxxx", "subnet-yyyyy"]
  public_subnet_ids  = ["subnet-aaaaa", "subnet-bbbbb"]

  # Domain and certificate
  domain_name     = "stategraph.example.com"
  certificate_arn = "arn:aws:acm:us-east-1:123456789012:certificate/xxxxx"

  # Image
  stategraph_image = "ghcr.io/stategraph/stategraph-server:2.5.7"

  environment = "production"

  tags = {
    Project   = "Stategraph"
    ManagedBy = "Terraform"
  }
}

output "alb_dns_name" {
  description = "Create a CNAME record pointing your domain here"
  value       = module.stategraph.alb_dns_name
}

output "database_endpoint" {
  value     = module.stategraph.database_endpoint
  sensitive = true
}

3. Apply

terraform init
terraform plan
terraform apply

The apply takes 10 to 15 minutes.

First startup

On the first deploy, migrations run for a few minutes before the server accepts requests. The module configures the ALB target group and the container health check for this window. See Health checks for the recommended intervals and grace period.

4. Configure DNS

Point your domain at the load balancer with a CNAME, or with a Route 53 alias record and the ALB zone ID:

terraform output alb_dns_name
terraform output -raw alb_zone_id

5. Open Stategraph

After DNS propagates, open https://stategraph.example.com and create the first admin account on the setup screen. The module sets STATEGRAPH_UI_BASE to https://<domain_name>, so cookies have the Secure flag. Keep domain_name in step with the certificate and DNS record.

Configuration

Required variables

Variable Description Example
vpc_id VPC for every resource vpc-xxxxx
private_subnet_ids Private subnets for the tasks and RDS ["subnet-xxxxx", "subnet-yyyyy"]
public_subnet_ids Public subnets for the ALB ["subnet-aaaaa", "subnet-bbbbb"]
domain_name Public host name stategraph.example.com
certificate_arn ACM certificate for HTTPS arn:aws:acm:...

Common optional variables

Variable Description Default
environment Environment name production
ecs_task_cpu Task CPU units 1024
ecs_task_memory Task memory in MB 2048
ecs_desired_count Number of tasks 2
enable_autoscaling Scale on CPU and memory true
database_instance_class RDS instance class db.t3.medium
database_multi_az Multi-AZ RDS true
stategraph_image Container image ghcr.io/stategraph/stategraph-server:latest

The module documentation lists every variable. Stategraph settings are environment variables on the server container in the task definition. Reference Secrets Manager secrets in the secrets list of the container definition, not in environment.

VPC requirements

  • At least two private subnets in different availability zones, for the tasks and RDS
  • At least two public subnets in different availability zones, for the ALB
  • A NAT gateway or NAT instances for internet access from the private subnets
  • DNS hostnames enabled

To create a VPC, use the terraform-aws-modules/vpc module, and pass its outputs:

module "vpc" {
  source  = "terraform-aws-modules/vpc/aws"
  version = "~> 5.0"

  name = "stategraph-vpc"
  cidr = "10.0.0.0/16"

  azs             = ["us-east-1a", "us-east-1b", "us-east-1c"]
  private_subnets = ["10.0.1.0/24", "10.0.2.0/24", "10.0.3.0/24"]
  public_subnets  = ["10.0.101.0/24", "10.0.102.0/24", "10.0.103.0/24"]

  enable_nat_gateway   = true
  enable_dns_hostnames = true
}

module "stategraph" {
  source = "<stategraph-ecs-module>?ref=v1.0.0"

  vpc_id             = module.vpc.vpc_id
  private_subnet_ids = module.vpc.private_subnets
  public_subnet_ids  = module.vpc.public_subnets

  # ... other variables
}

Using existing PostgreSQL

Set create_database = false and pass the connection details:

module "stategraph" {
  source = "<stategraph-ecs-module>?ref=v1.0.0"

  # ... other required variables

  create_database            = false
  external_database_host     = "postgres.example.com"
  external_database_port     = 5432
  external_database_name     = "stategraph"
  external_database_username = var.database_username
  external_database_password = var.database_password
}

The tasks need network access to the database server, and the stategraph database must exist before the first start.

Authentication

Local email and password sign-in is on by default. Turn on Google or OIDC sign-in through the module:

module "stategraph" {
  source = "<stategraph-ecs-module>?ref=v1.0.0"

  # ... other required variables

  oauth_enabled       = true
  oauth_provider      = "google"  # or "oidc"
  oauth_client_id     = var.oauth_client_id
  oauth_client_secret = var.oauth_client_secret

  # For an OIDC provider:
  # oauth_issuer_url = "https://your-provider.example.com"
}

In your provider, register the callback URL: https://stategraph.example.com/oauth2/google/callback for Google, https://stategraph.example.com/oauth2/oidc/callback for OIDC. Keep the client secret in an uncommitted terraform.tfvars file or a secrets manager.

The module runs ecs_desired_count tasks. Set STATEGRAPH_OAUTH_COOKIE_SECRET in the server container environment to the same 16, 24, or 32 character value on each task. Otherwise each task signs cookies with its own random secret, and tasks do not share sessions. See Access control.

Enable cost estimation

Set STATEGRAPH_COST_ENABLED=true in the server container environment of the task definition, so that each replacement task starts with cost on:

"containerDefinitions": [
  {
    "name": "stategraph",
    "environment": [
      { "name": "STATEGRAPH_COST_ENABLED", "value": "true" }
    ]
  }
]

At the first start, Stategraph loads the price book into cloud_pricing in the background. Cost estimation connects with PRICING_DB_*, not DB_*, so point PRICING_DB_HOST and PRICING_DB_PASSWORD at the RDS instance. See Enable cost estimation.

Enable Orchestration

Stategraph Orchestration is off by default. To turn it on, add STATEGRAPH_ORCHESTRATION_ENABLED=true to the server container environment, with the Orchestration public URLs. Keep each GitHub App or GitLab credential in Secrets Manager. For GitLab, these are GITLAB_APP_ID, GITLAB_APP_SECRET, GITLAB_ACCESS_TOKEN, and STATEGRAPH_FDW_PROVISIONER_PASSWORD.

Orchestration needs a second database, terrateam, on the RDS instance. The stategraph database reads it through postgres_fdw, which on Amazon RDS needs the rds_superuser role for the database user. Enable Orchestration lists the variables and the database steps.

To store a credential, for example the GitHub App private key:

aws secretsmanager put-secret-value \
  --secret-id <secret-arn> \
  --secret-string file://private-key.pem

After you change a secret value, force a new deployment so that tasks read it:

aws ecs update-service \
  --cluster $(terraform output -raw ecs_cluster_name) \
  --service $(terraform output -raw ecs_service_name) \
  --force-new-deployment

Then set the GitHub App webhook URL to https://stategraph.example.com/api/github/v1/events. For GitLab, connect each group from the Get Started page of the console. Its connect wizard shows the project webhook URL, https://stategraph.example.com/api/v1/gitlab/events, and the secret.

Scaling

To tune auto scaling on CPU and memory:

module "stategraph" {
  source = "<stategraph-ecs-module>?ref=v1.0.0"

  # ... other variables

  enable_autoscaling           = true
  ecs_min_count                = 1
  ecs_desired_count            = 2
  ecs_max_count                = 4
  autoscaling_cpu_threshold    = 70
  autoscaling_memory_threshold = 80
}

For a fixed count, set enable_autoscaling = false and ecs_desired_count. For more headroom per task, raise ecs_task_cpu and ecs_task_memory, for example to 2048 and 4096.

Upgrading

Change the image tag and apply. ECS does a rolling deployment, and the load balancer routes only to tasks that pass the health check:

module "stategraph" {
  # ... other variables
  stategraph_image = "ghcr.io/stategraph/stategraph-server:2.5.7"
}
terraform apply

To update the module, change ref in source, then run terraform init -upgrade, terraform plan, and terraform apply. Give the task an ECS stopTimeout of 60 seconds, so that Stategraph can finish requests in progress. See Upgrades.

Monitoring

Container logs:

aws logs tail $(terraform output -raw cloudwatch_log_group_name) --follow

Service status:

aws ecs describe-services \
  --cluster $(terraform output -raw ecs_cluster_name) \
  --services $(terraform output -raw ecs_service_name)

The module turns on Container Insights. In the CloudWatch console, Container Insights > ECS Clusters shows CPU and memory use, target response time, healthy host count, and request count. The database credentials are in Secrets Manager:

aws secretsmanager get-secret-value \
  --secret-id $(terraform output -raw database_secret_arn) \
  --query SecretString --output text | jq -r '.password'

See Observability for other signals.

Troubleshooting

Check the container log first:

aws logs tail $(terraform output -raw cloudwatch_log_group_name) --follow
  • Task stops right after it starts: a secret that the task references is empty, or a required variable is missing. The log says which.
  • 502 or 503 from the load balancer, and 502 from /health/ready: the task is still migrating. On the first deploy, wait a few minutes.
  • Sign-in redirects to the wrong URL: domain_name changed. Run terraform apply and force a new deployment, so that the tasks read the new STATEGRAPH_UI_BASE.

Next steps