I tried out ECS support for blue/green deployment of services exposed via VPC Lattice
This page has been translated by machine translation. View original
Introduction
In an announcement on October 2, 2026, Amazon ECS support for VPC Lattice gained built-in blue/green, linear, and canary deployments. This can be enabled for both new and existing ECS services.
You only need to configure a Lattice target group for your ECS service — there is no need to create a separate ALB. On the other hand, blue/green via CodeDeploy (the CODE_DEPLOY deployment controller) is not supported with Lattice. The documentation states that linear and canary deployments run on the same Lattice resources as blue/green.
In this article, I created a Nginx Fargate service without an ALB and performed a blue/green deployment from v1 to v2. I'll walk through a chronological observation of how ECS rewrites the weights of Lattice rules during deployment and how client responses change over time.
Lattice and ECS Configuration
Resources to Prepare on the Lattice Side
I prepared the following on the Lattice side. I used the same configuration examples from the official documentation for the creation commands.
- Two target groups (IP type, HTTP:80. Both kept ACTIVE)
- A Lattice service and listener (default action is a fixed 404)
- A production rule (prefix match on path
/) - A test rule (header
X-Environment: test, priority 10)
Rule actions must be forward only — ECS rejects rules with fixedResponse. Both rules must include both target groups, with the side currently receiving traffic set to weight 100 and the other to 0. The test rule is matched against a different request than the production rule and is assigned a lower priority number so it is evaluated first.
During deployment, ECS rewrites the forward action of the rules. Making changes externally will cause the deployment to fail, so leave it to ECS. The two target groups swap roles with each deployment. In the first staged deployment, traffic is shifted to the side that is not the one specified as targetGroupArn (the side with weight 100 in the production rule) — that is, to the alternate side. Hereinafter, the side specified as targetGroupArn is referred to as TG-A, and the alternate side as TG-B.
The only inbound traffic allowed in the task's security group was port 80 from the Lattice managed prefix list. The prefix list name is com.amazonaws.<region>.vpc-lattice, where <region> is replaced with the name of the region you are using.
ECS Configuration
On the ECS side, add advancedConfiguration to vpcLatticeConfigurations. For roleArn, specify the ARN of the infrastructure role that allows ECS to manage VPC Lattice resources. Specify the second target group for alternateTargetGroupArn, and the production and test rules for productionListenerRule and testListenerRule respectively. Set the strategy in deploymentConfiguration to BLUE_GREEN.
The official documentation describes how to switch a service currently running with rolling deployments. Add advancedConfiguration either in the same update that changes the strategy, or in an earlier update. In this case, I verified this with create-service to create a new service.
I saved the following JSON as svc.json and passed it to aws ecs create-service --cli-input-json file://svc.json. IDs have been replaced with example values.
{
"cluster": "lattice-bg",
"serviceName": "lattice-bg-web",
"taskDefinition": "lattice-bg-nginx:1",
"desiredCount": 2,
"launchType": "FARGATE",
"networkConfiguration": {
"awsvpcConfiguration": {
"subnets": [
"subnet-0123456789abcdef0",
"subnet-0fedcba9876543210"
],
"securityGroups": [
"sg-0123456789abcdef0"
],
"assignPublicIp": "ENABLED"
}
},
"vpcLatticeConfigurations": [
{
"roleArn": "arn:aws:iam::123456789012:role/lattice-bg-ecs-infra",
"targetGroupArn": "arn:aws:vpc-lattice:ap-northeast-1:123456789012:targetgroup/tg-0123456789abcdef0",
"portName": "web",
"advancedConfiguration": {
"alternateTargetGroupArn": "arn:aws:vpc-lattice:ap-northeast-1:123456789012:targetgroup/tg-0fedcba9876543210",
"productionListenerRule": "arn:aws:vpc-lattice:ap-northeast-1:123456789012:service/svc-0123456789abcdef0/listener/listener-0123456789abcdef0/rule/rule-0123456789abcdef0",
"testListenerRule": "arn:aws:vpc-lattice:ap-northeast-1:123456789012:service/svc-0123456789abcdef0/listener/listener-0123456789abcdef0/rule/rule-0fedcba9876543210"
}
}
],
"deploymentController": {
"type": "ECS"
},
"deploymentConfiguration": {
"strategy": "BLUE_GREEN",
"maximumPercent": 200,
"minimumHealthyPercent": 100,
"bakeTimeInMinutes": 2
}
}
Task Definition and Client
I used the official Nginx image for the task definition. At startup, the container retrieves the task ARN, revision, and Availability Zone from the ECS task metadata (ECS_CONTAINER_METADATA_URI_V4/task). It writes this to index.html, then starts Nginx. v1 uses nginx:1.27 and APP_VERSION=v1, v2 uses nginx:1.28 and APP_VERSION=v2, and everything else is the same. I saved the full v2 definition as td2.json and registered it with aws ecs register-task-definition --cli-input-json file://td2.json.
Task definition v2 (td2.json)
{
"family": "lattice-bg-nginx",
"networkMode": "awsvpc",
"requiresCompatibilities": [
"FARGATE"
],
"cpu": "256",
"memory": "512",
"executionRoleArn": "arn:aws:iam::123456789012:role/lattice-bg-task-exec",
"containerDefinitions": [
{
"name": "web",
"image": "public.ecr.aws/docker/library/nginx:1.28",
"essential": true,
"entryPoint": [
"/bin/sh",
"-c"
],
"command": [
"M=$(curl -s \"$ECS_CONTAINER_METADATA_URI_V4/task\")\ng(){ echo \"$M\" | sed -n \"s/.*\\\"$1\\\":\\\"\\([^\\\"]*\\)\\\".*/\\1/p\" | head -1; }\nprintf 'app_version=%s\\ntask_arn=%s\\ntaskdef_revision=%s\\navailability_zone=%s\\ncontainer_ip=%s\\nnginx=%s\\n' \"$APP_VERSION\" \"$(g TaskARN)\" \"$(g Revision)\" \"$(g AvailabilityZone)\" \"$(hostname -i)\" \"$(nginx -v 2>&1)\" > /usr/share/nginx/html/index.html\nexec nginx -g 'daemon off;'\n"
],
"portMappings": [
{
"name": "web",
"containerPort": 80,
"protocol": "tcp",
"appProtocol": "http"
}
],
"environment": [
{
"name": "APP_VERSION",
"value": "v2"
}
],
"logConfiguration": {
"logDriver": "awslogs",
"options": {
"awslogs-group": "/lattice-bg",
"awslogs-region": "ap-northeast-1",
"awslogs-stream-prefix": "web"
}
}
}
]
}
The Lattice service is invoked from clients through a service network or VPC association. The client in this case was a Lambda function placed in a VPC (the same VPC as the tasks) associated with the service network. Requests were sent at approximately 13-second intervals along two paths: the production path without a header, and the test path with X-Environment: test.
The environment was ap-northeast-1, with 2 Fargate tasks at 0.25 vCPU / 0.5 GB, and a target group health check interval of 10 seconds.
Initial Service Creation
Deployment completion was determined by the status in describe-service-deployments becoming SUCCESSFUL. All timestamps below are the times at which they were recorded.
Even for a newly created service, the initial deployment ran as blue/green. The service was created at 00:41:31. Both tasks were registered not in TG-A (specified as targetGroupArn) but in the alternate TG-B, and became HEALTHY. Both became HEALTHY at 00:42:45, and TG-A was empty from the earliest recorded time of 00:42:05.
The rolloutState of the ECS service (from describe-services) showed COMPLETED at 00:43:25. Even at that point, the production rule weights remained as set at creation: TG-A 100 / TG-B 0. The first observed stage at 00:45:49 was POST_SCALE_UP (no records of earlier stage transitions exist). After that, TEST_TRAFFIC_SHIFT (00:46:39) and PRODUCTION_TRAFFIC_SHIFT (00:48:18) were observed. The production rule changed to TG-A 0 / TG-B 100 at 00:48:20, and the service deployment became SUCCESSFUL at 00:49:55.
Deployment from v1 to v2
I passed the v2 task definition via update-service. bakeTimeInMinutes remained at 2. The v1 tasks were registered in TG-B, and the v2 tasks were registered in TG-A.
The timestamps in the table below are JST on 2026-10-04. ECS stage times are the first observed times when polling describe-service-deployments at 4-second intervals. Rule weights and client responses were recorded at approximately 13-second intervals in sync with Lambda requests.
| Time | ECS Stage | Observations from Lattice and Client |
|---|---|---|
| 01:12:39 | (update-service executed) |
Both production and test paths respond with v1 |
| 01:12:53 | SCALE_UP | v2 tasks registered in TG-A (first at 01:13:19, both HEALTHY at 01:13:59) |
| 01:14:07 | POST_SCALE_UP | Both production and test paths respond with v1 |
| 01:17:16 | TEST_TRAFFIC_SHIFT | Test rule changes to TG-A 100 / TG-B 0 (01:17:19). First v2 response on test path at 01:17:22. Production path remains v1 |
| 01:18:51 | POST_TEST_TRAFFIC_SHIFT | Test path serves v2, production path serves v1 |
| 01:18:56 | PRODUCTION_TRAFFIC_SHIFT | Production rule changes to TG-A 100 / TG-B 0 (01:19:05). First v2 response on production path at 01:19:08 |
| 01:20:31 | BAKE_TIME | Both production and test paths respond with v2 |
| 01:22:33 | CLEAN_UP | Both production and test paths respond with v2 |
| 01:22:38 | SUCCESSFUL | TG-B is DRAINING at 01:23:33, empty at 01:23:46 |
There was 1 minute and 46 seconds between the first v2 response on the test path at 01:17:22 and the first v2 response on the production path at 01:19:08.
From the execution of update-service (01:12:39) to the completion of the service deployment (completion time returned by the API: 01:22:37) was 9 minutes and 58 seconds. The PRIMARY rolloutState was recorded as COMPLETED at 01:24:26, which is after SUCCESSFUL was observed at 01:22:38.
The wait times between stages were: 3 minutes 9 seconds from POST_SCALE_UP to TEST_TRAFFIC_SHIFT; 1 minute 35 seconds each from TEST_TRAFFIC_SHIFT to POST_TEST_TRAFFIC_SHIFT and from PRODUCTION_TRAFFIC_SHIFT to BAKE_TIME; and 2 minutes 2 seconds from BAKE_TIME to CLEAN_UP. These correspond to the approximately 3 minutes after POST_SCALE_UP, approximately 90 seconds after each weight change, and the configured bakeTimeInMinutes of 2 minutes, as described in the official documentation.
Below are excerpts of one response body each from v1 and v2 (account ID, task ID, and container IP address have been replaced).
app_version=v1
task_arn=arn:aws:ecs:ap-northeast-1:<account-id>:task/lattice-bg/<task-id>
taskdef_revision=1
availability_zone=ap-northeast-1a
container_ip=<container-ip>
nginx=nginx version: nginx/1.27.5
app_version=v2
task_arn=arn:aws:ecs:ap-northeast-1:<account-id>:task/lattice-bg/<task-id>
taskdef_revision=2
availability_zone=ap-northeast-1c
container_ip=<container-ip>
nginx=nginx version: nginx/1.28.3
All 112 requests sent at approximately 13-second intervals (56 to production and 56 to test, for a total of 112) returned HTTP 200, with no error responses or connection errors. The task ARNs included in the production path responses were 2 for v1 and 2 for v2.
Summary
The built-in blue/green deployments in ECS, released in July 2025, now support Lattice integration with this update.
ECS built-in deployments have been improving with enhancements such as strengthened monitoring, making them easier to use.
If you have been using an ALB alongside your Lattice-based ECS services solely for blue/green deployments, please give this update a try.