I checked the reachable range of AWS services from Amazon MWAA Serverless tasks with and without a VPC
This page has been translated by machine translation. View original
Hello. This is Takeda from the Service Development Division.
PythonOperator and BashOperator are now available in Amazon MWAA Serverless. You can call AWS APIs with your own code from workflow tasks.
I confirmed the usage and snapshot behavior of code in the previous article.
During that validation, calling sts get-caller-identity via boto3 in a configuration without a VPC resulted in a ConnectTimeoutError. On the other hand, writing to S3 succeeded. MWAA Serverless makes VPC creation optional, so the question becomes: which AWS services can tasks reach without a VPC, and what needs to be set up in the VPC to reach services that are otherwise unreachable? I verified this by changing configurations.
What is reachable without a VPC
The documentation states that tasks run without internet access by default, and that if external connectivity is required, a VPC with internet access must be passed via NetworkConfiguration.
First, I checked whether each service was reachable without a VPC. I prepared Python tasks that call each service's API via boto3 once, with a connection timeout of 3 seconds and no retries, and BashOperator tasks that output name resolution results using getent hosts. Since the execution role does not have permissions for the target services, receiving AccessDenied means the service was reached.
# netprobe.py (excerpt)
import boto3
from botocore.config import Config
CFG = Config(connect_timeout=3, read_timeout=5, retries={"max_attempts": 0})
def probe():
calls = {
"sts": lambda: boto3.client("sts", config=CFG).get_caller_identity(),
"s3": lambda: boto3.client("s3", config=CFG).list_buckets(),
"sqs": lambda: boto3.client("sqs", config=CFG).list_queues(),
"logs": lambda: boto3.client("logs", config=CFG).describe_log_groups(limit=1),
"dynamodb": lambda: boto3.client("dynamodb", config=CFG).list_tables(Limit=1),
}
results = {}
for name, fn in calls.items():
try:
fn()
results[name] = "ok"
except Exception as e:
results[name] = f"{e.__class__.__name__}"
print(results)
return results
The results are as follows.
| Destination | Result | Name Resolution |
|---|---|---|
STS (sts.ap-northeast-1.amazonaws.com) |
ConnectTimeoutError |
Public IP |
STS (global sts.amazonaws.com) |
ConnectTimeoutError |
Public IP |
| S3 | AccessDenied (reached) |
Public IP |
| SQS | ConnectTimeoutError |
Public IP |
| CloudWatch Logs | AccessDeniedException (reached) |
Private IP (10.x) |
| DynamoDB | ConnectTimeoutError |
Public IP |
Internet (checkip.amazonaws.com) |
Timeout | Public IP |
Only CloudWatch Logs resolved to a private IP, while S3 resolved to a public IP but was still reachable. The worker's /etc/resolv.conf showed nameserver 10.0.0.2 and search ap-northeast-1.compute.internal. This is the same configuration seen with the VPC's Amazon-provided DNS. This result makes sense if the service management side has a gateway endpoint for S3 and an interface endpoint for CloudWatch Logs (with private DNS enabled).
If the required destination cannot be reached, specify your own VPC in NetworkConfiguration and prepare a route using a NAT gateway or VPC endpoints.
Passing your own VPC
Common requirements and routing options
The network requirements from the documentation are as follows.
The three common requirements are:
- Two or more private subnets in different AZs (must not have a route to an internet gateway)
- A security group (with self-referencing inbound allow and outbound allow for all traffic)
- A network ACL (the documentation's recommended example is to allow all inbound and outbound, which is the default ACL state)
Choose one of the following for routing:
- Public routing with a NAT gateway (the documentation recommends one per public subnet)
- Private routing with VPC endpoints per service used
The configuration I created has a VPC of 10.0.0.0/16, two public subnets (for NAT) and two private subnets (10.0.10.0/24, 10.0.11.0/24).
The security group allows self-referencing inbound traffic. The outbound rule, set to allow all traffic at creation time, is left as-is.
SG=$(aws ec2 create-security-group --group-name mwaa-sls-sg \
--description "mwaa serverless" --vpc-id vpc-xxxxxxxx \
--query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id "$SG" --protocol -1 --source-group "$SG"
Create a gateway endpoint for S3 and an interface endpoint for STS. Enable private DNS for the interface endpoint and associate it with the two private subnets and the security group above.
# S3 (gateway endpoint)
aws ec2 create-vpc-endpoint --vpc-id vpc-xxxxxxxx \
--service-name com.amazonaws.ap-northeast-1.s3 --vpc-endpoint-type Gateway \
--route-table-ids rtb-private-a rtb-private-c
# STS (interface endpoint, private DNS enabled)
aws ec2 create-vpc-endpoint --vpc-id vpc-xxxxxxxx \
--service-name com.amazonaws.ap-northeast-1.sts --vpc-endpoint-type Interface \
--subnet-ids subnet-private-a subnet-private-c \
--security-group-ids "$SG" --private-dns-enabled
The endpoint policy is left as the default (full access) in this case. Be careful if you want to restrict the S3 gateway endpoint policy to only your own buckets. The policy example in the documentation states that access to prod-<region>-starport-layer-bucket (a bucket used for retrieving Amazon ECR image layers) is required. Restricting to only your own buckets may prevent the worker from pulling container images.
Passing the VPC to the workflow
Specify the created private subnets and security group in --network-configuration of update-workflow (or create-workflow). The other workflow settings (definition, code, and role) must be specified each time as well.
aws mwaa-serverless update-workflow \
--workflow-arn arn:aws:airflow-serverless:ap-northeast-1:123456789012:workflow/net-probe-xxxxxxxxxx \
--definition-s3-location '{"Bucket":"amzn-s3-demo-bucket","ObjectKey":"definitions/net.yaml"}' \
--code '{"S3Location":{"Bucket":"amzn-s3-demo-bucket","ObjectKey":"code/netprobe.zip"}}' \
--role-arn arn:aws:iam::123456789012:role/mwaa-sls-exec \
--network-configuration '{"SecurityGroupIds":["sg-xxxxxxxx"],"SubnetIds":["subnet-private-a","subnet-private-c"]}'
Incidentally, running update-workflow with an empty array specified in NetworkConfiguration for cleanup resulted in a ValidationException. I was unable to confirm how to revert an existing workflow to no VPC.
Verifying with different configurations
I ran the same workflow with only the VPC-side configuration changed.
| Configuration | STS | S3 | SQS / DynamoDB | CloudWatch Logs API / Task logs | Internet |
|---|---|---|---|---|---|
| No VPC | ✗ | ○ | ✗ | ○ / present | ✗ |
| Own VPC + NAT + STS endpoint | ○ (via endpoint) | ○ | ○ | ○ / present | ○ |
| Own VPC + NAT only | ○ (via NAT) | ○ | ○ | ○ / present | ○ |
| Own VPC + STS endpoint + S3 gateway only (no NAT) | ○ (via endpoint) | ○ | ✗ | ✗ / absent | ✗ |
| Above + CloudWatch Logs endpoint | ○ | ○ | ✗ | ○ / present | ✗ |
NAT + STS endpoint
In the first run with the VPC passed, the worker was running at 10.0.10.243, within the specified private subnet. sts.ap-northeast-1.amazonaws.com resolved to the IPs of the endpoint's ENIs (10.0.10.161 and 10.0.11.213). In this state, get_caller_identity succeeded. SQS, DynamoDB, and the internet were reached via NAT.
NAT only
After deleting the STS endpoint and re-running, sts.ap-northeast-1.amazonaws.com resolved to a public IP, and get_caller_identity succeeded. Even without an endpoint, NAT is sufficient to reach it.
Endpoint only (no NAT)
I removed the 0.0.0.0/0 route from the private route table, recreated the STS endpoint, and ran it again. The only routes were local and the S3 gateway.
Even with this configuration, the workflow completed with SUCCESS, STS was reached via the endpoint, and XCom returned values. The documentation lists VPC endpoints for Apache Airflow in addition to the services used by the workflow as requirements for private routing. However, the endpoint service names for Serverless are not published, and status updates and XCom returns completed without creating them. They appear to be unnecessary at this time.
On the other hand, only this run did not create a task log stream in CloudWatch Logs. The run succeeded, but there were no logs. After adding the CloudWatch Logs interface endpoint and re-running, logs.ap-northeast-1.amazonaws.com resolved to the endpoint's IP. In that case, the log stream was created. Based on the two results, it can be determined that a route from your own VPC to CloudWatch Logs is required to output task logs.
In the no-NAT configuration I tested, the three things I set up were: the S3 gateway endpoint, the CloudWatch Logs interface endpoint, and the endpoint for the service called from the code (STS). Since the workflow succeeds even without logs appearing, the CloudWatch Logs endpoint is easy to forget.
That requirement is not explicitly stated in CloudWatch Logs. The documentation for traditional MWAA (provisioned type) lists com.amazonaws.<region>.logs in its list of required endpoints. I read this as outputting task logs being included in "services used by the workflow."
NAT or interface endpoint?
After deciding to pass a VPC, whether to use a NAT gateway or interface endpoints (PrivateLink) for routing depends on the destination and cost.
The decision order is as follows:
- If the required destinations can be reached without a VPC, do not pass a VPC (no additional cost)
- If internet access is needed, or if the destination does not support VPC endpoints, use a NAT gateway
- If destinations are fixed and all support VPC endpoints, use interface endpoints (including the CloudWatch Logs endpoint)
- If data processing volume is large, compare both hourly costs and data processing costs
Let me compare pricing using Tokyo region values. Monthly figures assume 730 hours and are rough estimates excluding data processing charges proportional to volume and standard data transfer fees.
| Route | Hourly charge | Data processing charge | Estimated monthly cost |
|---|---|---|---|
| Interface endpoint | 0.014 USD/hour/AZ | 0.01 USD/GB | ~20 USD/service (2 AZs) |
| NAT gateway | 0.062 USD/hour/unit | 0.062 USD/GB | ~90 USD (2 units), ~45 USD (1 unit) |
| S3 gateway endpoint | No charge | No charge | 0 USD |
Interface endpoints for 2 services (STS and CloudWatch Logs) cost approximately 41 USD/month, and 3 services cost approximately 61 USD/month. The break-even point against the documentation's recommended 2-NAT configuration (~90 USD/month) falls between 4 and 5 interface endpoint services. Compared to 1 NAT (~45 USD/month), the break-even falls between 2 and 3 services, but 1 NAT differs in availability from 2-AZ endpoints.
Since the data processing charge is lower for interface endpoints (0.01 USD/GB vs. 0.062 USD/GB), interface endpoints become more advantageous as traffic to supported services increases.
In this case, following the documentation, I associated interface endpoints with two private subnets.
Summary
- Among the destinations tested, S3 and CloudWatch Logs were reachable without a VPC; STS, SQS, DynamoDB, and the internet were not confirmed reachable
- When a VPC is passed, the worker runs within the specified private subnets. STS was reached both via NAT and via interface endpoint
- In private routing without NAT, task logs are not retained without a CloudWatch Logs interface endpoint. This is easy to miss since the workflow still succeeds
- Excluding data processing volume, 2-AZ interface endpoints for 4–5 services are roughly equivalent in cost to 2 NAT gateways. The break-even against 1 NAT is between 2–3 services, but availability differs
What was confirmed in this article about reachability without a VPC covers only the 5 services tested and the internet. Test reachability to the services your tasks will call before deciding whether to set up your own VPC and routes.
