AWS Containerized Deployment
Deploy agents to AWS ECS Fargate for consistent, low-latency execution.
Two topologies are supported:
| Mode | Containers | Selected by | Best for |
|---|---|---|---|
| Simple REST | One (runs RESTAPI) | queue_mode = false (default) in Terraform; app entrypoint RESTAPI.run | Moderate traffic, simplest setup |
| Scalable queue mode | Two (IO container + Agent Runner service) | queue_mode = true in Terraform; entrypoints ECSIOHandler.run + ECSAgentRunner.run; execution.queues.* config | High throughput, long-running agents, backpressure control |
ECS containers serve JSON REST. SSE token streaming (execution.mode: stream) is available in the Simple REST topology (single container running RESTAPI) with streaming-capable frameworks. WebSocket async/stream execution modes are not available on ECS (AWS Lambda only). The scalable queue mode supports rest_sync and rest_async only.
Simple REST Architecture
A single ECS service runs the built-in FastAPI server; request handling and agent execution happen in the same container.
Entrypoint (app.py):
from agentkernel.api import RESTAPI
from agentkernel.openai import OpenAIModule
OpenAIModule([...])
runner = RESTAPI.run
if __name__ == "__main__":
runner()
Prerequisites
- Docker installed
- AWS CLI configured
- ECR repository created
- Agent Kernel with AWS extras
Deployment
Refer to example ECS implementation which leverages Agent Kernel's terraform module for ECS deployment.
Scalable Queue Mode
For high-throughput or long-running agents, use the two-container queue architecture. The IO container and the Agent Runner are separate ECS services sharing SQS FIFO queues, so request ingestion and agent execution scale independently.
Multi-Threading Design
Both containers are internally multi-threaded, managed by ThreadRunner (one daemon threading.Thread per task, gated by a semaphore, with uniform crash/shutdown handling):
- Each
sqs-consumer-*thread runs an independent blocking long-poll loop (poll → process → delete), checking a sharedThreadRunner.shutdown_eventbetween iterations. - Consumer thread counts are configured per queue:
execution.queues.input.no_of_consumers(default 5, Agent Runner) andexecution.queues.output.no_of_consumers(default 2, IO container). These settings are ECS-only; Lambda ignores them. execution.queues.batch_sizecontrolsMaxNumberOfMessagesper SQS receive call; it is injected by Terraform and should never be set inconfig.yaml.
Failure and shutdown propagation: if any consumer thread crashes, ThreadRunner triggers a graceful shutdown: it sets the shared shutdown_event, the sibling consumer threads finish their in-flight message and exit their loops, and then os._exit(1) is called so ECS restarts the task cleanly. The rest-api thread (uvicorn) never checks shutdown_event and is marked awaited_on_shutdown=False, so the drain doesn't wait on it; it is simply terminated when os._exit(1) fires. Per-message processing errors do not kill a consumer thread; the message is left undeleted and retried after the SQS visibility timeout.
Container 1: IO container (ECSIOHandler)
ECSIOHandler.run() starts two peer tasks via ThreadRunner:
rest-apithread:RESTAPI.run(handlers=[ECSQueueRequestHandler()]), FastAPI/uvicorn.POST /api/v1/chat: validatessession_id+prompt, generates arequest_id(UUID), and enqueues to the Input Queue withMessageGroupId = session_idandMessageDeduplicationId = request_id. Inrest_syncmode it then polls the response store for thatrequest_idand returns the reply on the same connection (504 if it never arrives); inrest_asyncmode it returns{"status": "ACCEPTED", "request_id": ...}immediately.GET /api/v1/chat/{session_id}?request_id=...(rest_asynconly, 404 otherwise): reads the response store, validates the stored reply'ssession_idmatches the path (so a reply can't be read under the wrong session), and returns the body orNOT_FOUNDwhile still processing.- The IO container registers no agents; agent validation and execution happen only in the Agent Runner.
output-queue-consumerthread:ECSOutputConsumer.run()spawnsexecution.queues.output.no_of_consumers(default 2) long-poll threads on the Output Queue, each writing{session_id, request_id, body}records to the response store. On permanent failure (message exceededmax_receive_count), it writes an error record to the store so the waiting HTTP caller gets an error instead of hanging.
Entrypoint (app_rest_service.py) (no agent definitions):
from agentkernel.aws import ECSIOHandler
runner = ECSIOHandler.run
if __name__ == "__main__":
runner()
Container 2: Agent Runner (ECSAgentRunner)
Extends ECSSQSConsumer (which in turn extends the shared QueueConsumer base also used by Lambda's LambdaSQSConsumer): runs execution.queues.input.no_of_consumers (default 5) independent threads, each polling the Input Queue in a blocking loop, executing the agent through the full Runtime.run() pipeline (hooks, guardrails, session persistence), and putting the result on the Output Queue with the same request_id. On permanent failure it forwards an error body to the Output Queue so the client still receives a response.
Entrypoint (app_agent_runner.py):
from agentkernel.aws import ECSAgentRunner
from agentkernel.openai import OpenAIModule
OpenAIModule([...]) # register agents here only
handler = ECSAgentRunner.run
if __name__ == "__main__":
handler()
Terraform
Enable queue mode in the yaalalabs/ak-containerized/aws module:
queue_mode = true
execution_mode = "sync" # or "async"
rest_service = {
package_path = "../dist-rest-service"
command = ["python", "app_rest_service.py"]
cpu = 256
memory = 512
desired_count = 1
}
agent_runner = {
package_path = "../dist-agent-runner"
command = ["python", "app_agent_runner.py"]
cpu = 1024
memory = 2048
desired_count = 1
environment_variables = {
OPENAI_API_KEY = var.openai_api_key
}
}
scaling_config = {
enabled = true
min_count = 1
max_count = 10
backlog_target = 5
scale_in_cooldown = 180
scale_out_cooldown = 60
}
For the full example see examples/aws-containerized/openai-dynamodb-scalable.
For queue mode internals see Queue Mode Guide.
Advantages
- No cold starts - containers always warm
- Consistent performance - predictable latency
- Better for high traffic - efficient resource usage
- Full control - customize container, resources, etc.
- High availability - multi-AZ deployment with automatic failover
- Fault tolerant - automatic recovery and health-based routing
Fault Tolerance
AWS ECS deployment provides comprehensive fault tolerance features with extensive configurability.
Multi-AZ Architecture
Tasks are automatically distributed across multiple Availability Zones:
Benefits:
- Survives entire AZ failures
- No single point of failure
- Automatic traffic distribution
- Geographic redundancy
Automatic Task Recovery
ECS Service maintains desired task count with automatic recovery:
Features:
- Failed tasks automatically restarted
- Desired count maintained at all times
- Rolling deployments with zero downtime
- Gradual task replacement during updates
Health Check Configuration
Application Load Balancer performs continuous health monitoring:
How it works:
- ALB sends requests to
/healthendpoint every 30 seconds - Unhealthy tasks removed from load balancer rotation
- Traffic routed only to healthy tasks
- Failed tasks replaced automatically
- Connection draining ensures graceful shutdown
Auto-Scaling for Resilience
ECS Service auto-scaling maintains capacity during failures and load spikes.
In queue mode, the Agent Runner scales automatically based on queue depth using a custom CloudWatch metric (Custom/ECS/BacklogPerTask). A Lambda function runs every minute, computes BacklogPerTask = QueueDepth / max(RunningTasks, 1), and a Target Tracking policy adjusts the task count to keep this metric at or below backlog_target.
Enable this in the scaling_config block; see Scalable Queue Mode for configuration details.
Auto-scaling triggers:
- Queue backlog per task (queue mode, recommended)
- CPU utilization
- Memory utilization
- Custom CloudWatch metrics
Benefits:
- Automatic capacity adjustment
- Handles traffic spikes
- Compensates for task failures
- Cost optimization during low traffic
Network Resilience
Connection Draining:
- Existing connections complete before task termination
- Configurable timeout (default 30 seconds)
- Prevents abrupt connection drops
- Graceful shutdown process
Load Balancer Features:
- Sticky sessions (optional) for stateful apps
- Cross-zone load balancing enabled
- Health-based routing
- Automatic DNS failover
Recovery Time Objectives
Typical Recovery Times:
- Task failure detection: 5-30 seconds (health check interval)
- Task replacement: 30-60 seconds (container startup)
- Traffic rerouting: Immediate (ALB handles)
- Total RTO: < 2 minutes for most failures
Recovery Point Objectives:
- With DynamoDB: Continuous (multi-AZ replication)
- With Redis Cluster: < 1 second (automatic failover)
Configuration Best Practices
Minimum Task Count: Run at least 2 tasks (3+ recommended)
ecs_desired_count = 3
Testing Fault Tolerance
Simulate Failures:
# Stop a task to test auto-recovery
aws ecs stop-task --cluster my-cluster --task task-id
# Kill a container to test health checks
docker stop container-id
# Simulate AZ failure (in test environment)
# Manually stop all tasks in one AZ
Validate:
- Tasks automatically restarted
- No service interruption
- Load balanced across remaining tasks
- Metrics show recovery
Learn more about fault tolerance →
Session Storage
For containerized deployments, use Redis, Valkey, or DynamoDB for session persistence.
For detailed session storage configuration and best practices, see the Session Management documentation.
ElastiCache Redis (Recommended for Performance)
export AK_SESSION__TYPE=redis
export AK_SESSION__REDIS__URL=redis://elasticache-endpoint:6379
export AK_SESSION__CACHE__SIZE=256 # Enable in-memory caching
Benefits:
- High performance (sub-millisecond latency)
- Low latency
- In-memory speed
- Shared cache across tasks
Use when:
- You need sub-millisecond latency
- High throughput requirements
- Already using Redis infrastructure
See Redis configuration details →
ElastiCache for Valkey
export AK_SESSION__TYPE=valkey
export AK_SESSION__VALKEY__URL=valkey://elasticache-endpoint:6379 # valkeys:// for SSL
export AK_SESSION__CACHE__SIZE=256 # Enable in-memory caching
Provision the cluster with create_valkey_cluster = true; the module injects
AK_SESSION__VALKEY__URL into the task definition. Note that the env var alone does not switch the
backend: the config.yaml baked into the image must set session.type: valkey. Valkey is the
open-source Redis fork, offered on ElastiCache at a lower price point and wire-compatible with
Redis. Requires the agentkernel[valkey] extra.
See Valkey configuration details →
DynamoDB (Serverless Option)
export AK_SESSION__TYPE=dynamodb
export AK_SESSION__DYNAMODB__TABLE_NAME=agent-kernel-sessions
export AK_SESSION__DYNAMODB__TTL=604800 # 7 days
Benefits:
- Serverless, fully managed
- No infrastructure management
- Auto-scaling
- Multi-AZ by default
Use when:
- Simplicity and low operational overhead preferred
- Variable workload patterns
- AWS-native infrastructure
- Moderate latency is acceptable (single-digit milliseconds)
See DynamoDB configuration details →
Requirements:
- DynamoDB table with partition key
session_id(String) and sort keykey(String) - ECS Task IAM role with DynamoDB permissions
- The Terraform module automatically creates the table and configures permissions when
create_dynamodb_memory_table = true
Monitoring
Use CloudWatch Container Insights:
- CPU/Memory utilization
- Task count
- Network metrics
- Application logs
Health Checks
Agent Kernel provides a health endpoint:
# Automatically available at /health
# Returns 200 OK if healthy
Application Endpoints
Users can expose their own API endpoints alongside the Agent Kernel endpoints without having to do any custom implementation. Refer to example.
Authentication
For production deployments, implement authentication to secure your Agent Kernel endpoints.
Environment Variable Configuration
You can configure environment variables in your Terraform configuration like shown below:
module "container_app" {
source = "yaalalabs/ak-containerized/aws"
# ... other configuration ...
environment_variables = {
# Authentication configuration
SOME_VALUE = "some-value"
# Other app configuration
OPENAI_API_KEY = var.openai_api_key
}
}
Application Implementation
In your containerized application, implement your custom authentication logic by extending the AuthValidator class. You may use the environment variables:
NOTE: the
validate()function must return aValidationResultobject.
import os
import jwt
from agentkernel.api import RESTAPI
from agentkernel.auth import AuthValidator, ValidationContext, ValidationResult
from typing import Optional
class CustomAuthValidator(AuthValidator):
def validate(self, token: str, context: Optional[ValidationContext] = None) -> ValidationResult:
"""Validate authentication token and return result."""
payload = jwt.decode(token, options={"verify_signature": False})
print("Payload", payload)
email = payload.get("email", "")
if email == "test@test.com":
return ValidationResult(is_valid=True, subject="test_user")
return ValidationResult(is_valid=False, error_msg="Invalid token")
# Add authentication to REST API
RESTAPI.add_auth_handlers(auth_validators=[CustomAuthValidator()])
Example Implementation
See examples/aws-containerized/crewai-auth for a complete authentication example.
Best Practices
- Use at least 2 tasks for high availability
- Configure auto-scaling based on traffic
- Use Redis for session persistence when latency is critical
- Use DynamoDB for session persistence for serverless-style infrastructure
- Enable Container Insights for monitoring
- Set up log aggregation
- Use secrets manager for API keys
- Implement authentication for all production deployments
- Use environment-specific authentication configurations
- Monitor authentication failures for security incidents
