Skip to main content
The template names the Express service agent-os, the ECR repo agentos, and the RDS instance agentos-db. Secrets live under agentos/* in Secrets Manager, and the scripts record the service ARN and region in tmp/agentos-aws.state.

Manage

Production auth

Token-Based Authorization is on by default. Without a JWT_VERIFICATION_KEY or JWT_JWKS_FILE, the app refuses to serve traffic in production. The platform’s job is to keep your data private, so the safe default is refuse to start. Token-Based Auth gives you three things:
  1. No public access. The server rejects requests without a valid token.
  2. Per-request identity. Middleware parses the token and extracts the user_id, session_id, and custom claims. Each request is tied to a user and session, giving you auditability and traceability.
  3. Granular permissions. User tokens can run an agent and view their own sessions. Admin tokens read everyone’s sessions and test any agent.
Only /health stays open. AgentOS serves it unauthenticated even in production, so the ALB health checks pass. To opt out (not recommended), set authorization=False in app/main.py and redeploy. Use this only inside a private VPC behind another auth layer. Without it, anyone who guesses your service URL can access your platform.

Customize

Ask your coding agent to run /create-new-agent, or do it by hand. Create agents/my_agent.py:
Register it in app/main.py:
Local containers hot-reload on save. For production, run ./scripts/aws/redeploy.sh.
app/settings.py defines default_model(), used by every agent. Change it in one place:
Add anthropic to pyproject.toml, set the provider key in your env, and regenerate pins:
Rebuild locally with docker compose up -d --build. For production:
It rebuilds the image and syncs .env.production in one pass.
Agno ships 100+ toolkits. See Toolkits.
  1. Edit pyproject.toml.
  2. Regenerate pins: ./scripts/generate_requirements.sh (add upgrade to refresh every pin).
  3. Rebuild locally with docker compose up -d --build, or redeploy with ./scripts/aws/redeploy.sh.
Set both variables in your env file:
Sync with ./scripts/aws/env-sync.sh; both land in Secrets Manager. The interface activates automatically and routes messages to Agent Builder; change the agent= argument in app/main.py to point at another agent. See Slack setup.
The deployment check runs daily by default (ENABLE_DEPLOY_CHECK=True); it is deterministic and free. Scheduled evals are off by default (ENABLE_SCHEDULED_EVALS=False) because they use model calls. Both workflows stay runnable on demand regardless.
ARM cuts the Fargate line item from about 70toabout70 to about 57 per month. Edit runtimePlatform in scripts/aws/task-def.json, change docker build --platform linux/amd64 to linux/arm64 in both up.sh and redeploy.sh, then run ./scripts/aws/redeploy.sh.

Format, validate, and run evals

The format, validate, and eval scripts run on the host and need a venv. Set it up once:
./scripts/mcp_check.sh runs inside the container, so it needs no venv.

Environment variables

Troubleshooting

Upgrade the AWS CLI (for example brew upgrade awscli) until aws ecs create-express-gateway-service help works. If credentials are the problem instead, run aws configure and confirm aws sts get-caller-identity succeeds.
The RDS instance deploys into the region’s default VPC. Create one with aws ec2 create-default-vpc, or adapt scripts/aws/up.sh to your own VPC.
Expected. Mint the key at os.agno.com: connect your OS (Connect OSLive, enter your service URL), then turn on Token-Based Authorization (JWT) under SettingsOS & Security and paste the full PEM. To do it later, skip the prompt, add JWT_VERIFICATION_KEY or JWT_JWKS_FILE to .env.production, and run ./scripts/aws/env-sync.sh.
JWT auth is on whenever RUNTIME_ENV is not dev. Set JWT_VERIFICATION_KEY or JWT_JWKS_FILE and sync. To opt out inside a private VPC behind another auth layer, set authorization=False in app/main.py.
First-time provisioning of the ALB, certificate, and DNS takes 10-25 minutes; up.sh waits through it. Past that window, the known first-run cause is freshly created IAM roles: Express’s async infrastructure calls get denied before the role policies propagate, and ECS never retries. up.sh detects this and recreates the service once; the second attempt provisions reliably. If it still stalls, inspect with aws ecs monitor-express-gateway-service --region <region> --service-arn <arn>, look for an AccessDenied CreateLoadBalancer event in CloudTrail, then delete the service and re-run ./scripts/aws/up.sh.
The gateway is up; the app is still starting. First boot pulls the image and waits for the database. Wait a few minutes and check aws logs tail /ecs/agent-os --follow.
AGENTOS_URL is still the localhost default. up.sh sets it to your service URL automatically; for a custom domain or tunnel, set it by hand and run ./scripts/aws/env-sync.sh.
The scripts resolve the service ARN from tmp/agentos-aws.state first, then from a SERVICE_ARN= line in .env.production or .env. On a fresh clone or a new machine, write the ARN of your Express service into the state file: printf 'SERVICE_ARN=arn:aws:ecs:...' > tmp/agentos-aws.state.
The commands are hitting the wrong region. The scripts use AWS_REGION if set, then the region recorded in tmp/agentos-aws.state, then us-east-1. Set AWS_REGION to the region you deployed to and re-run ./scripts/aws/down.sh.