Skip to main content
The Cloud Run service is agent-os, the Cloud SQL instance is agentos-db, and the Artifact Registry repo is agentos. The scripts target the current gcloud project and us-central1; override them with GCP_PROJECT_ID and GCP_REGION.

Manage

Production auth

Token-Based Authorization is on by default. Without a JWT_VERIFICATION_KEY or JWT_JWKS_FILE, the app refuses to serve traffic in production. The platform’s job is to keep your data private, so the safe default is refuse to start. Token-Based Auth gives you three things:
  1. No public access. The server rejects requests without a valid token.
  2. Per-request identity. Middleware parses the token and extracts the user_id, session_id, and custom claims. Each request is tied to a user and session, giving you auditability and traceability.
  3. Granular permissions. User tokens can run an agent and view their own sessions. Admin tokens read everyone’s sessions and test any agent.
To opt out (not recommended), set authorization=False in app/main.py and redeploy. Use this only inside a private VPC behind another auth layer. Without it, anyone who guesses your Cloud Run URL can access your platform.

Customize

Ask your coding agent to run /create-new-agent, or do it by hand. Create agents/my_agent.py:
Register it in app/main.py:
Local containers hot-reload on save. For production, run ./scripts/gcp/redeploy.sh.
app/settings.py defines default_model(), used by every agent. Change it in one place:
Add anthropic to pyproject.toml, set the provider key in your env, and regenerate pins:
Rebuild locally with docker compose up -d --build. For production:
Agno ships 100+ toolkits. See Toolkits.
  1. Edit pyproject.toml.
  2. Regenerate pins: ./scripts/generate_requirements.sh (add upgrade to refresh every pin).
  3. Rebuild locally with docker compose up -d --build, or redeploy with ./scripts/gcp/redeploy.sh.
Set both variables in your env file:
Sync with ./scripts/gcp/env-sync.sh. The interface activates automatically and routes messages to Agent Builder; change the agent= argument in app/main.py to point at another agent. See Slack setup.
The deployment check runs daily by default (ENABLE_DEPLOY_CHECK=True); it is deterministic and free. Scheduled evals are off by default (ENABLE_SCHEDULED_EVALS=False) because they use model calls. Both workflows stay runnable on demand regardless.

Format, validate, and run evals

The format, validate, and eval scripts run on the host and need a venv. Set it up once:
./scripts/mcp_check.sh runs inside the container, so it needs no venv.

Environment variables

Troubleshooting

Install the Google Cloud SDK, then run gcloud auth login and gcloud config set project <id>.
Expected. The script deploys first because Cloud Run only reveals the URL once the service exists, then pauses so you can mint the key against it. Mint the key at os.agno.com: connect your OS (Connect OSLive, enter your Cloud Run URL), then turn on Token-Based Authorization (JWT) under SettingsOS & Security and paste the full PEM. To do it later, skip the prompt, add JWT_VERIFICATION_KEY or JWT_JWKS_FILE to .env.production, and run ./scripts/gcp/env-sync.sh.
JWT auth is on whenever RUNTIME_ENV is not dev. Set JWT_VERIFICATION_KEY or JWT_JWKS_FILE and sync. To opt out inside a private VPC behind another auth layer, set authorization=False in app/main.py.
Your organization likely enforces Domain Restricted Sharing (constraints/iam.allowedPolicyMemberDomains), which silently rejects the allUsers binding that --allow-unauthenticated needs. The deploy still succeeds; the service just ships private. up.sh prints a warning when it detects this. Grant roles/run.invoker to specific principals, or ask an org admin for an allUsers exception.
Check that billing is enabled on the project. up.sh warns when it can’t confirm billing; Cloud SQL and Cloud Run creation both fail without it.
AGENTOS_URL is still the localhost default. up.sh sets it to your Cloud Run URL automatically; for a custom domain or tunnel, set it by hand and run ./scripts/gcp/env-sync.sh.
down.sh is targeting a different region than the one you deployed to. up.sh records GCP_REGION in your env file and down.sh reads it from there; if the file is gone, rerun with GCP_REGION=<region> ./scripts/gcp/down.sh. A wrong-region teardown looks clean while the real resources keep billing.