Production readiness outcome
A protected, recoverable deployment with known limits
Complete this guide before admitting production users or running business-critical Agents.
- Production and staging state are isolated.
- Deployment, runtime access, and stable keys have explicit owners.
- Data can be restored from an externally managed recovery set.
- Releases pass through staging, backup, and post-deploy verification.
- Operators understand the default single-replica availability boundary.
Overview
A running Kubernetes installation is the starting point. Production readiness adds controls that the default Helm charts do not create automatically.
This guide assumes you have completed the self-hosted configuration, deployed Agent Barn to Kubernetes, and verified the API, UI, databases, ingress, and Agent lifecycle.
Prepare your installation
Use values from your own infrastructure when configuring Agent Barn. The example domains, registry addresses, and storage classes in repository configuration files may describe an AAI Labs environment.
Before deploying, gather the following information with your infrastructure administrator:
| What you need | Where it comes from |
|---|---|
| Kubernetes access and the target namespace | Your cluster administrator. The default namespace is agent-farm; using another namespace also requires compatible bootstrap resources. |
| Container image locations, versions, and any required pull credentials | The release or image distribution path you will use. Public source repositories do not guarantee anonymous registry access. |
| API and web application hostnames | Domains you control, with DNS directed to your ingress. |
| TLS certificate configuration | A cert-manager ClusterIssuer configured for your cluster and domains. |
| Persistent storage | A StorageClass available in your cluster that meets your data durability requirements. |
| Database passwords, signing and encryption keys, and the initial administrator account | Values created and managed for this installation. |
| Model-provider credentials | Your provider account and the model settings you intend to use. |
| Email delivery settings | Your configured sending account and domain, if the installation will deliver invitations and password-reset emails. |
Configure the deployment
- Open the deployment configuration for the release or checkout you are using. For the repository's
deploy.shpath, copy.env.deploy.specto.env.deployand fill in that file. - Set the cluster access, image locations and versions, application hostnames, certificate issuer, and storage values for your installation. Do not leave a company example address in place unless you intentionally use that service and have access to it.
- Configure the application secrets and initial administrator account using the requirements in Deployment configuration. Keep signing and encryption keys stable across ordinary upgrades.
- Review the identity the API will use to manage Agent workloads. A separate restricted identity is not created automatically by leaving the pod kubeconfig setting empty.
- Continue with Deploy to Kubernetes for the deployment procedure and its prerequisites.
These steps prepare configuration; they do not create the cluster, issue registry credentials, or provision a Kubernetes identity.
For the AAI Labs cluster and release workflow, see AAI Labs hosted-service operations. Independent installations should use their own cluster, domains, storage, and credentials.
Production baseline
The repository supplies a functional single-cluster platform. The matrix makes the remaining operator responsibilities explicit.
| Requirement | Default support | Operator action |
|---|---|---|
| Environment separation | Production and staging namespaces | Supply distinct keys, databases, hosts, and kubeconfigs |
| TLS ingress | Traefik and cert-manager integration | Configure DNS, issuer, and certificate monitoring |
| Persistent databases | Single-replica PostgreSQL StatefulSets | Select durable storage and external backups |
| Agent workspaces | One PVC per Agent | Back up PVCs when continuity is required |
| Background processing | Redis, worker, and reconciliation CronJob | Monitor worker, Redis, and Event Delivery health |
| Monitoring | Prometheus, Grafana, and Alertmanager | Route alerts and add cluster-capacity monitoring |
| Stable secrets | Kubernetes Secrets | Keep authoritative copies in a protected secret manager |
| High availability | Not provided | Design and validate a custom HA architecture if required |
| Autoscaling | Not provided | Capacity-plan and add tested controls if required |
| Network isolation | Internal ClusterIP Services | Add cluster-compatible NetworkPolicies or equivalent controls |
| Database recovery | Not provided | Configure retention, encryption, backups, and restore tests |
| Deployment approval | Branch-triggered workflow | Protect main and require review before merge |
Before you begin
Assign a production owner, security contact, and on-call operator. Document availability, recovery point, recovery time, retention, maintenance-window, and provider-budget requirements.
Estimate concurrent Agents, message and Tool Call volume, storage growth, and provider rate limits. Prepare:
- A staging namespace and environment-specific kubeconfigs
- A protected secret manager
- A backup system compatible with PostgreSQL and the selected StorageClass
- Production DNS control and an operator-monitored alert destination
- A tested release artifact or protected deployment workflow
Choose your deployment environment
Choose the cluster and namespace where this installation will run. Record its API and web application hostnames, storage class, image registry, and credential owners.
If you maintain separate test and production installations, give each an explicit configuration and identify which data and credentials belong to it. The supplied deploy.sh uses bootstrap resources for agent-farm; choosing another namespace requires matching bootstrap resources and permissions.
Use Deployment configuration to set the inputs for your installation and Deploy to Kubernetes for the deployment procedure.
Keep each environment on its own deployment target, with its own namespace, credentials, hosts, and kubeconfig. Do not share signing keys, encryption keys, or database passwords between a test environment and production. Namespace isolation is not a separate cluster or provider security boundary: use distinct application and LiteLLM database passwords, signing and encryption keys, bootstrap administrators, kubeconfigs, hosts, and sender subdomains.
Verify that each installation's API receives its own release namespace as K8S_NAMESPACE; otherwise a test installation can create Agent workloads in the production namespace.
kubectl get deployment agentbarn-api \
--namespace <your-namespace> \
-o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="K8S_NAMESPACE")].value}{"\n"}'Expected: the namespace of the installation you are checking.
Use an authenticated image registry you control and pin production versions rather than deploying from a moving tag.
AAI Labs' cluster names, release workflows, and PUBLIC_* or STAGING_* variables describe its hosted environments. Maintainers of those installations should use AAI Labs hosted-service operations.
Protect stable keys and credentials
Keep authoritative platform deployment Secrets in a protected secret manager: database passwords, signing and encryption keys, registry password, LiteLLM master key, OpenRouter and Firecrawl keys, Platform Administrator bootstrap credentials, Grafana administrator password, and alerting webhooks. These are distinct from Communication Connection credentials.
| Stable secret | Restore with |
|---|---|
SECRET_SIGNING_KEY | Active Agent Barn authentication sessions |
AGENT_TOKEN_ENCRYPTION_KEY | The application database containing encrypted credentials and Agent keys |
LITELLM_MASTER_KEY | The LiteLLM database and virtual-key state |
| PostgreSQL credentials | Existing initialized PostgreSQL volumes |
Registry, OpenRouter, Cloudflare, Google OAuth, Slack webhook, and Grafana credentials are usually independently rotatable. Test emergency rotation without changing the stable recovery set.
Secret isolation and Connection credentials
Give each installation its own PostgreSQL passwords, signing and encryption keys, registry password, LiteLLM master key, Firecrawl key, Platform Administrator credentials, Grafana administrator password, alerting webhook, and model-provider key where quota isolation is required. Share an email account and token, OAuth application credentials, database usernames and names, model defaults, or allowlist configuration between installations only deliberately.
Slack bot and app-level tokens, Microsoft Teams App ID/client secret/Tenant ID, Telegram bot tokens, and Discord bot tokens are Communication Connection credentials. They are created through Agent Barn after deployment, encrypted in PostgreSQL, never returned by read APIs, and are not GitHub Actions deployment Secrets, Agent Runtime variables, Agent Secrets, or Shared Credentials.
AGENT_TOKEN_ENCRYPTION_KEY must be a valid Fernet key, stored securely, stable across ordinary deployments, shared by Product API and Communications, and included in recovery planning. Rotation is a data migration because Agent Secrets, Connection credential envelopes, and Runtime-related encrypted values depend on the current key.
Bootstrap administrator
PLATFORM_ADMIN_CREDENTIALS creates an administrator only when none exists. Updating the Secret later does not reset an existing account password. Create at least two independently controlled, named Platform Administrators and avoid routine use of the bootstrap account.
Restrict Kubernetes access
Use different identities for deployment and API runtime orchestration.
- Installs and upgrades releases
- Manages chart-owned resources
- Runs migrations and hooks
- Manages namespaced Agent workloads
- Reads Pods, health, and logs
- Supports approved exec and port-forward operations
The API identity needs namespaced access to create Agent Deployments, Services, Secrets, ConfigMaps, and PVCs, and to inspect Pods, health, logs, exec, and port-forward operations. It should not access staging, unrelated namespaces, cluster Secrets, or nodes.
Prepare the hook's ServiceAccount
The LiteLLM key setup hook runs under a configured Kubernetes ServiceAccount. In the supplied chart, its default name is agent-farm-user. Your cluster administrator must provision the account and the Secret permissions required by the hook in the installation's namespace.
This hook identity is separate from the kubeconfig used by the API to manage Agent workloads. Follow Arrange Kubernetes access before deployment; do not assume that provisioning the hook account creates the API credential or all of its workload permissions.
Verify that the account configured through LITELLM_KEY_SERVICE_ACCOUNT exists in the installation's namespace.
kubectl get serviceaccount \
--namespace agent-farm \
agent-farm-userAudit effective permissions
kubectl auth can-i create deployments --namespace agent-farm
kubectl auth can-i create services --namespace agent-farm
kubectl auth can-i create secrets --namespace agent-farm
kubectl auth can-i create persistentvolumeclaims --namespace agent-farm
kubectl auth can-i get pods --namespace agent-farm
kubectl auth can-i get pods/log --namespace agent-farmRepeat the checks with impersonation or the intended kubeconfig, then confirm that the same identity cannot read another namespace.
Plan storage and backups
The default deployment requests at least 19 GiB before any Agents: 5 GiB for application PostgreSQL, 2 GiB each for LiteLLM and Firecrawl PostgreSQL, and 10 GiB for Prometheus. Every Agent PVC requests another 1 GiB.
A node-local StorageClass such as local-path keeps data on a single node. Choose storage that meets the production durability target. Changing STORAGE_CLASS on an existing PostgreSQL StatefulSet is not an in-place migration because its volume claim template is immutable.
Set STORAGE_CLASS to a replicated, node-independent StorageClass for production. Durable PostgreSQL storage must survive node failure; treat local-path as temporary where one node is a failure domain. Communication data durability follows application PostgreSQL persistence, while Agent workspace PVC behavior remains separate from Communication Delivery persistence.
Back up the complete recovery set
Agent Barn does not install a database or Agent PVC backup controller. Integrate an external backup system, define frequency, retention, off-cluster storage, encryption, access, RPO/RTO, restore order, verification, and failure ownership.
Redis has no persistent volume in the current chart. Durable Domain Event intent remains in PostgreSQL, and reconciliation can republish eligible Event Deliveries when Redis returns.
Secure networking and TLS
| Surface | Exposure | Route |
|---|---|---|
| Web application | Public | / |
| Product API | Public | /api |
| Grafana | Operator-only when needed | Dedicated host |
| Metrics and Ingest | Internal | No public ingress |
| PostgreSQL, Redis, LiteLLM, Firecrawl | Internal | ClusterIP Services |
Public hosts and Communications paths
| Setting | Purpose |
|---|---|
API_HOST | Public Product API and provider-webhook hostname |
UI_HOST | Agent Barn web application hostname |
WEB_APP_URL | Full HTTPS URL of the web application |
GRAFANA_HOST | Product monitoring hostname |
SENDER_EMAIL | Transactional email sender |
The Product API’s API_EXTERNAL_URL resolves from the public API host. Connection webhook URLs use https://<public-api-host>/communications/v1/webhooks/<connection-id>; never configure provider webhooks against the UI hostname or an internal ClusterIP address.
| Source | Destination | Requirement |
|---|---|---|
| Browser | UI | Public HTTPS |
| UI and API clients | Product API :8000/api/v1 | Public HTTPS or trusted application network |
| Agent Runtime | Ingest :8001/ingest/v1 | Private cluster network |
| Agent Runtime | Communications :8002/communications/v1 | Private cluster network |
| Provider | Communications webhook prefix | Public HTTPS |
| Communications | PostgreSQL | Private database network |
| Communications | Provider APIs and Gateways | Controlled outbound internet |
| Prometheus | Communications :8002/metrics | Private monitoring network |
| Kubernetes probes | Communications :8002/health | Internal |
Expose only /communications/v1/webhooks from Communications. Runtime Delivery and Platform Driver routes, /health, /metrics, and Ingest routes remain internal. Provider webhooks require public DNS, valid certificates, stable ingress routing, the correct public API origin, provider authentication, and Connection-scoped paths. Microsoft Teams currently depends on this authenticated webhook path.
UI_HOST=agentbarn.example.com
API_HOST=api.agentbarn.example.com
GRAFANA_HOST=grafana.agentbarn.example.com
WEB_APP_URL=https://agentbarn.example.comConfirm DNS resolution, ports 80 and 443, certificate issuance and renewal monitoring, API-to-Service routing, and restricted Grafana authentication.
Introduce network isolation in staging
The charts do not create NetworkPolicies. If the cluster enforces them, allow only required UI-to-API, service-to-database, worker-to-Redis, LiteLLM-to-provider, Firecrawl, Agent, and Prometheus flows. Test egress carefully: a bad policy can leave Agents running but unable to reach model or messaging providers.
Configure production providers
OpenRouter and LiteLLM
Use a production-dedicated OpenRouter key when isolation is required. Assign its budget, credit limit, rate limits, billing alerts, owner, rotation path, and emergency replacement procedure. The OpenRouterCreditsLow alert is useful only when the provider key has a credit limit; an unlimited key reports infinite remaining credit.
Keep LITELLM_MASTER_KEY stable and separate from the OpenRouter provider key.
Transactional email
SENDER_EMAIL=noreply@mail.agentbarn.example.comTransactional email requires the Cloudflare account ID, API token, and sender address, plus a verified sending domain. Verify invitations, password recovery, lifecycle notifications, reputation, and quota. Use a different sender subdomain for staging.
Production and staging share the Cloudflare account, token, and account quota in the current workflow, so rotation and quota exhaustion can affect both.
Google Workspace
https://agentbarn.example.com/api/v1/integrations/google/callbackRegister the exact production callback. If an OAuth client is shared with staging, register that callback too; separate clients provide stronger failure and consent-screen isolation.
Firecrawl
Treat the platform Firecrawl key as a production credential. Its Services remain internal, but Firecrawl makes outbound requests, so enforce appropriate egress, destination, abuse-prevention, and capacity policies.
Plan capacity and availability
API, UI, Communications, worker, LiteLLM, Redis, all PostgreSQL databases, each Firecrawl component, Prometheus, Grafana, and Alertmanager run as single replicas by default.
Each Agent adds a Deployment, pod, Service, 1 GiB PVC, CPU and memory demand, model usage, platform API traffic, and monitoring traffic. Agent resource profiles are not configurable through .env.deploy.
Load-test representative Hermes and OpenClaw workloads in staging, including concurrent Agents, peak tools, large model responses, browser or Firecrawl work, Template and Skill loading, and restart behavior.
Production Platform messaging requires a separate Communications Deployment and ClusterIP Service on port 8002, with /health probes, internal /metrics, a provider ingress supervisor, and an outbound Delivery worker. It uses the same API image tag as Product API while running a separate process.
Communications replicas coordinate supervised provider ingress through PostgreSQL-backed leases: one replica owns a supervised Connection at a time, webhook traffic may load-balance, and lease expiry enables failover. Every replica shares the same database and encryption configuration. Use queue depth, oldest queued age, and provider status to guide scaling; Connection revision changes reconcile provider sessions without restarting Agents.
If the availability target requires multiple replicas, replication, or failover, treat that as custom architecture and validation work, not an environment-variable toggle.
Protect Communications persistence
Production backups must cover the application PostgreSQL database containing Communication Connections, encrypted credential envelopes, Communication Deliveries, canonical Conversation Messages, operational journal entries, and Connection health and lease state. Durable Deliveries survive Communications pod replacement. Redis is not the Communication Delivery source of truth.
COMMUNICATION_JOURNAL_RETENTION_DAYS=31Values range from 1 through 3650; the supervisor prunes older content-free journal entries. Set retention for troubleshooting, storage, and compliance needs. Journal retention and Conversation history retention are separate, and restoring the database without the matching encryption key leaves credentials unreadable.
Configure monitoring and alerts
SLACK_ALERTS_WEBHOOK_URL=REPLACE_WITH_PRODUCTION_WEBHOOK
GRAFANA_ADMIN_PASSWORD=REPLACE_WITH_STRONG_PASSWORD
GRAFANA_HOST=grafana.agentbarn.example.comFor every alert, document severity, on-call owner, response target, diagnostic link, escalation, recovery action, and user-notification threshold. Trigger a controlled staging alert before go-live.
Understand monitoring scope
Prometheus requests a 10 GiB PVC and retains 15 days by default. The namespace-scoped configuration does not collect cluster-wide node, kubelet, or cAdvisor metrics.
Add cluster monitoring for nodes, filesystems, volumes, container CPU and memory, the control plane, ingress, and cert-manager. Grafana dashboards are operational views, not audit logs or business records.
Prometheus should scrape Communications internally from port 8002. Monitor error and degraded Connection counts, dead-lettered Delivery outcomes, queue growth, oldest queued age, Delivery latency, reconnect frequency, policy-rejection changes, and Communications pod readiness and restarts.
agentbarn_communication_connection_statusagentbarn_communication_delivery_outcomesagentbarn_communication_queue_depthagentbarn_communication_oldest_queued_age_secondsagentbarn_communication_delivery_latency_secondsagentbarn_communication_reconnectsagentbarn_communication_policy_dispositions
These metrics intentionally exclude Organization, Agent, Connection, Conversation, and User identifiers. Prometheus and alerts detect platform-wide changes; Communication diagnostics supports Agent-scoped investigation, the Communications journal provides content-free stage history, Agent logs describe Runtime behavior, Conversation history contains accepted message content, and Security Audit records selected recovery and health events. Do not place message content, provider tokens, or raw exception text in metrics or alerts.
Prepare disaster recovery
Write and schedule a recovery exercise using this minimum sequence:
- Provision an isolated recovery namespace or cluster.
- Restore stable keys and the three PostgreSQL databases.
- Restore required Agent PVCs and compatible deployment artifacts.
- Deploy without changing the restored keys and verify the expected schema.
- Verify authentication, LiteLLM virtual keys, and a model request.
- Verify Shared Credentials and Agent credentials can decrypt.
- Start a test Agent and verify messages, Tool Calls, and costs.
- Verify Event Delivery reconciliation, monitoring, and alerts.
Preserve version compatibility
Record the git commit, API/UI and runtime image tags, chart versions, Alembic revision, PostgreSQL and LiteLLM versions, backup time, and secret-manager references with every recovery set. Do not restore into an arbitrary application version.
Plan migration rollback
Database migrations run as a Helm pre-install/pre-upgrade hook. Helm rollback changes manifests and images; it does not reverse Alembic migrations. Before schema changes, review compatibility, back up the application database, test upgrade and recovery in staging, and decide whether rollback uses an older compatible image or a database restore.
Rollback planning also covers release-tagged Product and Communications images, Runtime base-image versions, Communications protocol compatibility, Platform Plugin behavior, pending and dead-lettered Deliveries, and provider-session reconciliation. API and Communications normally roll back together because they use the same API image tag; durable Deliveries remain in PostgreSQL across replacement. Do not delete Connection or Conversation data, reuse a release tag for different image contents, or assume database migrations are reversible.
Stage and verify releases
In staging, verify API and UI health, administrator login, Organization isolation, invitations and email, OAuth, Agent creation, Hermes and OpenClaw startup, messaging platforms, LiteLLM requests, Activity, Tool Calls, logs, costs, Shared Credentials, Event Delivery processing, dashboards, alerts, migrations, and restart recovery.
Use dedicated staging chat applications, credentials, and data. Do not test with production workspaces or customer conversations.
Immediately before production
kubectl config current-context
kubectl get pods --namespace agent-farm
kubectl get pvc --namespace agent-farm
kubectl get certificate --namespace agent-farmConfirm the production backup completed, is readable, and is associated with the matching stable-key set.
Production Communications canary
- Confirm migrations and Product, Ingest, Communications, worker, and UI readiness.
- Confirm Communications
/healthand Prometheus scraping internally, then verify Runtime DNS to its ClusterIP Service. - Create or use a dedicated canary Agent and add a canary Communication Connection.
- Confirm provider credential validation, start the Agent, and send an inbound message that passes policy.
- Confirm the pipeline advances from Provider observed through Provider delivered, with the Conversation under the correct Connection.
- Confirm a Tool Call appears through Ingest separately and no queue-age, dead-letter, or Connection-error alert is introduced.
- Where the Platform uses webhooks, confirm the public provider webhook path works.
A Product API health check alone is not a sufficient canary. Hermes and OpenClaw share the Communications protocol: test at least one supported Runtime after Runtime or adapter changes, both after shared protocol changes, the affected Platform Plugin after provider changes, supervised ingress after Slack/Telegram/Discord changes, webhook ingress after Teams or ingress changes, and outbound delivery after provider-client changes.
Complete production go-live
Verify releases, workloads, storage, ingress, certificates, and the public API.
helm list --namespace agent-farmkubectl get deployments,statefulsets,pods,jobs,cronjobs \
--namespace agent-farmkubectl get pvc,ingress,certificate \
--namespace agent-farmcurl --fail https://api.agentbarn.example.com/api/v1/healthThen sign in with a named Platform Administrator, select the production Organization, invite a controlled test user, verify email enrollment, and start a non-sensitive test Agent. Send a message and verify the response, Activity, Tool Calls, cost attribution, Agent metrics, and alert state. Retire the test Agent according to policy and record the deployed commit and verification result.
Current availability constraints
| Area | Current default |
|---|---|
| API, UI, worker, LiteLLM | One replica each; replacements are non-overlapping |
| PostgreSQL | Three independent, single-replica StatefulSets |
| Redis | One replica without persistent Kubernetes storage |
| Firecrawl | One replica per component |
| Prometheus | One replica with 15-day retention |
| Grafana and Alertmanager | One replica each |
| Agent storage | One ReadWriteOnce PVC per Agent |
| Autoscaling and disruption budgets | Not configured |
| NetworkPolicies | Not configured |
| Backups and database replication | Not configured |
| Cross-cluster failover | Not configured |
These defaults fit a lean self-hosted platform whose operators accept single-node and maintenance risk. Stricter availability objectives require explicit architecture work.
Production checklist
Use these cards during the change review and final go-live call.
-
Identity and access
-
Secrets
-
Data and recovery
-
Networking
-
Providers
-
Operations
Security considerations
- Treat deployment credentials, workflow changes, stable keys, and backups as production security surfaces.
- Never mount cluster-admin access into the API or expose metrics, Ingest, databases, Redis, LiteLLM, or Firecrawl publicly.
- Use separate environment credentials when shared provider quota, revocation, or access is unacceptable.
- Restrict Grafana, encrypt backups, and test recovery access.
- Do not place credentials in Templates, Skill files, logs, alert annotations, or privilege reasons.
- Review chat-platform access, model allowlists, and provider budgets before connecting production channels.
- Restart Agents deliberately only when generated Runtime configuration changes; Communications reconciles Connection and provider-session changes independently.
- Treat stable-key rotation and StorageClass changes as migrations.
Planned Organization suspension and platform-audit capabilities are not substitutes for these production controls.
Troubleshooting
| Symptom | Likely cause | Resolution |
|---|---|---|
| A test installation creates Agent resources in production | K8S_NAMESPACE or the API kubeconfig points to the production namespace | Stop the affected Agents and correct namespace and kubeconfig wiring. |
| The production API can access unrelated namespaces | The mounted kubeconfig is too privileged | Replace it with a dedicated namespace-scoped API identity. |
| Database authentication fails after deployment | The configured password differs from the initialized volume password | Restore the prior value or perform a coordinated database rotation. |
| Restored credentials cannot decrypt | The wrong Agent encryption key was restored | Restore the key belonging to that application database backup. |
| Restored LiteLLM Agent keys fail | The wrong master key or LiteLLM database was restored | Restore the matching LiteLLM master key and database. |
| PVCs remain Pending | The StorageClass is unavailable or incompatible | Select a supported ReadWriteOnce StorageClass. |
| Data is lost after node failure | Node-local storage was used without backup or replication | Restore from backup and adopt an appropriate durable storage design. |
| Model requests fail during deployment | LiteLLM is being replaced without an overlapping replica | Wait for readiness and schedule future changes in a maintenance window. |
| An Agent still uses old runtime configuration | Existing Agent workloads were not rebuilt | Stop and start the affected Agent deliberately. |
| Event Deliveries remain Pending | Redis, the worker, or reconciliation is unavailable | Inspect Redis, worker readiness, and the reconciliation CronJob. |
| Prometheus lacks CPU or memory metrics | Namespace monitoring does not discover node or cAdvisor metrics | Add cluster-level infrastructure monitoring. |
| OpenRouter credit alerts are not useful | The provider key has no credit limit | Configure a provider limit and validate the metric. |
| Staging exhausts the email quota | Both environments share the Cloudflare account quota | Limit staging sends or separate provider accounts. |
| Helm rollback does not restore old behavior | The database schema remained migrated | Use the documented compatibility or database-restore procedure. |
| TLS remains unready | DNS, ingress, or the fixed ClusterIssuer is incorrect | Inspect Certificate, Challenge, Ingress, and ClusterIssuer resources. |
Next steps
- Manage database migrations before the next schema-changing release.
- Monitor the platform and assign alert ownership.
- Upgrade Agent Barn with staged verification and tested backups.
- Troubleshoot self-hosting when a production check fails.
- Operate the Communications service for provider ingress, Delivery processing, and diagnostics.
- Review Agent health and logs after starting production Agents.
Cost synchronization
Agent Barn's cost reports read stored cost records. In Helm deployments, the cost-sync CronJob imports LiteLLM spend logs and recovers eligible missing charges through OpenRouter. It is enabled by default with a 15-minute schedule and concurrencyPolicy: Forbid.
The job requires database access, the credential-encryption key used to attribute Agent keys, access to the LiteLLM master key through the configured Kubernetes Secret and kubeconfig, and OpenRouter access for missing-cost recovery. A LiteLLM virtual key is not sufficient for the spend-log endpoint.
Configure the schedule through costs.sync.schedule and enablement through costs.sync.enabled. The implementation's maximum run time is 600 seconds; keep the schedule interval longer than that bound. These are chart values and a code constant respectively, not interchangeable environment variables.
On an empty cost table, synchronization starts from historical logs. Later runs rewind the latest stored timestamp by one hour to include late-arriving records. Recovered charges are preserved when later LiteLLM pages still report zero. A report can therefore lag recent usage, and historical totals can rise as recovery progresses.
The job entrypoint, inside an appropriately configured application environment, is:
python -c "from api.domains.costs.sync import main; main()"
Use that entrypoint rather than python -m api.domains.costs.sync. Docker Compose and run.sh do not schedule this job automatically. Local cost reporting needs an explicit sync invocation with the same required connectivity and credentials; refreshing the Costs page does not perform synchronization.
See Costs and Spend Attribution for the stored cost model.
Agent Email
Inbound-secret rotation must update the Product API service and the inbound Worker for the same environment. The current single-secret contract has an interruption window; a secret change alone does not select the path-filtered Worker publication. See Configure Agent Email for the coordinated deployment procedure.
Agent Email needs Cloudflare Email Routing, an inbound Worker, sending access, and matching environment-specific inbound secrets. Transactional email alone does not enable it. Follow Configure Agent Email for setup and routing-rule ownership.