Self-hosting
Guide

Deploy Agent Barn to Kubernetes

Deploy Product, Ingest, Communications, workers, Agent Runtimes, and monitoring to Kubernetes with Helmfile.

For
Platform administrators, DevOps engineers, and Kubernetes operators
On this page
  1. Agent Email
  2. What you will deploy
  3. Before you begin
  4. 1. Verify the cluster
  5. 2. Get the deployment files
  6. 3. Configure the deployment
  7. 4. Configure DNS and TLS
  8. 5. Deploy Agent Barn
  9. 6. Verify the deployment
  10. 7. Complete the first-time setup
  11. Operate Communications replicas and monitoring
  12. Upgrade the deployment
  13. Troubleshooting
  14. Deployment constraints

Deploy Agent Barn and its supporting services to an existing Kubernetes cluster using the Helm charts and Helmfile configuration included with the project.

This guide uses the production namespace agent-farm; staging uses agent-farm-staging. The repository and product are named Agent Barn, but these Kubernetes namespaces deliberately retain the earlier names as frozen infrastructure identifiers. The Communications Deployment and Service are created in the Helm release namespace.

What you will deploy

By the end of this guide, you will have:

  • Product API, Ingest API, Communications, the Domain Event worker, and the web application
  • PostgreSQL databases for Agent Barn, LiteLLM, and Firecrawl
  • Redis for background processing
  • LiteLLM connected to OpenRouter
  • A self-hosted Firecrawl service
  • Prometheus, Grafana, Alertmanager, and Agent Barn dashboards
  • TLS ingress for the web application, API, and Grafana
  • A Kubernetes namespace in which Agent Barn can create agent workloads
  • Persistent storage for application databases, monitoring, and agents

Starting an Agent later creates a dedicated Deployment, Service, Secret, ConfigMap, PVC, and related runtime configuration inside the same namespace.

The agentbarn-api chart deploys Product API, Ingest API, Communications, the Domain Event worker, the Domain Event reconciliation CronJob, the Restore Point reconciliation CronJob, migration hook resources, and their supporting Services and Secrets. Communications uses the same API image tag as Product and Ingest but runs independently as a Deployment using api.communications_main:app on port 8002. It is neither a Product API sidecar nor the Domain Event worker.

BoundaryTraffic and responsibility
Public ingress/api/* → Product API :8000; /communications/v1/webhooks/* → Product API :8000; /agent-hooks/v1/* → Product API :8000
Cluster networkAgent Runtime → Ingest :8001; Agent Runtime → Communications :8002; Prometheus → Communications /metrics, when a scrape target is configured; Kubernetes → Communications /health
CommunicationsPostgreSQL-backed Deliveries and leases; provider ingress supervisor; outbound Delivery worker

Conversations do not flow through Ingest. Communications writes canonical Conversation Messages, while Ingest records Runtime Tool Call telemetry.

Before you begin

Kubernetes cluster

You need an existing Kubernetes cluster with:

  • A working kubectl context
  • A default StorageClass, or the name of a StorageClass you can select
  • Support for dynamically provisioned ReadWriteOnce volumes
  • A Traefik ingress controller with the traefik IngressClass
  • cert-manager
  • A cert-manager ClusterIssuer named letsencrypt-http01
  • DNS control for the web application, API, and Grafana hostnames
  • Enough capacity for the platform services and the agents you plan to run
WorkloadDefault storage
Agent Barn PostgreSQL5 GiB
LiteLLM PostgreSQL2 GiB
Firecrawl PostgreSQL2 GiB
Prometheus10 GiB
Each running Agent1 GiB

Plan additional capacity for database growth, monitoring retention, backups, and simultaneous agent workloads.

Command-line tools

Install these tools on the machine from which you will deploy:

Confirm that they are available:

Shell
kubectl version --client
helm version
helmfile --version
helm plugin list

If diff is not listed as a Helm plugin, install it:

Shell
helm plugin install https://github.com/databus23/helm-diff

External accounts and credentials

You will need:

  • An OpenRouter API key
  • Credentials for the container registry holding the Agent Barn images
  • A Slack incoming-webhook URL for deployment alerts
  • Three DNS hostnames pointing to your ingress endpoint

The following integrations are optional:

  • Cloudflare Email Sending for invitations, password resets, and notifications
  • A Google OAuth web client for Google Workspace authentication

Arrange Kubernetes access

Ask the administrator of your Kubernetes cluster to prepare access for the following tasks:

Task Access used
Prepare the namespace and Kubernetes permission bindingsBootstrap access that can manage the required namespace and RBAC resources.
Install or update Agent BarnThe kubeconfig supplied through KUBECONFIG on the machine running deployment tools.
Let the API create and manage Agent workloadsThe kubeconfig supplied through POD_KUBECONFIG_B64 and mounted into the API.
Run the LiteLLM key setup hookThe configured hook ServiceAccount, normally agent-farm-user.

For the API, request an identity whose permissions are limited to the Agent workload operations required in the installation's namespace. Have your cluster administrator supply a kubeconfig usable by the API container, together with its expiry and renewal instructions. An interactive login on your laptop is not a complete credential setup for a server process.

The kubeconfig mounted into the Agent Barn API must allow it to manage agent Deployments, Services, PVCs, Secrets, ConfigMaps, and Pods in the target namespace. It also needs access to pod logs, exec, and port-forward operations used by Agent health and log features.

Keep the deployment kubeconfig and API kubeconfig clearly named so that you can select the intended file for each setting.

Understand the shipped bootstrap step

The supplied deploy.sh always applies k8s/agent-farm-user.yaml before running Helmfile. This manifest creates or updates the agent-farm namespace and its hook/monitoring ServiceAccount permissions. It also sets Pod Security labels on that namespace to privileged.

Your cluster administrator should review these resources against the cluster's policies before you use this deployment path. A deployment identity restricted to ordinary application resources may not be allowed to apply them. Precreating the resources does not remove the script's apply step.

The manifest does not generate a kubeconfig and is not a complete permission policy for managing Agent workloads. Do not use its ServiceAccount as the API identity solely because its name appears in the deployment instructions.

The default namespace is agent-farm. Choosing another NAMESPACE also requires corresponding bootstrap resources and permissions; the existing manifest is not rewritten automatically.

Verify the cluster

Confirm that kubectl is connected to the intended cluster:

Shell
kubectl config current-context
kubectl cluster-info
kubectl get nodes

Check the required cluster services:

Shell
kubectl get ingressclass traefik
kubectl get clusterissuer letsencrypt-http01
kubectl get storageclass

Check whether your deployment identity can create and manage resources:

Shell
kubectl auth can-i create namespaces
kubectl auth can-i create deployments --namespace agent-farm
kubectl auth can-i create services --namespace agent-farm
kubectl auth can-i create secrets --namespace agent-farm
kubectl auth can-i create persistentvolumeclaims --namespace agent-farm
kubectl auth can-i create ingresses --namespace agent-farm

These examples cover selected permissions; they are not a complete permission checklist for the bootstrap manifest and Helm releases.

Before running deploy.sh, ask your cluster administrator to confirm that the deployment identity can apply the namespace and RBAC resources in k8s/agent-farm-user.yaml, as well as install the Helm releases. The script attempts this bootstrap apply on every run, even when the resources already exist. Creating the namespace in advance does not make a restricted deployment identity sufficient. If your identity cannot perform that step, arrange the deployment with your cluster administrator before continuing.

Get the deployment files

Download the release source

  1. Open Agent Barn releases and select the release you intend to install.
  2. Read its release notes, then download Source code (zip) or Source code (tar.gz) from the release assets.
  3. Extract the archive into a new directory and open the extracted repository directory in your terminal. Use the directory created by your extraction tool; its name is not necessarily agent-barn-deploy.
  4. Copy .env.deploy.spec to .env.deploy, then fill in the deployment settings described in Deployment configuration.

The source archive contains application source and deployment definitions. It does not contain built container images. Before deploying, obtain a compatible image set and any required registry credentials, or build and publish images through a separately documented process for your release.

As of September 5, 2026, v0.16.1 and v0.16.2 provide source downloads but have no separate deployment bundle attached. Older releases contain agent-farm-named bundles; do not use one as a substitute for the deployment files of a newer release.

If your release includes a deployment bundle

This section applies only when a separate deployment bundle has been provided as an additional asset alongside the source download. Skip it if your release has only source archives.

In a terminal with tar available, open the directory containing the download. Replace the example tag with the tag you selected:

Shell
RELEASE_TAG='vX.Y.Z'
tar -xzf "agent-barn-deploy-${RELEASE_TAG}.tar.gz"
cd agent-barn-deploy

A bundle extracts to an agent-barn-deploy directory containing deployment charts, Kubernetes prerequisites, helmfile.yaml.gotmpl, deploy.sh, and a generated .env.deploy. Edit that included .env.deploy; you do not need to create it from a spec file when using a bundle.

Keep the extracted files together and run subsequent bundle commands from this directory. The archive contains deployment files, not the container images themselves.

Obtain access to the container images

The deployment needs API, UI, Hermes, and OpenClaw images.

Container image access is separate from source downloads. You need the repository names, version tags, and any required credentials for the registry you will use. If these have not been provided for your installation, resolve image distribution before running the deployment.

Public access to Agent Barn's source code does not grant access to a container registry. Use credentials for the registry named in your deployment configuration; a GitHub token is not a substitute for those credentials.

Before continuing, you need:

Information Where it belongs
Registry host used for authenticationREGISTRY_SERVER
Registry prefix, including any shared repository pathREGISTRY_PREFIX
Registry username and password for image accessREGISTRY_USERNAME and REGISTRY_PASSWORD
Repository names for the four Agent Barn imagesAPI_IMAGE_REPOSITORY, UI_IMAGE_REPOSITORY, HERMES_IMAGE_REPOSITORY, OPENCLAW_IMAGE_REPOSITORY
Version tags for those repositoriesAPI_IMAGE_TAG, UI_IMAGE_TAG, HERMES_IMAGE_TAG, OPENCLAW_IMAGE_TAG

A generated deployment bundle, where one is provided, uses clients.registry.k8s.aai-labs.com as its registry host and clients.registry.k8s.aai-labs.com/agent-barn as its image prefix, with repository names api, ui, hermes-base, and openclaw-base. These values describe that distribution; they do not establish that your account has access.

Keep the supplied version tags unless you are deliberately selecting another compatible image set. API and UI use the product release tag. Hermes and OpenClaw have separate runtime version tags; do not give every component the product tag.

If you do not yet have image access, obtain it before running the deployment. Changing the hostname to a different Agent Barn registry does not make the same repositories or tags available there.

Deploy from a Git checkout

Instead of a release archive, you can use the deployment files directly from a checkout of the public repository:

Shell
git clone https://github.com/aai-labs/agent-barn.git
cd agent-barn
git checkout RELEASE_TAG
cp .env.deploy.spec .env.deploy

Use a released tag or a commit whose API, UI, Hermes, and OpenClaw image tags are available in your registry.

Configure the deployment

Open .env.deploy in a text editor and replace every required blank value.

The file is sourced as a shell environment file. Use plain KEY=value entries and avoid spaces around the equals sign.

Cluster and namespace

VariableRequirementDescription
KUBECONFIGRequiredAbsolute path to the kubeconfig used by kubectl, Helm, and Helmfile
NAMESPACERequiredTarget namespace; use agent-farm for the production deployment
POD_KUBECONFIG_B64OptionalSingle-line base64 kubeconfig used by the API to manage Agent workloads; when omitted, deploy.sh encodes and uses KUBECONFIG
STORAGE_CLASSOptionalStorageClass for databases, Prometheus, and Agent PVCs; leave empty to use the cluster default

Set the deployment and API kubeconfigs

  1. Obtain the deployment kubeconfig and the API workload kubeconfig from your cluster administrator. Confirm which installation and namespace each is intended for.
  2. In .env.deploy, set KUBECONFIG to the deployment kubeconfig's path on the machine where you will run deployment tools.
  3. Encode the administrator-supplied API kubeconfig as a single line. In the following command, replace /path/to/agentbarn-api.kubeconfig with that file's actual path:
Encode the API kubeconfig
base64 < /path/to/agentbarn-api.kubeconfig | tr -d '\n'
  1. Copy the complete output into POD_KUBECONFIG_B64 in .env.deploy. This value contains Kubernetes credentials; keep it with your deployment secrets and do not paste it into support messages.
  2. Keep the credential's expiry and renewal instructions with the installation records. When it changes, update POD_KUBECONFIG_B64 and roll the API through your deployment process so it loads the new credential.

If POD_KUBECONFIG_B64 is missing or empty, the current deploy.sh encodes and uses the deployment kubeconfig for the API. It does not create a separate restricted identity. Base64 encoding changes the representation of a credential, not its permissions.

Use an absolute kubeconfig path:

Environment
KUBECONFIG=/home/your-user/.kube/agent-barn-production.yaml
NAMESPACE=agent-farm
STORAGE_CLASS=REPLACE_WITH_STORAGE_CLASS

Container registry

VariableRequirementDescription
REGISTRY_PREFIXRequiredRegistry path prepended to all Agent Barn image repositories
REGISTRY_SERVERRequiredRegistry hostname used for authentication
REGISTRY_USERNAMERequiredRegistry username
REGISTRY_PASSWORDRequiredRegistry password or access token
API_IMAGE_REPOSITORYRequiredAPI image repository
UI_IMAGE_REPOSITORYRequiredUI image repository
HERMES_IMAGE_REPOSITORYRequiredHermes runtime image repository
OPENCLAW_IMAGE_REPOSITORYRequiredOpenClaw runtime image repository
API_IMAGE_TAGRequiredReleased API image tag
UI_IMAGE_TAGRequiredReleased UI image tag
HERMES_IMAGE_TAGRequiredCompatible Hermes image tag
OPENCLAW_IMAGE_TAGRequiredCompatible OpenClaw image tag

If you downloaded a release bundle, the registry paths and image tags are already populated. Add the registry username and password for your image access without changing the pinned versions. See Obtain access to the container images for what you need and where to get it.

A GitHub token is not required to read the public agent-barn or aai-cli repositories. Kubernetes only needs credentials for the registry from which it pulls the deployment images.

Database configuration

Agent Barn deploys three independent PostgreSQL instances:

VariableRequirementDefault user or database
POSTGRES_APP_USERRequiredagentfarm
POSTGRES_APP_PASSWORDRequiredGenerate a strong password
POSTGRES_APP_DBRequiredagentfarm
POSTGRES_LITELLM_USERRequiredlitellm
POSTGRES_LITELLM_PASSWORDRequiredGenerate a different strong password
POSTGRES_LITELLM_DBRequiredlitellm
POSTGRES_FIRECRAWL_USERRequiredfirecrawl
POSTGRES_FIRECRAWL_PASSWORDRequiredGenerate a different strong password
POSTGRES_FIRECRAWL_DBRequiredfirecrawl

Generate independent passwords:

Shell
openssl rand -hex 24
openssl rand -hex 24
openssl rand -hex 24

Store these values in your secret manager before deploying.

LiteLLM and OpenRouter

VariableRequirementDescription
LITELLM_MASTER_KEYRequiredStable LiteLLM administrative key beginning with sk-
OPENROUTER_API_KEYRequiredAPI key issued by OpenRouter
AGENT_DEFAULT_MODELOptionalDefault in litellm/openrouter/<model> format
AGENT_MODEL_ALLOWLISTOptionalComma-separated model patterns

Generate a LiteLLM master key:

Shell
echo "sk-$(openssl rand -hex 24)"

Keep this key stable. LiteLLM uses it when managing the virtual keys assigned to Agents.

Agent Barn application secrets

VariableRequirementDescription
SECRET_SIGNING_KEYRequiredSigns Agent Barn authentication tokens
AGENT_TOKEN_ENCRYPTION_KEYRequiredFernet-compatible key used to encrypt stored credentials
PLATFORM_ADMIN_CREDENTIALSRequiredInitial platform administrator in email:password format
ENVIRONMENTRequiredDeployment name stamped into the application and alerts
FIRECRAWL_API_KEYRequiredShared key protecting the internal Firecrawl service

Generate the signing key:

Shell
openssl rand -hex 32

Generate the Fernet-compatible encryption key:

Shell
openssl rand -base64 32 | tr '+/' '-_' | tr -d '\n'

Generate the Firecrawl key:

Shell
openssl rand -hex 24

Configure the administrator and environment:

Environment
PLATFORM_ADMIN_CREDENTIALS=admin@example.com:REPLACE_WITH_STRONG_PASSWORD
ENVIRONMENT=production

The administrator password must contain at least eight characters, including an uppercase letter, a lowercase letter, and a digit.

Communications values and shared configuration

Communications values
communications:
  enabled: true
  replicaCount: 1
  service:
    port: 8002
  resources:
    requests:
      memory: 256Mi
      cpu: 100m
    limits:
      memory: 512Mi
      cpu: 500m
  • communications.enabled controls rendering of the Communications Deployment and Service.
  • communications.replicaCount controls process replicas.
  • communications.service.port controls the ClusterIP Service port; the container target port remains 8002.
  • Communications resource requests and limits are independent of the Product API container. It uses the chart’s API_IMAGE_TAG; there is no separate Communications image repository or tag.

The chart renders <release-name>-communications as both Deployment and ClusterIP Service. Its pod runs uvicorn api.communications_main:app --host 0.0.0.0 --port 8002 with app.kubernetes.io/component: communications. It uses the API release ServiceAccount and image pull Secrets, loads common application settings from the API Secret, connects to the shared PostgreSQL database, uses the shared Agent token encryption key, and starts the ingress supervisor and outbound Delivery worker during its lifespan.

The shared Secret includes DB_CONNECTION_URL, AGENT_TOKEN_ENCRYPTION_KEY, environment identity, common application settings, and provider-specific global settings where configured. Product API and Communications must use the same database and a key consistent across both workloads and all replicas. Connection credentials remain encrypted in PostgreSQL; provider credentials do not belong in Deployment manifests, release environment files, or Agent Runtime pods.

Public hostnames

VariableRequirementExample
UI_HOSTRequiredagentbarn.example.com
API_HOSTRequiredapi.agentbarn.example.com
WEB_APP_URLRequiredhttps://agentbarn.example.com
GRAFANA_HOSTRequiredgrafana.agentbarn.example.com

Configure hostnames without URL schemes in the *_HOST values:

Environment
UI_HOST=agentbarn.example.com
API_HOST=api.agentbarn.example.com
WEB_APP_URL=https://agentbarn.example.com
GRAFANA_HOST=grafana.agentbarn.example.com

Monitoring

The current Helmfile requires:

VariableRequirementDescription
SLACK_ALERTS_WEBHOOK_URLRequiredSlack incoming webhook used by Alertmanager
GRAFANA_ADMIN_PASSWORDRequiredInitial Grafana administrator password
GRAFANA_HOSTRequiredPublic Grafana hostname

Add these entries to .env.deploy if they are not already present:

Environment
SLACK_ALERTS_WEBHOOK_URL=REPLACE_WITH_SLACK_WEBHOOK
GRAFANA_ADMIN_PASSWORD=REPLACE_WITH_STRONG_PASSWORD
GRAFANA_HOST=grafana.agentbarn.example.com

Optional email delivery

Leave all three values empty to disable transactional email:

Environment
CLOUDFLARE_ACCOUNT_ID=
CLOUDFLARE_API_TOKEN=
SENDER_EMAIL=

To enable it:

  • Create or select a Cloudflare Email Sending account.
  • Verify the sending domain.
  • Give the API token the Email Sending: Edit permission.
  • Set SENDER_EMAIL to an address on the verified domain.

Use an environment-specific mail. subdomain, such as:

Environment
SENDER_EMAIL=noreply@mail.agentbarn.example.com

All three values must be configured for delivery to be enabled.

Optional Google Workspace authentication

Leave these values empty to disable Google OAuth:

Environment
GOOGLE_CLOUD_CLIENT_ID=
GOOGLE_CLOUD_CLIENT_SECRET=

To enable it, create a Google OAuth 2.0 Web application client and register this redirect URI:

Text
https://agentbarn.example.com/api/v1/integrations/google/callback

Then set the client ID and client secret in .env.deploy.

Configure DNS and TLS

Point the following DNS records to the public address of your Traefik ingress controller:

  1. UI_HOST
  2. API_HOST
  3. GRAFANA_HOST

Find the ingress address using the command appropriate for your cluster. For example:

Shell
kubectl get services --all-namespaces
kubectl get ingressclass traefik

Confirm that DNS resolves before deploying:

Shell
dig +short agentbarn.example.com
dig +short api.agentbarn.example.com
dig +short grafana.agentbarn.example.com

Confirm that the expected ClusterIssuer is ready:

Shell
kubectl get clusterissuer letsencrypt-http01
kubectl describe clusterissuer letsencrypt-http01

The Product API remains exposed at /api. The only public webhook prefix is /communications/v1/webhooks, routed to <release-name>-api:8000. Keep Runtime Delivery claim, reply, and completion routes, Platform Driver events, /health, and /metrics internal.

Webhook-based Platform Connections receive URLs shaped like https://<api-host>/communications/v1/webhooks/<connection-id>. Public DNS, valid TLS, provider-to-ingress network access, ingress routing, and the correct API_EXTERNAL_URL are required. Microsoft Teams currently uses authenticated webhook ingress; the Connection ID scopes the endpoint and the Platform Plugin authenticates the provider request. A public endpoint must never bypass provider authentication.

Deploy Agent Barn

Build the pinned monitoring chart dependencies:

Shell
helm dependency build helm/monitoring

Review the active Kubernetes context one final time:

Shell
kubectl config current-context

Deploy the stack:

Shell
./deploy.sh

The script:

  1. Loads values from .env.deploy.
  2. Verifies that helmfile and kubectl are available.
  3. Derives the API pod kubeconfig when POD_KUBECONFIG_B64 is not supplied.
  4. Applies the agent-farm namespace and bootstrap RBAC manifest.
  5. Runs helmfile sync --wait.
  6. Applies database migrations through an API chart hook.
  7. Generates the LiteLLM API key used by Agent Barn.
  8. Waits for the Helm releases to become ready.

Helmfile deploys agentbarn-api after its supporting application PostgreSQL, LiteLLM, Firecrawl, and Redis releases. Communications is rendered by agentbarn-api, not as a separate Helmfile release. PostgreSQL must be ready before it processes durable data; Redis is not the Communication Delivery queue. Product, Ingest, and Communications use one API image version supplied by API_IMAGE_TAG.

The releases are installed in dependency order:

ReleasePurpose
postgres-appAgent Barn application database
postgres-litellmLiteLLM database
postgres-firecrawlFirecrawl database
redisBackground task transport
litellmModel proxy and per-Agent virtual keys
firecrawlSelf-hosted web retrieval
agentbarn-apiProduct API, Ingest API, Communications, Domain Event worker, migrations, and reconciliation
agentbarn-uiAgent Barn web application
monitoringPrometheus, Grafana, Alertmanager, and dashboards

Communication Connection, Delivery, journal, and Conversation schema changes use the normal API Alembic migrations. The migration hook runs before install or upgrade; wait for it to complete before evaluating Communications health. Do not run a newer Communications process against an older database schema, and do not add a separate Communications migration command.

Verify the deployment

  • Helm releases are installed.
  • Workloads become ready.
  • Persistent volume claims are bound.
  • Ingress and certificate resources are ready.
  • The public API and web application respond.

Check the Helm releases:

Shell
helm list --namespace agent-farm

Check the workloads:

Shell
kubectl get deployments,statefulsets,pods --namespace agent-farm

Check persistent volumes:

Shell
kubectl get pvc --namespace agent-farm

Check ingress and certificate resources:

Shell
kubectl get ingress --namespace agent-farm
kubectl get certificate --namespace agent-farm
kubectl get certificaterequest --namespace agent-farm

Expected: All long-running pods eventually report Running, and their ready-container counts are complete.

Verify the public API:

Shell
curl --fail https://api.agentbarn.example.com/api/v1/health

A healthy response resembles:

JSON
{
  "status": "ok",
  "db": "connected"
}

Verify the web application:

Shell
curl --head https://agentbarn.example.com

Open these URLs in a browser:

Text
https://agentbarn.example.com
https://grafana.agentbarn.example.com

Verify Communications

Use the namespace and release name for the environment you deployed:

Deployment
kubectl -n <namespace> get deployment <release-name>-communications
Rollout
kubectl -n <namespace> rollout status deployment/<release-name>-communications
Service
kubectl -n <namespace> get service <release-name>-communications
Pods
kubectl -n <namespace> get pods \
  -l app.kubernetes.io/component=communications
Logs
kubectl -n <namespace> logs \
  deployment/<release-name>-communications

Port-forward the internal Service to inspect process availability and the Prometheus surface:

Port-forward
kubectl -n <namespace> port-forward \
  service/<release-name>-communications 8002:8002
Internal checks
curl http://127.0.0.1:8002/health
curl http://127.0.0.1:8002/metrics

/health checks process availability and /metrics confirms that the metrics endpoint is exposed, not that Prometheus is collecting from it. Neither proves PostgreSQL Delivery progress, Runtime claim health, provider connectivity, individual Connection health, or successful outbound provider delivery. Use Communication diagnostics and metrics for those checks; do not expose either endpoint publicly.

Agent Runtime pods must resolve <release-name>-communications and reach http://<release-name>-communications:8002/communications/v1, sending the Runtime Communications bearer credential and protocol-version header. DNS or connection failures indicate cluster networking or Service configuration; 401 Unauthorized indicates a credential problem, 426 Upgrade Required a protocol-version mismatch, and 204 No Content from a valid claim request means no Delivery is pending. Do not put a Runtime bearer credential in this guide.

Confirm public DNS and valid TLS route /communications/v1/webhooks/* and /api/* to Product API. Confirm /communications/v1/agents/* and /metrics are not public. An unauthenticated webhook request should be rejected, not treated as a functional message test; validate the complete webhook through the intended provider’s setup flow.

If a pod is not ready, inspect it before retrying the deployment:

Shell
kubectl describe pod POD_NAME --namespace agent-farm
kubectl logs POD_NAME --namespace agent-farm

Complete the first-time setup

Sign in at WEB_APP_URL using the address and password from PLATFORM_ADMIN_CREDENTIALS.

A fresh Agent Barn database contains the platform administrator but no Organization. Complete the initial setup in this order:

  1. Sign in as the platform administrator.
  2. Create an Organization.
  3. Add or invite Organization members.
  4. Configure shared credentials and integrations.
  5. Hire an Agent from a predefined template.
  6. Configure its Communication Connections and model.
  7. Start the Agent.

After starting the first Agent, verify that Kubernetes created its resources:

Shell
kubectl get deployments,pods,services,pvc \
  --namespace agent-farm \
  --selector agentbarn.io/component=agent

Each Agent receives a 1 GiB PVC by default. Its generated resources remain isolated by Agent identity within the agent-farm namespace.

Complete a Communications check

  1. Confirm Product, Ingest, and Communications workloads are healthy and migrations completed.
  2. Create or open a stopped headless Agent, then add one Communication Connection.
  3. Confirm the Connection reports provider health independently from Agent lifecycle, then start the Agent.
  4. Send a provider message that passes Connection policy and confirm an inbound Conversation Message appears.
  5. Confirm the Runtime claims and completes the Delivery, then confirm the outbound reply reaches the same Connection.
  6. Review Connection diagnostics for the complete pipeline and confirm Tool Calls still appear through Ingest separately.
Deployment complete

Agent Barn and its supporting services are running, and the cluster is ready to host Agent workloads.

Next guide Configure a production deployment

Operate Communications replicas and monitoring

Communications readiness uses GET /health on port 8002 after a 5-second initial delay, checking every 5 seconds. Liveness uses the same endpoint after 15 seconds, checking every 15 seconds. A successful response is equivalent to {"status":"ok"}; it is process health only, not a Delivery, Runtime, provider, or Connection-health guarantee.

Multiple Communications replicas coordinate supervised provider ingress through PostgreSQL leases. One replica owns a supervised Connection at a time; lease expiry permits takeover after failure, while webhook traffic can load-balance across replicas. Durable Communication Deliveries remain in PostgreSQL, outbound claims preserve per-Conversation ordering, and Connection revision changes trigger provider-session reconciliation. Ownership is not process-local, and every replica must share the same database and encryption configuration.

Communications monitoring

The Communications service exposes /metrics internally on port 8002, but the default monitoring chart does not configure Prometheus to scrape it.

Deploying the service therefore makes its metrics endpoint available without automatically adding Communications metrics to Prometheus or Grafana. Configure a scrape target in your monitoring installation if you want to collect these metrics.

For the standard API Helm release, the internal endpoint is:

Text
http://agentbarn-api-communications:8002/metrics

If your API Helm release has a different name, use <your-api-release-name>-communications as the Service hostname.

Keep this endpoint internal. See Monitor the platform for configuration details and Communication Diagnostics for inspecting individual connections and deliveries.

The Communications endpoint exposes the following metric families:

  • agentbarn_communication_connection_status
  • agentbarn_communication_delivery_outcomes
  • agentbarn_communication_queue_depth
  • agentbarn_communication_oldest_queued_age_seconds
  • agentbarn_communication_delivery_latency_seconds
  • agentbarn_communication_reconnects
  • agentbarn_communication_policy_dispositions

These metrics use low-cardinality labels and must not include Organization, Agent, Connection, Conversation, or User identifiers.

For service details, see Self-hosting Communications; use self-hosting configuration, database migrations, and monitoring for their respective operating boundaries.

Upgrade the deployment

Before upgrading:

  • Back up all three PostgreSQL databases.
  • Back up the stable application and LiteLLM keys.
  • Review the release notes.
  • Confirm that all four Agent Barn image tags belong to the target release.
  • Review chart or configuration changes.
  • Plan for a brief LiteLLM interruption during replacement.

Update the pinned image tags in .env.deploy, then run:

Shell
helm dependency build helm/monitoring
./deploy.sh

The API chart runs database migrations before installation or upgrade.

LiteLLM uses a non-overlapping update strategy because two 2 GiB LiteLLM pods may not fit inside the namespace quota simultaneously. Its replacement can briefly interrupt model requests.

Evaluate rollback compatibility across the Product API image, Communications image, database schema, Runtime Communications protocol, and Platform Plugin behavior. Product and Communications normally use the same API image tag, migrations may not be automatically reversible, provider sessions reconcile as the Communications Deployment rolls, and durable Deliveries survive pod replacement in PostgreSQL. Include pending and dead-lettered Deliveries in the rollback decision; never delete Communication tables or Connection data as a rollback step.

Troubleshooting

Helm reports missing chart dependencies

helm dependency build helm/monitoring

Build the monitoring dependencies and retry:

Shell
helm dependency build helm/monitoring
./deploy.sh

The namespace or RBAC bootstrap fails

kubectl auth can-i create namespaces

Check your active identity:

Shell
kubectl auth whoami
kubectl auth can-i create namespaces
kubectl auth can-i create roles --namespace agent-farm
kubectl auth can-i create rolebindings --namespace agent-farm

The default deploy.sh applies k8s/agent-farm-user.yaml, including the Namespace, Role, and RoleBinding.

If you only have namespace-scoped access, the namespace and bootstrap identity must be provisioned out of band, and the bootstrap step in deploy.sh must be adapted accordingly.

A pod is stuck in ImagePullBackOff

kubectl describe pod POD_NAME --namespace agent-farm

Inspect the pod:

Shell
kubectl describe pod POD_NAME --namespace agent-farm

Confirm:

  • The registry hostname is correct.
  • The registry username and password are valid.
  • The image repository and tag exist.
  • The generated registry pull Secret contains credentials for REGISTRY_SERVER.

Do not replace a missing release image with an unrelated latest image.

A PVC remains Pending

kubectl describe pvc PVC_NAME --namespace agent-farm

Check the claim and available StorageClasses:

Shell
kubectl describe pvc PVC_NAME --namespace agent-farm
kubectl get storageclass

Confirm that STORAGE_CLASS exists and supports dynamically provisioned ReadWriteOnce volumes.

TLS certificates are not ready

kubectl get certificate,challenge,order --namespace agent-farm

Inspect the ingress, certificates, and cert-manager challenges:

Shell
kubectl describe ingress --namespace agent-farm
kubectl get certificate,certificaterequest,challenge,order --namespace agent-farm

Confirm that:

  • The three hostnames resolve to the ingress endpoint.
  • Traefik accepts the traefik IngressClass.
  • The letsencrypt-http01 ClusterIssuer exists and is ready.
  • Ports 80 and 443 are reachable where required by the issuer.

The migration hook fails

kubectl get jobs --namespace agent-farm

List jobs and inspect the failed migration pod:

Shell
kubectl get jobs --namespace agent-farm
kubectl get pods --namespace agent-farm
kubectl logs JOB_POD_NAME --namespace agent-farm

Confirm that the application PostgreSQL pod is ready and that the configured database password still matches the initialized database.

The LiteLLM key hook fails

kubectl logs JOB_POD_NAME --namespace agent-farm

Inspect the hook job:

Shell
kubectl get jobs --namespace agent-farm
kubectl logs JOB_POD_NAME --namespace agent-farm

Confirm that:

  • LiteLLM is ready.
  • LITELLM_MASTER_KEY is correct.
  • The configured hook ServiceAccount can create and update Secrets.
  • The ServiceAccount exists in agent-farm.

Agent Barn loads, but an Agent cannot start

kubectl logs deployment/agentbarn-api --namespace agent-farm

Inspect the API logs:

Shell
kubectl logs deployment/agentbarn-api \
  --namespace agent-farm \
  --container api

Confirm that the kubeconfig mounted into the API:

  • Targets the correct cluster
  • Can manage resources in agent-farm
  • Uses a reachable in-cluster Kubernetes API endpoint
  • Has not expired
  • Can create Deployments, Services, PVCs, Secrets, and ConfigMaps
  • Can read Pods and pod logs

Also verify that the Hermes and OpenClaw image tags exist in the configured registry.

Grafana starts but shows no Agents

agentbarn.io/component=agent

Agent metrics appear after an Agent has been started with the current resource labels.

For an Agent that predates the monitoring deployment, stop and start it once so Agent Barn recreates its runtime resources and monitoring metadata.

Deployment constraints

Account for these current deployment-tooling constraints when preparing an environment:

  1. .env.deploy.spec defines INGRESS_CLUSTER_ISSUER, but helmfile.yaml.gotmpl does not currently pass it into the API or UI charts. The charts therefore use letsencrypt-http01.
  2. helmfile.yaml.gotmpl requires SLACK_ALERTS_WEBHOOK_URL, GRAFANA_ADMIN_PASSWORD, MONITORING_WEB_PASSWORD, and GRAFANA_HOST, but those variables are not currently listed in .env.deploy.spec. Add them manually.
  3. deploy.sh does not currently run helm dependency build helm/monitoring. Run it before the deployment.
  4. deploy.sh always applies the production bootstrap manifest at k8s/agent-farm-user.yaml. Changing only NAMESPACE is not sufficient for a staging or custom-namespace deployment.
  5. The API-facing kubeconfig defaults to the same kubeconfig used for the deployment. Production operators should replace this with a dedicated namespace-scoped identity.
  6. The repository and GitHub project are named agent-barn, while the production and staging namespaces intentionally remain agent-farm and agent-farm-staging.

Agent Email

Agent Email needs Cloudflare Email Routing, an inbound Worker, sending access, and matching environment-specific inbound secrets. Transactional email alone does not enable it. Follow Configure Agent Email for setup and routing-rule ownership.

Documentation