Self-hosting
How-to

Upgrade Agent Barn

Plan compatible application and runtime rollouts, preserve session history, and distinguish client-image publication from hosted deployment and recovery.

For
Platform engineers, self-hosted operators, and release managers
On this page
  1. Overview
  2. Understand the upgrade boundaries
  3. Choose an upgrade path
  4. Before you begin
  5. 1. Record the current state
  6. 2. Review the release
  7. 3. Back up state
  8. 4. Test in staging
  9. 5. Preview the upgrade
  10. 6. Run the upgrade
  11. 7. Watch upgrade hooks
  12. Verify the Communications upgrade surface
  13. 8. Verify the upgrade
  14. 9. Upgrade running Agents
  15. Rollback and recovery
  16. Recover the correct component
  17. Troubleshooting
  18. Next steps
  19. Roll out initiated delivery
  20. Hermes session continuity
  21. Client image release

Upgrade outcome

A staged release with a deliberate recovery path

Complete this guide to move the platform and selected Agents forward without losing track of data, keys, images, or partial deployment state.

  • Staging proves the complete release and every affected runtime/platform pair.
  • Backups, stable secrets, image references, and Alembic state are recorded.
  • Hooks and rollouts are watched before product and provider verification.
  • Agent runtime adoption proceeds through canaries and controlled batches.

Overview

An Agent Barn upgrade can change the API, UI, worker, database schema, supporting services, monitoring, Agent builders, and runtime images.

ReviewBack upStageVerifyApproveDeployVerifyRestart selected Agents

GitHub branch deployments and versioned release bundles use the same Helmfile release graph. Neither path automatically rebuilds existing Agent workloads, so platform deployment and runtime rollout remain separate operations.

Understand the upgrade boundaries

Upgrade areaWhat can changePrimary risk
APIRoutes, services, workers, Agent builders, integrationsContract, migration, or background-processing failure
CommunicationsConnection supervision, ingress leases, Deliveries, journal, and provider transportsA healthy Product API does not prove provider delivery or Connection recovery
IngestRuntime Tool Call telemetry and Ingest API availabilityTelemetry can fail independently of Product and Communications
UIPages, schemas, queries, onboarding flowsUI and API incompatibility
Application databaseAlembic schema and data migrationsIrreversible data or compatibility change
Hermes runtimeBase image and runtime behaviorExisting Hermes Agents retain the old image until restarted
OpenClaw runtimeBase image and runtime behaviorExisting OpenClaw Agents retain the old image until restarted
LiteLLMProxy image, configuration, master key, virtual keysBrief interruption or invalid Agent and API keys
FirecrawlAPI, browser service, RabbitMQ, database integrationIntegration or upstream-image incompatibility
PostgreSQL and RedisStateful service images and configurationData compatibility and availability
Helm chartsDeployments, Services, Secrets, Jobs, ingress, storagePartial or incompatible Kubernetes rollout
MonitoringRules, dashboards, Grafana, AlertmanagerLost visibility or alert delivery
Environment configurationCredentials, URLs, models, storage, OAuth, emailCross-environment or secret mismatch

Changes that are not routine upgrades

  • Changing PostgreSQL passwords on initialized volumes.
  • Rotating SECRET_SIGNING_KEY, AGENT_TOKEN_ENCRYPTION_KEY, or LITELLM_MASTER_KEY.
  • Changing a StatefulSet’s existing StorageClass.
  • Renaming agent-farm or agent-farm-staging.
  • Moving persistent data to another cluster or storage provider.

Choose an upgrade path

Source-operated environments

GitHub branch deployment

Use the agent-barn repository when environments deploy directly from source.

BranchEnvironmentNamespaceAPI/UI tagsRuntime tags
stagingStaging on k3sagent-farm-staginglatest-staging<version>-staging
mainAAI Labs testing ground on k3sagent-farmlatest<version>
vX.Y.Z tagHosted public production on TalosDedicated Talos clustervX.Y.ZSelected runtime VERSION

Pushes and manual dispatch run only from staging or main; any other branch is rejected. Per-branch concurrency serializes runs and does not cancel an in-progress deployment. These k3s paths are isolated staging and the AAI Labs testing ground; main is not the hosted public-production source.

Change detection compares with the latest successful deployment on that branch. A failure does not advance the baseline. Manual dispatch, a missing baseline, or an unavailable or non-ancestor baseline rebuilds API, UI, Hermes, and OpenClaw.

Packaged self-hosting

Release-bundle deployment

This path applies only when a separate deployment bundle has been provided for your release. A source archive uses .env.deploy.spec instead.

A bundle supplies helm/, k8s/, helmfile.yaml.gotmpl, deploy.sh, and a generated .env.deploy.

ComponentVersion source
APIAPI_IMAGE_TAG
UIUI_IMAGE_TAG
Hermes runtimeHermes VERSION file
OpenClaw runtimeOpenClaw VERSION file
Helm chartChart version for chart packaging

Transfer environment-specific values into the new bundle’s file. Do not replace it wholesale with an older copy because requirements can change between releases.

Hosted public production on Talos

The dedicated Talos cluster deploys from a release tag matching vX.Y.Z through .github/workflows/deploy-public.yml. The workflow publishes API and UI images to registry.agentbarn.dev, pins both image inputs to that exact tag, and uses PUBLIC_-prefixed cluster configuration and secrets.

Hosted public release identity
Cluster:         dedicated Talos cluster
Workflow:        .github/workflows/deploy-public.yml
Release tag:     vX.Y.Z
API/UI tags:     vX.Y.Z (pinned)
Registry:        registry.agentbarn.dev
Configuration:   PUBLIC_ prefixed values and secrets

It does not update the k3s registry’s moving latest tags. A manual public deployment can target an existing tag and use skip_build when the tag-pinned images already exist.

Before you begin

Assign a release owner, database and backup owner, verification owner, incident decision-maker, Agent restart owner, and communication owner.

  • Confirm staging, production monitoring, DNS, certificates, storage, registries, and providers are healthy.
  • Confirm there is no active incident or overlapping automated or manual deployment.
  • Verify operators can reach Kubernetes, Helm history, application logs, and backups.
  • Confirm every required image exists before approval.

Prepare the target release

  1. Select the target release from Agent Barn releases and read its release notes.
  2. Download its source ZIP or tar.gz and extract it into a separate working directory. If a matching deployment bundle has been provided instead, extract that separately and use its included configuration.
  3. For a source archive, copy .env.deploy.spec to .env.deploy. Carry forward your installation's infrastructure settings and credentials, including stable signing and encryption keys. Account for settings added or changed in the target release.
  4. Obtain the image locations and component versions intended for the target release. Preserve separate API/UI and runtime version selections; source download availability does not establish image availability.
  5. Continue with the backup and migration preparation below before deploying.

API and UI use the selected product release tag. Runtime image tags remain separate; do not replace every component tag with the product tag or with latest.

Record the current state

Capture enough evidence to distinguish an upgrade regression from pre-existing state and to reproduce the current deployment.

Helm releases and history
helm list --namespace agent-farm

helm history agentbarn-api --namespace agent-farm
helm history agentbarn-ui --namespace agent-farm
helm history litellm --namespace agent-farm
helm history monitoring --namespace agent-farm
Workload image references
kubectl get deployments,statefulsets \
  --namespace agent-farm \
  -o custom-columns='KIND:.kind,NAME:.metadata.name,IMAGES:.spec.template.spec.containers[*].image'

kubectl get cronjobs \
  --namespace agent-farm \
  -o custom-columns='NAME:.metadata.name,IMAGES:.spec.jobTemplate.spec.template.spec.containers[*].image'
Current Alembic revision
kubectl exec \
  --namespace agent-farm \
  deployment/agentbarn-api \
  -- sh -c 'cd /app/api && alembic current'
Workloads and persistent claims
kubectl get pods,deployments,statefulsets,jobs,cronjobs,persistentvolumeclaims \
  --namespace agent-farm

Record the source commit or bundle version, image references, Helm revisions, Alembic revision, backup identifiers, Agent count by runtime and platform, restart candidates, and known warnings.

Review the release

Application contracts

  • API and UI schema changes
  • Authorization and tenancy
  • Alembic migrations
  • Workers and Event Deliveries
  • Agent builders, Templates, Skills, and credentials
  • Provider and monitoring changes

Deployment contracts

  • Charts, dependencies, images, and strategies
  • Resources, Services, ingress, and certificates
  • Secrets, ConfigMaps, Jobs, hooks, and CronJobs
  • PVC size and StorageClass
  • Helmfile dependencies and Kubernetes permissions

Understand version identifiers

IdentifierMeaning
Git commit or PRSource identifier for the k3s branch deployment
API_IMAGE_TAGExplicit API deployment input; latest-staging or latest on k3s, vX.Y.Z on hosted public production
UI_IMAGE_TAGExplicit UI deployment input; the hosted public release uses the same vX.Y.Z tag as the API
Hermes/OpenClaw VERSIONIndependent Runtime image version from each Runtime VERSION file
Helm chart versionPackaging version that changes with chart templates or values
Git release tagImmutable public-release identifier; not a moving k3s image tag
Alembic revisionApplication database schema state

API and UI images are explicit API_IMAGE_TAG and UI_IMAGE_TAG inputs, not values derived from chart appVersion. A chart version changes when its templates or values change; branch application updates and documentation-only changes do not require an appVersion or service-image-tag change. There is no shared API/UI application number outside selected deployment inputs; hosted public API and UI releases deliberately share the same immutable vX.Y.Z tag.

Back up state

Back up Agent Barn, LiteLLM, and Firecrawl PostgreSQL databases; Agent PVCs when workspace continuity matters; environment configuration; stable keys; database credentials; and external provider configuration needed for recovery.

Keep with the recovery set
SECRET_SIGNING_KEY
AGENT_TOKEN_ENCRYPTION_KEY
LITELLM_MASTER_KEY
POSTGRES_APP_PASSWORD
POSTGRES_LITELLM_PASSWORD
POSTGRES_FIRECRAWL_PASSWORD

A database backup without its matching encryption key can leave Agent credentials unreadable. A LiteLLM database backup without its master key can leave virtual keys unusable.

  • Confirm the backup is complete and encrypted.
  • Store it outside the affected namespace.
  • Record its identifier in the release record.
  • Verify restoration through a tested procedure.

Test in staging

Deploy the complete release to agent-farm-staging and wait for the staging workflow to finish.

Staging release identity
Namespace:      agent-farm-staging
Environment:    staging
API/UI tag:     latest-staging
Runtime suffix: -staging

Platform

  • API, UI, login, and Organization isolation
  • Invitations, email, and Google OAuth
  • Database migration and restart recovery
  • Workers and Event Deliveries

Agents

  • Create, start, stop, and restart
  • Hermes on every used platform
  • OpenClaw on every used platform
  • Messages, Tool Calls, costs, and credentials

Providers

  • LiteLLM and OpenRouter requests
  • Firecrawl and Shared Credentials
  • Grafana and alert delivery
  • Canary for every affected runtime/platform pair

Account for shared providers

Namespace separation does not guarantee provider isolation. Registry, OpenRouter, Google OAuth, Cloudflare, and Slack alert delivery may be shared. Avoid tests that exhaust quotas, revoke shared credentials, or send unwanted production-facing messages.

Preview the upgrade

For a manual deployment, load the protected environment and inspect the Helmfile diff.

Secret-suppressed Helmfile preview
set -a
source .env.deploy
set +a

export POD_KUBECONFIG_B64="$(base64 "$KUBECONFIG" | tr -d '\n')"

helmfile -f helmfile.yaml.gotmpl diff --suppress-secrets

unset POD_KUBECONFIG_B64
  • Review StatefulSet replacement, PVC, StorageClass, database Secret, and stable-key changes.
  • Review images, resources, ingress, ServiceAccounts, hooks, deletions, and namespace scope.

The GitHub path proceeds directly to helmfile sync --wait. Pull-request review, source diff, and staging provide its preview boundary.

Run the upgrade

Hosted public release

  1. Select a commit already present on main.
  2. Review migrations, required configuration, chart changes, API/UI changes, Runtime versions, and Platform Plugin changes.
  3. Confirm public-only configuration and secrets exist.
  4. Create a unique vX.Y.Z tag for that commit.
  5. Push the tag to start the public deployment workflow.
  6. Allow the workflow to build and publish tag-pinned API and UI images, or use manual skip_build only when the tagged images already exist.
  7. Allow Helmfile to apply releases in dependency order.
  8. Confirm the migration hook succeeds before relying on the new application processes.
  9. Check Product API, Ingest, Communications, workers, UI, and monitoring.
  10. Run change-specific Runtime, Connection, and Delivery canaries.
  11. Record the deployed tag and any migration or compatibility constraints.
Create and push a public release tag
git tag vX.Y.Z
git push origin vX.Y.Z

Release tags are immutable identifiers: make corrections with a new version tag, never by moving one already deployed.

Release-bundle deployment

  1. Fill the new .env.deploy with existing environment values.
  2. Preserve pinned image repositories and tags.
  3. Add newly required variables.
  4. Verify kubeconfig and NAMESPACE=agent-farm.
Shell
ENV_FILE=.env.deploy bash deploy.sh
Required monitoring values
SLACK_ALERTS_WEBHOOK_URL=
GRAFANA_ADMIN_PASSWORD=
GRAFANA_HOST=

Understand Helmfile ordering

  1. Application, LiteLLM, and Firecrawl PostgreSQL
  2. Redis
  3. LiteLLM and Firecrawl
  4. API hooks, Product API, Ingest, Communications, and workers
  5. Agent Barn UI
  6. Monitoring
Helmfile behavior
helmfile -f helmfile.yaml.gotmpl sync --wait

Default timeout: 600 seconds

Expect brief service interruption

API, UI, worker, and LiteLLM are single-replica, non-surge rollouts in the constrained namespace. LiteLLM specifically uses:

LiteLLM rollout strategy
maxSurge: 0
maxUnavailable: 1

Watch upgrade hooks

The API chart completes ordered pre-install and pre-upgrade work before its rollout.

1Secrets and hook resourcesweight −10 2LiteLLM virtual keyweight −5 3Alembic migrationweight −1 4API rolloutDeployment

Secret and configuration hooks

The API Secret, registry pull Secret, and key-generation script run at weight -10. Existing pods reread changed envFrom values only after replacement.

LiteLLM virtual-key hook

At weight -5, agentbarn-api-litellm-key waits for LiteLLM, deletes the agentbarn-api alias, generates a replacement key, and updates litellm-api-key. A later failure can leave old API and worker pods holding the invalid previous key.

Inspect the LiteLLM key Job
kubectl get job agentbarn-api-litellm-key \
  --namespace agent-farm

kubectl logs job/agentbarn-api-litellm-key \
  --namespace agent-farm

Database migration hook

The API chart runs agentbarn-api-migrate as a pre-install,pre-upgrade Helm hook at weight -1, using the target API image to execute cd /app/api && alembic upgrade head. Review every migration between the deployed and target tags, confirm the target database backup and recovery plan, and treat a failed hook as a failed release. Do not force the application rollout past it, edit Alembic state, or manually alter Communication records to get through an upgrade.

Inspect the migration Job
kubectl get job agentbarn-api-migrate \
  --namespace agent-farm

kubectl describe job agentbarn-api-migrate \
  --namespace agent-farm

kubectl logs job/agentbarn-api-migrate \
  --namespace agent-farm
Watch primary rollouts
kubectl get pods --namespace agent-farm --watch

kubectl rollout status deployment/agentbarn-api \
  --namespace agent-farm --timeout=10m

kubectl rollout status deployment/agentbarn-api-worker \
  --namespace agent-farm --timeout=10m

kubectl rollout status deployment/agentbarn-ui \
  --namespace agent-farm --timeout=10m
Inspect release status
helm list --namespace agent-farm
helm status agentbarn-api --namespace agent-farm
helm status agentbarn-ui --namespace agent-farm

When supplied, the deployment commit annotation rolls API, UI, and worker pods even though k3s branch tags are mutable. The API release also serves multiple independently checked processes: Product API, Ingest API, Communications, the Domain Event delivery worker, Domain Event reconciliation workload, and the migration hook.

Verify the Communications upgrade surface

Communications is a separately served process. Its Deployment, internal ClusterIP Service, port 8002, /health, /metrics, Runtime protocol base URL, Connection-scoped public webhook ingress, database-backed ingress leases, durable Deliveries, and journal retention must be checked independently. Product API readiness does not prove the Communications rollout or end-to-end provider delivery succeeded.

Check the Communications rollout and process health

Run these commands from a terminal with kubectl configured for the cluster you just upgraded.

The examples use Agent Barn's standard namespace, agent-farm, and API Helm release name, agentbarn-api. The Communications Deployment and Service are both named agentbarn-api-communications.

If your installation uses a different namespace, replace agent-farm. If you changed the API release name, replace agentbarn-api-communications with <your-api-release-name>-communications.

1. Wait for the Communications Deployment

Communications rollout status
kubectl rollout status deployment/agentbarn-api-communications \
  --namespace agent-farm

Wait for Kubernetes to report that the Deployment successfully rolled out. If it reports a failure, inspect the Communications workload before continuing with the upgrade checks.

2. Confirm that the Service exists

Communications Service
kubectl get service agentbarn-api-communications \
  --namespace agent-farm

The output should list the Communications Service with port 8002/TCP.

The Service is internal to the cluster. Its existence alone does not confirm that the application is healthy.

3. Check process health from inside the cluster

The following command creates a temporary pod in the same namespace, calls the internal health endpoint, and removes the pod when it exits.

Your Kubernetes account needs permission to create and attach to a pod in this namespace. The cluster also needs to be able to pull the curlimages/curl image.

Communications process health
kubectl run communications-health-check \
  --namespace agent-farm \
  --rm -i \
  --restart=Never \
  --image=curlimages/curl \
  -- curl --fail --show-error \
  http://agentbarn-api-communications:8002/health

A successful response contains:

Successful response
{"status":"ok"}

If you use a different API Helm release name, update the hostname in this command as well as the resource names in the earlier commands.

This response confirms that the Communications process is reachable and responding. It does not confirm that every Slack, Microsoft Teams, Telegram, or Discord connection is working.

4. Check an actual connection

After the process check succeeds:

  1. Open Agent Barn and select a running Agent with a configured chat connection.
  2. Review that connection's health.
  3. Send a message from a location and user allowed by the connection's settings. Mention the bot if its policy requires it.
  4. Confirm that the Agent replies through the same connection.

If the health endpoint succeeds but messaging fails, inspect the affected connection and its delivery diagnostics. See Communication Diagnostics.

Also confirm on the Communications surface

  • /metrics remains internally scrapeable.
  • Webhook ingress exposes only the required Communications prefix; webhook Connections retain their generated Connection-specific URLs.
  • Supervised Connections reconcile, statuses do not remain unexpectedly PENDING, CONNECTING, DEGRADED, or ERROR, and ingress leases do not show persistent contention.
  • Queue depth, oldest queued Delivery age, dead letters, latency, reconnects, and policy dispositions remain within the normal profile.

After migration

Confirm existing Connections remain scoped to their correct Agents and Organizations, encrypted credentials remain readable without exposure, enabled Connections are reconciled, and status transitions resume. Account for pending or processing Deliveries, dead-letter history, and journal entries; keep Conversation Messages scoped by Connection and provider location. Regenerate Runtime protocol credentials through normal Agent start behavior when required; do not move Connection credentials into Agent Secrets or Runtime configuration.

Replica coordination

Communications replicas coordinate supervised provider ingress through database leases. Expect leases to transfer while old replicas terminate and new replicas become ready; watch for reconnect spikes, sustained duplicate sessions, lease contention, and Connections stuck in CONNECTING or ERROR. Do not manually assign provider sessions to pods or run independent Telegram pollers for the same bot outside Agent Barn.

Verify the upgrade

Do not declare success from a green workflow or Helm result alone.

Infrastructure state
kubectl get pods,deployments,statefulsets,jobs,cronjobs,persistentvolumeclaims \
  --namespace agent-farm
API health example
curl --fail https://api.agentbarn.example.com/api/v1/health
Target database revision
kubectl exec \
  --namespace agent-farm \
  deployment/agentbarn-api \
  -- sh -c 'cd /app/api && alembic current'

Infrastructure

  • Pods ready; no unexplained crashes
  • PVCs retained
  • Migration and Domain Event Jobs completed
  • Ingress, certificates, and monitoring ready

Product and data

  • Product API, Ingest, UI, authentication, and Platform access
  • Organizations and owned resources present
  • Alembic matches release head
  • Workers, Redis, and Event Deliveries healthy

Providers and signals

  • Communications target availability and Connection status counts
  • Delivery queue depth, oldest age, outcomes, latency, reconnects, and policy dispositions
  • LiteLLM, OpenRouter, Firecrawl, email, and OAuth canaries
  • Agent Runtime health, scrape targets, dashboards, and alerts

Record final source or tag, images, Helm revisions, Alembic revision, backup identifiers, results, Agent restart status, and follow-up work.

Upgrade running Agents

A platform upgrade changes the runtime references used for future builds; it does not replace existing Agent Deployments. Each running Agent retains its image, ConfigMap, Secret, ingest identity, and Deployment.

Stop and start an Agent through Agent Barn only when its Runtime, shared protocol, or Agent resources need replacement. That rebuilds its resources from the selected Runtime, pinned Template or Override, Skills, credentials, current builders and Runtime image, and a fresh protocol credential. A Connection-only change does not require restarting every Agent.

Runtime canaries

  1. Start a Hermes Agent and an OpenClaw Agent when the Runtime or Runtime-neutral protocol changes.
  2. Confirm each Runtime Communications adapter reaches the internal Communications Service.
  3. Confirm an inbound Delivery can be claimed, a reply can be submitted against its source Delivery, and the Delivery reaches its terminal state.
  4. Confirm Tool Call telemetry reaches Ingest separately.

Platform Plugin canaries

Platform Plugins are code-owned artifacts in the Agent Barn API release; Hermes and OpenClaw use the same versioned Runtime-neutral Communications protocol. When Communications or a Platform Plugin changes, test the affected transport independently: Slack supervised Socket Mode, Telegram supervised polling, Discord supervised Gateway, or Microsoft Teams Connection-scoped webhook ingress.

  1. Confirm the Connection reconciles and a provider event is observed.
  2. Confirm policy admits or intentionally rejects the event; an admitted event becomes a durable inbound Delivery.
  3. Confirm the Runtime claims and processes it, the reply becomes a durable outbound Delivery, and the provider accepts it.
  4. Confirm the Communications journal shows the expected stages. Use dedicated test Connections or non-production provider locations where possible.

Choose canaries by the changed boundary

Changed areaRequired canary focus
Product API onlyProduct operations, authorization, and Activity reads
IngestTool Call telemetry from affected Runtime adapters
Communications coreRepresentative supervised and webhook Connections, queue processing, and journal
Runtime-neutral protocolHermes and OpenClaw independently
One Platform PluginThat Platform’s Connection setup, ingress, policy, and outbound Delivery
Connection schemaExisting and newly created Connections for affected Platforms
Delivery persistencePending, retrying, succeeded, and dead-lettered behavior
Runtime imageThe affected Runtime without assuming a Platform pairing
MonitoringChanged scrape targets, dashboards, and alert behavior

Do not require a complete Runtime/Platform cross-product unless the shared protocol or Platform Plugin boundary itself changed. Observe canaries before moving to batches, and pause for increased errors, spend, restarts, provider failures, or Delivery backlog.

Rollback and recovery

Failure statePreferred response
Build or test failed before deploymentFix the release; production is unchanged
Helmfile failed before API hooksInspect every dependency release that may already have changed
LiteLLM key rotated but API rollout failedRestart or redeploy a compatible API and worker with the new key
Migration failed and rolled back cleanlyCorrect the migration and redeploy
Migration succeeded but API failedDeploy an image compatible with the migrated schema
API/UI regression with compatible schemaRevert source or redeploy the previous pinned bundle
Runtime regressionRestore the previous runtime reference, then restart only affected Agents
Irreversible data changeStop writes and restore the tested recovery set
Secret mismatchRestore the exact stable value belonging to the current data
Storage migration failureFollow the storage recovery plan; Helm rollback does not move data
Monitoring failureKeep application recovery separate and restore external visibility

Application and database rollback are separate decisions

Before any rollback, confirm the previous API image operates against the migrated schema, the migration is backward-compatible, and new Connection, Delivery, journal, Template, Skill, or Agent data remains readable. Also confirm the previous Runtime image supports the active Communications protocol and the previous Platform Plugin understands persisted Connection settings and schema versions. Preserve the deployed tag and diagnostic evidence. If schema compatibility is uncertain, stop and assess the migration rather than blindly downgrading Alembic or restoring the database.

k3s branch rollback

k3s API and UI use moving latest tags. A Helm rollback can restore an old manifest revision while still pulling the current image. Prefer a reviewed corrective or revert commit that rebuilds, republishes, rolls out with a new commit annotation, and passes verification.

Hosted public and release-bundle rollback

Where schema compatibility permits, redeploy the previous immutable API and UI tags; restore compatible Hermes and OpenClaw versions if they changed. Retain stable secrets unless restoring matching data and keys, then recheck Product API, Ingest, Communications, and workers independently. Confirm Connection reconciliation and Delivery processing resume, using the Communications journal to identify Deliveries that require operator action.

Database rollback

Helm rollback does not execute alembic downgrade. Use a tested, data-safe downgrade only when compatible with the restored application; otherwise use a forward fix or restore the database recovery set.

Agent runtime rollback

Restoring the old platform runtime reference does not change Agents already restarted. Identify affected Agents, verify one canary on the restored reference, then proceed in batches.

Inspect partial Helmfile state

Shell
helm list --namespace agent-farm

helm status RELEASE_NAME --namespace agent-farm
helm history RELEASE_NAME --namespace agent-farm

Releases have independent histories. Do not assume they share a Helm revision or upgrade status.

Recover the correct component

  • Restart or roll Communications: replaces the service processes.
  • Request a Connection reconnect: increments that Connection’s revision and recreates its provider session.
  • Retry a Delivery: requeues an eligible dead-lettered outbound Delivery only after its provider, credential, policy, configuration, or Runtime failure is corrected.
  • Restart an Agent: replaces Runtime resources and protocol credentials.

Use the smallest recovery action that addresses the fault. Do not restart every Agent for Connection-only changes.

Troubleshooting

SymptomLikely causeResolution
Deployment from a feature branch fails immediatelyOnly staging and main may deployPromote through a supported branch
An unchanged component was rebuiltManual dispatch or no valid successful baseline was availableConfirm the workflow warning and complete full verification
A changed component was not rebuiltChange detection missed its source pathStop deployment and review the detected-component outputs
Helmfile reports a required value missingThe release introduced a configuration requirementCompare the new environment with the bundle and documentation
PostgreSQL authentication failsA password changed while the initialized volume retained the old valueRestore the old value or run a coordinated credential migration
Existing credentials cannot decryptAGENT_TOKEN_ENCRYPTION_KEY changedRestore the key matching the application database
Sessions fail after upgradeSECRET_SIGNING_KEY changedRestore the old key or complete a planned session migration
Agent model calls fail after an API hook failureThe LiteLLM key changed while old pods retain the previous valueRedeploy compatible API and worker pods so they load the Secret
Migration Job failsThe target revision cannot apply to the databaseInspect the retained Job, logs, current revision, and backup
Helm rollback does not restore behaviorDatabase state or moving image tags stayed on the new releaseDeploy a compatible explicit image or follow the recovery plan
LiteLLM briefly becomes unavailableIts non-surge single-replica rollout replaced the old pod firstWait for readiness and schedule future work in a maintenance window
Upgrade times out after ten minutesA hook or workload exceeded the 600-second timeoutInspect Jobs, pods, events, and release status before retrying
Only some platform releases upgradedHelmfile failed after earlier dependencies completedInventory every release and recover from actual state
Running Agents still use the old runtimePlatform deployment does not rebuild Agent workloadsStop and start Agents through a canary and batch rollout
Restarted Agents behave differentlyThe new image or generated configuration was appliedCompare canary logs, settings, pinned versions, and integrations
Agent PVC data is missingThe upgrade recreated or changed storageStop further restarts and follow the tested PVC recovery plan
UI and API disagree on a contractIncompatible images were deployedDeploy the matching API/UI pair
Grafana loses visibilityMonitoring or target discovery failedUse Kubernetes and external monitoring while restoring the release
New image cannot be pulledTag, credentials, repository, or pull Secret is wrongVerify the pinned image and registry hook before retrying
Staging creates Agent resources in productionK8S_NAMESPACE or API kubeconfig is wrongStop staging and correct namespace-scoped configuration

Upgrade checklist

  • Before staging

  • Before production

  • After production

Next steps

Keep production configuration, Kubernetes deployment, migrations, self-hosting Communications, and monitoring available during the rollout. For operational follow-up, use Communication Connections, Communication diagnostics, Runtime deployment, and the Slack, Microsoft Teams, Telegram, and Discord setup guides.

Continue the self-hosting sequence Troubleshoot self-hosting → Diagnose deployment, database, runtime, provider, and monitoring failures.

Roll out initiated delivery

Initiated delivery depends on compatible application schema/code, runtime-generated scripts, and runtime image hooks. Roll the application through its normal migration workflow, deploy the matching runtime images, and recreate affected Agent runtime configuration through the supported stop/start flow. Existing pods do not acquire new image hooks or generated scripts merely because a new application image was built. Preserve the Agent's persistent state when performing a normal restart.

Hermes session continuity

Agent Barn's Hermes adapter sends resume_session: true with a stable Connection/location/thread session identity. The shipped Hermes image patches its run endpoint to reload persisted conversation history, including compaction lineage and tool-call metadata. Deploy the compatible patched image and adapter together when adopting this behavior.

Existing session state remains on the Agent's persistent volume. New sessions can legitimately have no history. Do not recreate storage as a routine response to missing continuity, and do not assume updating the application automatically replaces every running Agent pod or image. Recreate the affected runtime through the normal Agent lifecycle after selecting the intended compatible runtime image.

Runtime images have their own version sources. A product release tag is not a substitute for the Hermes or OpenClaw image version. Preserve persistent state when rolling runtime configuration.

See Choose an Agent Runtime for transport and replacement boundaries.

Client image release

The manual Client Release workflow publishes customer API, UI, Hermes-base, and OpenClaw-base images without deploying a customer cluster. It is separate from k3s and public hosted deployments. See Client image release for registry settings, independent Runtime versions, and backend topology assumptions.

Documentation