Upgrade outcome
A staged release with a deliberate recovery path
Complete this guide to move the platform and selected Agents forward without losing track of data, keys, images, or partial deployment state.
- Staging proves the complete release and every affected runtime/platform pair.
- Backups, stable secrets, image references, and Alembic state are recorded.
- Hooks and rollouts are watched before product and provider verification.
- Agent runtime adoption proceeds through canaries and controlled batches.
Overview
An Agent Barn upgrade can change the API, UI, worker, database schema, supporting services, monitoring, Agent builders, and runtime images.
GitHub branch deployments and versioned release bundles use the same Helmfile release graph. Neither path automatically rebuilds existing Agent workloads, so platform deployment and runtime rollout remain separate operations.
Understand the upgrade boundaries
| Upgrade area | What can change | Primary risk |
|---|---|---|
| API | Routes, services, workers, Agent builders, integrations | Contract, migration, or background-processing failure |
| Communications | Connection supervision, ingress leases, Deliveries, journal, and provider transports | A healthy Product API does not prove provider delivery or Connection recovery |
| Ingest | Runtime Tool Call telemetry and Ingest API availability | Telemetry can fail independently of Product and Communications |
| UI | Pages, schemas, queries, onboarding flows | UI and API incompatibility |
| Application database | Alembic schema and data migrations | Irreversible data or compatibility change |
| Hermes runtime | Base image and runtime behavior | Existing Hermes Agents retain the old image until restarted |
| OpenClaw runtime | Base image and runtime behavior | Existing OpenClaw Agents retain the old image until restarted |
| LiteLLM | Proxy image, configuration, master key, virtual keys | Brief interruption or invalid Agent and API keys |
| Firecrawl | API, browser service, RabbitMQ, database integration | Integration or upstream-image incompatibility |
| PostgreSQL and Redis | Stateful service images and configuration | Data compatibility and availability |
| Helm charts | Deployments, Services, Secrets, Jobs, ingress, storage | Partial or incompatible Kubernetes rollout |
| Monitoring | Rules, dashboards, Grafana, Alertmanager | Lost visibility or alert delivery |
| Environment configuration | Credentials, URLs, models, storage, OAuth, email | Cross-environment or secret mismatch |
Changes that are not routine upgrades
- Changing PostgreSQL passwords on initialized volumes.
- Rotating
SECRET_SIGNING_KEY,AGENT_TOKEN_ENCRYPTION_KEY, orLITELLM_MASTER_KEY. - Changing a StatefulSet’s existing StorageClass.
- Renaming
agent-farmoragent-farm-staging. - Moving persistent data to another cluster or storage provider.
Choose an upgrade path
Source-operated environments
GitHub branch deployment
Use the agent-barn repository when environments deploy directly from source.
| Branch | Environment | Namespace | API/UI tags | Runtime tags |
|---|---|---|---|---|
staging | Staging on k3s | agent-farm-staging | latest-staging | <version>-staging |
main | AAI Labs testing ground on k3s | agent-farm | latest | <version> |
vX.Y.Z tag | Hosted public production on Talos | Dedicated Talos cluster | vX.Y.Z | Selected runtime VERSION |
Pushes and manual dispatch run only from staging or main; any other branch is rejected. Per-branch concurrency serializes runs and does not cancel an in-progress deployment. These k3s paths are isolated staging and the AAI Labs testing ground; main is not the hosted public-production source.
Change detection compares with the latest successful deployment on that branch. A failure does not advance the baseline. Manual dispatch, a missing baseline, or an unavailable or non-ancestor baseline rebuilds API, UI, Hermes, and OpenClaw.
Packaged self-hosting
Release-bundle deployment
This path applies only when a separate deployment bundle has been provided for your release. A source archive uses .env.deploy.spec instead.
A bundle supplies helm/, k8s/, helmfile.yaml.gotmpl, deploy.sh, and a generated .env.deploy.
| Component | Version source |
|---|---|
| API | API_IMAGE_TAG |
| UI | UI_IMAGE_TAG |
| Hermes runtime | Hermes VERSION file |
| OpenClaw runtime | OpenClaw VERSION file |
| Helm chart | Chart version for chart packaging |
Transfer environment-specific values into the new bundle’s file. Do not replace it wholesale with an older copy because requirements can change between releases.
Hosted public production on Talos
The dedicated Talos cluster deploys from a release tag matching vX.Y.Z through .github/workflows/deploy-public.yml. The workflow publishes API and UI images to registry.agentbarn.dev, pins both image inputs to that exact tag, and uses PUBLIC_-prefixed cluster configuration and secrets.
Cluster: dedicated Talos cluster
Workflow: .github/workflows/deploy-public.yml
Release tag: vX.Y.Z
API/UI tags: vX.Y.Z (pinned)
Registry: registry.agentbarn.dev
Configuration: PUBLIC_ prefixed values and secretsIt does not update the k3s registry’s moving latest tags. A manual public deployment can target an existing tag and use skip_build when the tag-pinned images already exist.
Before you begin
Assign a release owner, database and backup owner, verification owner, incident decision-maker, Agent restart owner, and communication owner.
- Confirm staging, production monitoring, DNS, certificates, storage, registries, and providers are healthy.
- Confirm there is no active incident or overlapping automated or manual deployment.
- Verify operators can reach Kubernetes, Helm history, application logs, and backups.
- Confirm every required image exists before approval.
Prepare the target release
- Select the target release from Agent Barn releases and read its release notes.
- Download its source ZIP or tar.gz and extract it into a separate working directory. If a matching deployment bundle has been provided instead, extract that separately and use its included configuration.
- For a source archive, copy
.env.deploy.specto.env.deploy. Carry forward your installation's infrastructure settings and credentials, including stable signing and encryption keys. Account for settings added or changed in the target release. - Obtain the image locations and component versions intended for the target release. Preserve separate API/UI and runtime version selections; source download availability does not establish image availability.
- Continue with the backup and migration preparation below before deploying.
API and UI use the selected product release tag. Runtime image tags remain separate; do not replace every component tag with the product tag or with latest.
Record the current state
Capture enough evidence to distinguish an upgrade regression from pre-existing state and to reproduce the current deployment.
helm list --namespace agent-farm
helm history agentbarn-api --namespace agent-farm
helm history agentbarn-ui --namespace agent-farm
helm history litellm --namespace agent-farm
helm history monitoring --namespace agent-farmkubectl get deployments,statefulsets \
--namespace agent-farm \
-o custom-columns='KIND:.kind,NAME:.metadata.name,IMAGES:.spec.template.spec.containers[*].image'
kubectl get cronjobs \
--namespace agent-farm \
-o custom-columns='NAME:.metadata.name,IMAGES:.spec.jobTemplate.spec.template.spec.containers[*].image'kubectl exec \
--namespace agent-farm \
deployment/agentbarn-api \
-- sh -c 'cd /app/api && alembic current'kubectl get pods,deployments,statefulsets,jobs,cronjobs,persistentvolumeclaims \
--namespace agent-farmRecord the source commit or bundle version, image references, Helm revisions, Alembic revision, backup identifiers, Agent count by runtime and platform, restart candidates, and known warnings.
Review the release
Application contracts
- API and UI schema changes
- Authorization and tenancy
- Alembic migrations
- Workers and Event Deliveries
- Agent builders, Templates, Skills, and credentials
- Provider and monitoring changes
Deployment contracts
- Charts, dependencies, images, and strategies
- Resources, Services, ingress, and certificates
- Secrets, ConfigMaps, Jobs, hooks, and CronJobs
- PVC size and StorageClass
- Helmfile dependencies and Kubernetes permissions
Understand version identifiers
| Identifier | Meaning |
|---|---|
| Git commit or PR | Source identifier for the k3s branch deployment |
| API_IMAGE_TAG | Explicit API deployment input; latest-staging or latest on k3s, vX.Y.Z on hosted public production |
| UI_IMAGE_TAG | Explicit UI deployment input; the hosted public release uses the same vX.Y.Z tag as the API |
| Hermes/OpenClaw VERSION | Independent Runtime image version from each Runtime VERSION file |
| Helm chart version | Packaging version that changes with chart templates or values |
| Git release tag | Immutable public-release identifier; not a moving k3s image tag |
| Alembic revision | Application database schema state |
API and UI images are explicit API_IMAGE_TAG and UI_IMAGE_TAG inputs, not values derived from chart appVersion. A chart version changes when its templates or values change; branch application updates and documentation-only changes do not require an appVersion or service-image-tag change. There is no shared API/UI application number outside selected deployment inputs; hosted public API and UI releases deliberately share the same immutable vX.Y.Z tag.
Back up state
Back up Agent Barn, LiteLLM, and Firecrawl PostgreSQL databases; Agent PVCs when workspace continuity matters; environment configuration; stable keys; database credentials; and external provider configuration needed for recovery.
SECRET_SIGNING_KEY
AGENT_TOKEN_ENCRYPTION_KEY
LITELLM_MASTER_KEY
POSTGRES_APP_PASSWORD
POSTGRES_LITELLM_PASSWORD
POSTGRES_FIRECRAWL_PASSWORDA database backup without its matching encryption key can leave Agent credentials unreadable. A LiteLLM database backup without its master key can leave virtual keys unusable.
- Confirm the backup is complete and encrypted.
- Store it outside the affected namespace.
- Record its identifier in the release record.
- Verify restoration through a tested procedure.
Test in staging
Deploy the complete release to agent-farm-staging and wait for the staging workflow to finish.
Namespace: agent-farm-staging
Environment: staging
API/UI tag: latest-staging
Runtime suffix: -stagingPlatform
- API, UI, login, and Organization isolation
- Invitations, email, and Google OAuth
- Database migration and restart recovery
- Workers and Event Deliveries
Agents
- Create, start, stop, and restart
- Hermes on every used platform
- OpenClaw on every used platform
- Messages, Tool Calls, costs, and credentials
Providers
- LiteLLM and OpenRouter requests
- Firecrawl and Shared Credentials
- Grafana and alert delivery
- Canary for every affected runtime/platform pair
Account for shared providers
Namespace separation does not guarantee provider isolation. Registry, OpenRouter, Google OAuth, Cloudflare, and Slack alert delivery may be shared. Avoid tests that exhaust quotas, revoke shared credentials, or send unwanted production-facing messages.
Preview the upgrade
For a manual deployment, load the protected environment and inspect the Helmfile diff.
set -a
source .env.deploy
set +a
export POD_KUBECONFIG_B64="$(base64 "$KUBECONFIG" | tr -d '\n')"
helmfile -f helmfile.yaml.gotmpl diff --suppress-secrets
unset POD_KUBECONFIG_B64- Review StatefulSet replacement, PVC, StorageClass, database Secret, and stable-key changes.
- Review images, resources, ingress, ServiceAccounts, hooks, deletions, and namespace scope.
The GitHub path proceeds directly to helmfile sync --wait. Pull-request review, source diff, and staging provide its preview boundary.
Run the upgrade
Hosted public release
- Select a commit already present on
main. - Review migrations, required configuration, chart changes, API/UI changes, Runtime versions, and Platform Plugin changes.
- Confirm public-only configuration and secrets exist.
- Create a unique
vX.Y.Ztag for that commit. - Push the tag to start the public deployment workflow.
- Allow the workflow to build and publish tag-pinned API and UI images, or use manual
skip_buildonly when the tagged images already exist. - Allow Helmfile to apply releases in dependency order.
- Confirm the migration hook succeeds before relying on the new application processes.
- Check Product API, Ingest, Communications, workers, UI, and monitoring.
- Run change-specific Runtime, Connection, and Delivery canaries.
- Record the deployed tag and any migration or compatibility constraints.
git tag vX.Y.Z
git push origin vX.Y.ZRelease tags are immutable identifiers: make corrections with a new version tag, never by moving one already deployed.
Release-bundle deployment
- Fill the new
.env.deploywith existing environment values. - Preserve pinned image repositories and tags.
- Add newly required variables.
- Verify kubeconfig and
NAMESPACE=agent-farm.
ENV_FILE=.env.deploy bash deploy.shSLACK_ALERTS_WEBHOOK_URL=
GRAFANA_ADMIN_PASSWORD=
GRAFANA_HOST=Understand Helmfile ordering
- Application, LiteLLM, and Firecrawl PostgreSQL
- Redis
- LiteLLM and Firecrawl
- API hooks, Product API, Ingest, Communications, and workers
- Agent Barn UI
- Monitoring
helmfile -f helmfile.yaml.gotmpl sync --wait
Default timeout: 600 secondsExpect brief service interruption
API, UI, worker, and LiteLLM are single-replica, non-surge rollouts in the constrained namespace. LiteLLM specifically uses:
maxSurge: 0
maxUnavailable: 1Watch upgrade hooks
The API chart completes ordered pre-install and pre-upgrade work before its rollout.
Secret and configuration hooks
The API Secret, registry pull Secret, and key-generation script run at weight -10. Existing pods reread changed envFrom values only after replacement.
LiteLLM virtual-key hook
At weight -5, agentbarn-api-litellm-key waits for LiteLLM, deletes the agentbarn-api alias, generates a replacement key, and updates litellm-api-key. A later failure can leave old API and worker pods holding the invalid previous key.
kubectl get job agentbarn-api-litellm-key \
--namespace agent-farm
kubectl logs job/agentbarn-api-litellm-key \
--namespace agent-farmDatabase migration hook
The API chart runs agentbarn-api-migrate as a pre-install,pre-upgrade Helm hook at weight -1, using the target API image to execute cd /app/api && alembic upgrade head. Review every migration between the deployed and target tags, confirm the target database backup and recovery plan, and treat a failed hook as a failed release. Do not force the application rollout past it, edit Alembic state, or manually alter Communication records to get through an upgrade.
kubectl get job agentbarn-api-migrate \
--namespace agent-farm
kubectl describe job agentbarn-api-migrate \
--namespace agent-farm
kubectl logs job/agentbarn-api-migrate \
--namespace agent-farmkubectl get pods --namespace agent-farm --watch
kubectl rollout status deployment/agentbarn-api \
--namespace agent-farm --timeout=10m
kubectl rollout status deployment/agentbarn-api-worker \
--namespace agent-farm --timeout=10m
kubectl rollout status deployment/agentbarn-ui \
--namespace agent-farm --timeout=10mhelm list --namespace agent-farm
helm status agentbarn-api --namespace agent-farm
helm status agentbarn-ui --namespace agent-farmWhen supplied, the deployment commit annotation rolls API, UI, and worker pods even though k3s branch tags are mutable. The API release also serves multiple independently checked processes: Product API, Ingest API, Communications, the Domain Event delivery worker, Domain Event reconciliation workload, and the migration hook.
Verify the Communications upgrade surface
Communications is a separately served process. Its Deployment, internal ClusterIP Service, port 8002, /health, /metrics, Runtime protocol base URL, Connection-scoped public webhook ingress, database-backed ingress leases, durable Deliveries, and journal retention must be checked independently. Product API readiness does not prove the Communications rollout or end-to-end provider delivery succeeded.
Check the Communications rollout and process health
Run these commands from a terminal with kubectl configured for the cluster you just upgraded.
The examples use Agent Barn's standard namespace, agent-farm, and API Helm release name, agentbarn-api. The Communications Deployment and Service are both named agentbarn-api-communications.
If your installation uses a different namespace, replace agent-farm. If you changed the API release name, replace agentbarn-api-communications with <your-api-release-name>-communications.
1. Wait for the Communications Deployment
kubectl rollout status deployment/agentbarn-api-communications \
--namespace agent-farmWait for Kubernetes to report that the Deployment successfully rolled out. If it reports a failure, inspect the Communications workload before continuing with the upgrade checks.
2. Confirm that the Service exists
kubectl get service agentbarn-api-communications \
--namespace agent-farmThe output should list the Communications Service with port 8002/TCP.
The Service is internal to the cluster. Its existence alone does not confirm that the application is healthy.
3. Check process health from inside the cluster
The following command creates a temporary pod in the same namespace, calls the internal health endpoint, and removes the pod when it exits.
Your Kubernetes account needs permission to create and attach to a pod in this namespace. The cluster also needs to be able to pull the curlimages/curl image.
kubectl run communications-health-check \
--namespace agent-farm \
--rm -i \
--restart=Never \
--image=curlimages/curl \
-- curl --fail --show-error \
http://agentbarn-api-communications:8002/healthA successful response contains:
{"status":"ok"}If you use a different API Helm release name, update the hostname in this command as well as the resource names in the earlier commands.
This response confirms that the Communications process is reachable and responding. It does not confirm that every Slack, Microsoft Teams, Telegram, or Discord connection is working.
4. Check an actual connection
After the process check succeeds:
- Open Agent Barn and select a running Agent with a configured chat connection.
- Review that connection's health.
- Send a message from a location and user allowed by the connection's settings. Mention the bot if its policy requires it.
- Confirm that the Agent replies through the same connection.
If the health endpoint succeeds but messaging fails, inspect the affected connection and its delivery diagnostics. See Communication Diagnostics.
Also confirm on the Communications surface
/metricsremains internally scrapeable.- Webhook ingress exposes only the required Communications prefix; webhook Connections retain their generated Connection-specific URLs.
- Supervised Connections reconcile, statuses do not remain unexpectedly
PENDING,CONNECTING,DEGRADED, orERROR, and ingress leases do not show persistent contention. - Queue depth, oldest queued Delivery age, dead letters, latency, reconnects, and policy dispositions remain within the normal profile.
After migration
Confirm existing Connections remain scoped to their correct Agents and Organizations, encrypted credentials remain readable without exposure, enabled Connections are reconciled, and status transitions resume. Account for pending or processing Deliveries, dead-letter history, and journal entries; keep Conversation Messages scoped by Connection and provider location. Regenerate Runtime protocol credentials through normal Agent start behavior when required; do not move Connection credentials into Agent Secrets or Runtime configuration.
Replica coordination
Communications replicas coordinate supervised provider ingress through database leases. Expect leases to transfer while old replicas terminate and new replicas become ready; watch for reconnect spikes, sustained duplicate sessions, lease contention, and Connections stuck in CONNECTING or ERROR. Do not manually assign provider sessions to pods or run independent Telegram pollers for the same bot outside Agent Barn.
Verify the upgrade
Do not declare success from a green workflow or Helm result alone.
kubectl get pods,deployments,statefulsets,jobs,cronjobs,persistentvolumeclaims \
--namespace agent-farmcurl --fail https://api.agentbarn.example.com/api/v1/healthkubectl exec \
--namespace agent-farm \
deployment/agentbarn-api \
-- sh -c 'cd /app/api && alembic current'Infrastructure
- Pods ready; no unexplained crashes
- PVCs retained
- Migration and Domain Event Jobs completed
- Ingress, certificates, and monitoring ready
Product and data
- Product API, Ingest, UI, authentication, and Platform access
- Organizations and owned resources present
- Alembic matches release head
- Workers, Redis, and Event Deliveries healthy
Providers and signals
- Communications target availability and Connection status counts
- Delivery queue depth, oldest age, outcomes, latency, reconnects, and policy dispositions
- LiteLLM, OpenRouter, Firecrawl, email, and OAuth canaries
- Agent Runtime health, scrape targets, dashboards, and alerts
Record final source or tag, images, Helm revisions, Alembic revision, backup identifiers, results, Agent restart status, and follow-up work.
Upgrade running Agents
A platform upgrade changes the runtime references used for future builds; it does not replace existing Agent Deployments. Each running Agent retains its image, ConfigMap, Secret, ingest identity, and Deployment.
Stop and start an Agent through Agent Barn only when its Runtime, shared protocol, or Agent resources need replacement. That rebuilds its resources from the selected Runtime, pinned Template or Override, Skills, credentials, current builders and Runtime image, and a fresh protocol credential. A Connection-only change does not require restarting every Agent.
Runtime canaries
- Start a Hermes Agent and an OpenClaw Agent when the Runtime or Runtime-neutral protocol changes.
- Confirm each Runtime Communications adapter reaches the internal Communications Service.
- Confirm an inbound Delivery can be claimed, a reply can be submitted against its source Delivery, and the Delivery reaches its terminal state.
- Confirm Tool Call telemetry reaches Ingest separately.
Platform Plugin canaries
Platform Plugins are code-owned artifacts in the Agent Barn API release; Hermes and OpenClaw use the same versioned Runtime-neutral Communications protocol. When Communications or a Platform Plugin changes, test the affected transport independently: Slack supervised Socket Mode, Telegram supervised polling, Discord supervised Gateway, or Microsoft Teams Connection-scoped webhook ingress.
- Confirm the Connection reconciles and a provider event is observed.
- Confirm policy admits or intentionally rejects the event; an admitted event becomes a durable inbound Delivery.
- Confirm the Runtime claims and processes it, the reply becomes a durable outbound Delivery, and the provider accepts it.
- Confirm the Communications journal shows the expected stages. Use dedicated test Connections or non-production provider locations where possible.
Choose canaries by the changed boundary
| Changed area | Required canary focus |
|---|---|
| Product API only | Product operations, authorization, and Activity reads |
| Ingest | Tool Call telemetry from affected Runtime adapters |
| Communications core | Representative supervised and webhook Connections, queue processing, and journal |
| Runtime-neutral protocol | Hermes and OpenClaw independently |
| One Platform Plugin | That Platform’s Connection setup, ingress, policy, and outbound Delivery |
| Connection schema | Existing and newly created Connections for affected Platforms |
| Delivery persistence | Pending, retrying, succeeded, and dead-lettered behavior |
| Runtime image | The affected Runtime without assuming a Platform pairing |
| Monitoring | Changed scrape targets, dashboards, and alert behavior |
Do not require a complete Runtime/Platform cross-product unless the shared protocol or Platform Plugin boundary itself changed. Observe canaries before moving to batches, and pause for increased errors, spend, restarts, provider failures, or Delivery backlog.
Rollback and recovery
| Failure state | Preferred response |
|---|---|
| Build or test failed before deployment | Fix the release; production is unchanged |
| Helmfile failed before API hooks | Inspect every dependency release that may already have changed |
| LiteLLM key rotated but API rollout failed | Restart or redeploy a compatible API and worker with the new key |
| Migration failed and rolled back cleanly | Correct the migration and redeploy |
| Migration succeeded but API failed | Deploy an image compatible with the migrated schema |
| API/UI regression with compatible schema | Revert source or redeploy the previous pinned bundle |
| Runtime regression | Restore the previous runtime reference, then restart only affected Agents |
| Irreversible data change | Stop writes and restore the tested recovery set |
| Secret mismatch | Restore the exact stable value belonging to the current data |
| Storage migration failure | Follow the storage recovery plan; Helm rollback does not move data |
| Monitoring failure | Keep application recovery separate and restore external visibility |
Application and database rollback are separate decisions
Before any rollback, confirm the previous API image operates against the migrated schema, the migration is backward-compatible, and new Connection, Delivery, journal, Template, Skill, or Agent data remains readable. Also confirm the previous Runtime image supports the active Communications protocol and the previous Platform Plugin understands persisted Connection settings and schema versions. Preserve the deployed tag and diagnostic evidence. If schema compatibility is uncertain, stop and assess the migration rather than blindly downgrading Alembic or restoring the database.
k3s branch rollback
k3s API and UI use moving latest tags. A Helm rollback can restore an old manifest revision while still pulling the current image. Prefer a reviewed corrective or revert commit that rebuilds, republishes, rolls out with a new commit annotation, and passes verification.
Hosted public and release-bundle rollback
Where schema compatibility permits, redeploy the previous immutable API and UI tags; restore compatible Hermes and OpenClaw versions if they changed. Retain stable secrets unless restoring matching data and keys, then recheck Product API, Ingest, Communications, and workers independently. Confirm Connection reconciliation and Delivery processing resume, using the Communications journal to identify Deliveries that require operator action.
Database rollback
Helm rollback does not execute alembic downgrade. Use a tested, data-safe downgrade only when compatible with the restored application; otherwise use a forward fix or restore the database recovery set.
Agent runtime rollback
Restoring the old platform runtime reference does not change Agents already restarted. Identify affected Agents, verify one canary on the restored reference, then proceed in batches.
Inspect partial Helmfile state
helm list --namespace agent-farm
helm status RELEASE_NAME --namespace agent-farm
helm history RELEASE_NAME --namespace agent-farmReleases have independent histories. Do not assume they share a Helm revision or upgrade status.
Recover the correct component
- Restart or roll Communications: replaces the service processes.
- Request a Connection reconnect: increments that Connection’s revision and recreates its provider session.
- Retry a Delivery: requeues an eligible dead-lettered outbound Delivery only after its provider, credential, policy, configuration, or Runtime failure is corrected.
- Restart an Agent: replaces Runtime resources and protocol credentials.
Use the smallest recovery action that addresses the fault. Do not restart every Agent for Connection-only changes.
Troubleshooting
| Symptom | Likely cause | Resolution |
|---|---|---|
| Deployment from a feature branch fails immediately | Only staging and main may deploy | Promote through a supported branch |
| An unchanged component was rebuilt | Manual dispatch or no valid successful baseline was available | Confirm the workflow warning and complete full verification |
| A changed component was not rebuilt | Change detection missed its source path | Stop deployment and review the detected-component outputs |
| Helmfile reports a required value missing | The release introduced a configuration requirement | Compare the new environment with the bundle and documentation |
| PostgreSQL authentication fails | A password changed while the initialized volume retained the old value | Restore the old value or run a coordinated credential migration |
| Existing credentials cannot decrypt | AGENT_TOKEN_ENCRYPTION_KEY changed | Restore the key matching the application database |
| Sessions fail after upgrade | SECRET_SIGNING_KEY changed | Restore the old key or complete a planned session migration |
| Agent model calls fail after an API hook failure | The LiteLLM key changed while old pods retain the previous value | Redeploy compatible API and worker pods so they load the Secret |
| Migration Job fails | The target revision cannot apply to the database | Inspect the retained Job, logs, current revision, and backup |
| Helm rollback does not restore behavior | Database state or moving image tags stayed on the new release | Deploy a compatible explicit image or follow the recovery plan |
| LiteLLM briefly becomes unavailable | Its non-surge single-replica rollout replaced the old pod first | Wait for readiness and schedule future work in a maintenance window |
| Upgrade times out after ten minutes | A hook or workload exceeded the 600-second timeout | Inspect Jobs, pods, events, and release status before retrying |
| Only some platform releases upgraded | Helmfile failed after earlier dependencies completed | Inventory every release and recover from actual state |
| Running Agents still use the old runtime | Platform deployment does not rebuild Agent workloads | Stop and start Agents through a canary and batch rollout |
| Restarted Agents behave differently | The new image or generated configuration was applied | Compare canary logs, settings, pinned versions, and integrations |
| Agent PVC data is missing | The upgrade recreated or changed storage | Stop further restarts and follow the tested PVC recovery plan |
| UI and API disagree on a contract | Incompatible images were deployed | Deploy the matching API/UI pair |
| Grafana loses visibility | Monitoring or target discovery failed | Use Kubernetes and external monitoring while restoring the release |
| New image cannot be pulled | Tag, credentials, repository, or pull Secret is wrong | Verify the pinned image and registry hook before retrying |
| Staging creates Agent resources in production | K8S_NAMESPACE or API kubeconfig is wrong | Stop staging and correct namespace-scoped configuration |
Upgrade checklist
-
Before staging
-
Before production
-
After production
Next steps
Keep production configuration, Kubernetes deployment, migrations, self-hosting Communications, and monitoring available during the rollout. For operational follow-up, use Communication Connections, Communication diagnostics, Runtime deployment, and the Slack, Microsoft Teams, Telegram, and Discord setup guides.
Continue the self-hosting sequence Troubleshoot self-hosting → Diagnose deployment, database, runtime, provider, and monitoring failures.Roll out initiated delivery
Initiated delivery depends on compatible application schema/code, runtime-generated scripts, and runtime image hooks. Roll the application through its normal migration workflow, deploy the matching runtime images, and recreate affected Agent runtime configuration through the supported stop/start flow. Existing pods do not acquire new image hooks or generated scripts merely because a new application image was built. Preserve the Agent's persistent state when performing a normal restart.
Hermes session continuity
Agent Barn's Hermes adapter sends resume_session: true with a stable Connection/location/thread session identity. The shipped Hermes image patches its run endpoint to reload persisted conversation history, including compaction lineage and tool-call metadata. Deploy the compatible patched image and adapter together when adopting this behavior.
Existing session state remains on the Agent's persistent volume. New sessions can legitimately have no history. Do not recreate storage as a routine response to missing continuity, and do not assume updating the application automatically replaces every running Agent pod or image. Recreate the affected runtime through the normal Agent lifecycle after selecting the intended compatible runtime image.
Runtime images have their own version sources. A product release tag is not a substitute for the Hermes or OpenClaw image version. Preserve persistent state when rolling runtime configuration.
See Choose an Agent Runtime for transport and replacement boundaries.
Client image release
The manual Client Release workflow publishes customer API, UI, Hermes-base, and OpenClaw-base images without deploying a customer cluster. It is separate from k3s and public hosted deployments. See Client image release for registry settings, independent Runtime versions, and backend topology assumptions.