Self-hosting
Concept Available

Runtime Assembly and Deployment

Agent Barn renders an Agent's exact configuration and builds Kubernetes resources for Hermes or OpenClaw, keeping Runtime execution independent from Platform delivery, which the Communications Gateway supervises through shipped Platform Plugins. Deployment spans the API, web app, workers, and Kubernetes resources.

For
Runtime and platform engineers
On this page
  1. Protocol v2 transport and delivery leases
  2. Startup assembly
  3. Runtimes and Platforms are independent
  4. Runtime-neutral Communications
  5. Scheduled execution and chat delivery
  6. Conversations and Tool Calls have different writers
  7. When a diagnostic journal entry cannot be saved
  8. Choose the right troubleshooting path
  9. Deployed services and boundaries
  10. Deployment topology and releases
  11. Workers and monitoring
  12. Change checklist
01

Protocol v2 transport and delivery leases

Both runtimes consume Agent Barn's shared Communications protocol, but the adapter invokes them differently. Hermes uses the asynchronous /v1/runs API and its event/approval endpoints. OpenClaw uses /v1/chat/completions. Provider credentials remain inside Communications and are not sent to either runtime.

  1. 01.1

    Protocol version 2 uses an authenticated outbound Server-Sent Events control stream to notify the runtime adapter of durable work. The adapter claims an inbound Delivery, invokes the selected runtime, submits the reply against the source Delivery, and completes the Delivery.

  2. 01.2

    PostgreSQL owns delivery state, claims, leases, idempotency, and cancellation. Redis Streams provide content-free wakeups rather than authoritative message state. Reconnect replay and bounded fallback wakeups recover durable work when a signal is missed or Redis is unavailable. The adapter retains a five-second safety claim poll.

  3. 01.3

    Claims last 120 seconds. Hermes renews its live claim every 60 seconds while its asynchronous run or approval is active. OpenClaw retains the bounded-turn behavior of its chat-completions path. Do not infer that both runtimes have the same progress, approval, or abort transport merely because they share Communications.

  4. 01.4

    POST .../deliveries/{delivery_id}/renew?awaiting_input=true renews the live claim while reporting that its run is parked for a human answer. The flag defaults to false and is reported again on lease heartbeats. A waiting claim does not block the next inbound answer on the same ordering key. A newly claimed Delivery starts with the flag false. This is processing state on a claim, not a new terminal Delivery status, an authorization bypass, or a revival of an expired claim.

  5. 01.5

    Version-1 claim/reply/complete routes remain accepted while older Agents are rebuilt. Updating a Connection reconciles provider connectivity separately from restarting the runtime.

  6. 01.6

    Communications → shared adapter → Hermes /v1/runs; shared adapter → OpenClaw /v1/chat/completions. The SSE control stream carries work notifications, not token-by-token model responses.

02

Startup assembly

Starting an Agent is a Product API orchestration flow that renders the Agent's exact configuration and builds its Kubernetes resources.

  1. 02.1

    Communication Connection credentials are not decrypted into the Runtime, and provider tokens for Slack, Microsoft Teams, Telegram, or Discord are never materialized into Hermes or OpenClaw.

  2. 02.2

    A Runtime startup or Kubernetes failure can move an Agent to ERROR. A provider-session or Connection failure changes Connection health instead of Agent lifecycle.

  3. 02.3

    Agent Barn starts the Hermes boot runner alongside the gateway. When BOOT.md is non-empty, the runner waits for the local API to accept a /v1/runs submission. The run uses a reserved no-conversation session with session resume disabled and is intended for repeatable startup work. The runner can wait up to 300 seconds for readiness.

  4. 02.4

    The startup log confirms that the boot run was submitted. The runner does not wait for the checklist to finish, so that log is not proof that setup succeeded. Inspect the gateway log for the outcome. A comments-only but non-empty BOOT.md is still non-empty and is submitted.

  5. 02.5

    Loads the Organization-owned Agent and its pinned Template or Agent Template Override Version.

  6. 02.6

    Resolves and renders the Agent's current configuration.

  7. 02.7

    Decrypts Agent Secrets used by tool Integrations.

  8. 02.8

    Selects the Hermes or OpenClaw Runtime builder.

  9. 02.9

    Combines explicitly assigned Skill Versions with eligible built-in provider Skills.

  10. 02.10

    Materializes supported Integration configuration.

  11. 02.11

    Appends tool pointers and unconditional runtime behavior policies.

  12. 02.12

    Generates fresh Ingest and Communications protocol credentials.

  13. 02.13

    Builds the ConfigMap, Secret, PVC, Service, and Deployment resources, including the runtime-neutral Communications adapter.

  14. 02.14

    Applies the Kubernetes resources and transitions the Agent to running when startup succeeds.

03

Runtimes and Platforms are independent

Hermes and OpenClaw are Runtimes; Slack, Microsoft Teams, Telegram, and Discord are Platforms, and the two are independent. Platform support is supplied by shipped, code-owned Platform Plugins rather than by branches inside each Runtime builder, and both Runtimes consume the same versioned, runtime-neutral Communications protocol.

  1. 03.1

    Platform selection is not part of Agent creation. An Agent may be headless or own zero or many Communication Connections, including multiple Connections to the same Platform; see /guides/agents/communication-connections for Connection lifecycle. Adding a Platform should add one Platform Plugin and provider adapter rather than changes to both Runtime builders.

  2. 03.2

    Each shipped Platform Plugin owns:

  3. 03.3

    Typed Connection settings

  4. 03.4

    Typed credential schemas

  5. 03.5

    External credential validation

  6. 03.6

    Credential identity and uniqueness rules

  7. 03.7

    Provider ingress or supervised sessions

  8. 03.8

    Message normalization and admission policy

  9. 03.9

    Optional directory and name enrichment

  10. 03.10

    Outbound provider delivery

  11. 03.11

    Optional processing feedback

  12. 03.12

    Provider health observation

  13. 03.13

    Slack ingress: supervised Socket Mode

  14. 03.14

    Telegram ingress: supervised polling

  15. 03.15

    Discord ingress: supervised Gateway session

  16. 03.16

    Microsoft Teams ingress: authenticated provider webhook

04

Runtime-neutral Communications

A Runtime-side adapter claims a durable inbound Communication Delivery, invokes the local Hermes or OpenClaw endpoint, and submits the reply against the source delivery; the Communications Gateway then sends the reply through the same source Connection and Platform Plugin.

  1. 04.1

    The Runtime uses Connection-scoped session identity but does not receive provider credentials. Durable delivery, leases, ordering, idempotency, and provider retries belong to Communications, not to the Runtime.

  2. 04.2

    Communication Connection settings and credentials have an independent lifecycle: updating a Connection increments its revision, and the Communications Gateway reconciles the provider session without restarting a running Agent. Connection health and Agent lifecycle remain separate, and one failed Connection does not stop another Connection or automatically move the Agent to ERROR.

  3. 04.3

    Inbound: Communication Platform

  4. 04.4

    Inbound: Platform Plugin

  5. 04.5

    Inbound: Communications Gateway

  6. 04.6

    Inbound: Durable Communication Delivery

  7. 04.7

    Inbound: Runtime-neutral adapter

  8. 04.8

    Inbound: Hermes or OpenClaw

  9. 04.9

    Outbound: Hermes or OpenClaw

  10. 04.10

    Outbound: Runtime-neutral adapter

  11. 04.11

    Outbound: Durable Communication Delivery

  12. 04.12

    Outbound: Communications Gateway

  13. 04.13

    Outbound: Platform Plugin

  14. 04.14

    Outbound: Communication Platform

05

Scheduled execution and chat delivery

Agent Barn Communications supports ordinary replies and a separate initiated-message acceptance path. Interactive sends carry server-issued context for a live inbound execution and resolve only on that execution's Connection. Scheduled submissions use a recorded origin or the Agent's configured default. The model does not select a Connection ID. Only Slack currently advertises initiated delivery.

  1. 05.1

    Hermes scheduled jobs created from a conversation retain its Connection, channel, and thread. Startup-created jobs use the default. OpenClaw uses the default only when a completion has no recorded origin. Its pinned cron hook can expose a delivery-channel label instead of the creating conversation; an unmappable origin is refused instead of being sent to the default.

  2. 05.2

    Hermes captures scheduled final responses at a scheduler completion boundary in its image and suppresses native provider delivery. OpenClaw captures completions through its in-process agent_end hook. Both write eligible completions to a runtime-local SQLite spool before returning from the hook. A separate drain process submits them to Communications with a stable run identity.

  3. 05.3

    Communications resolves the destination and atomically persists a canonical outbound message and durable Delivery before acknowledging acceptance. The resolved destination is retained for retries. The Platform Plugin checks the Connection's current outbound policy again before provider delivery. Acceptance is not provider success.

  4. 05.4

    For a job created in a channel thread, an explicit delivery setting may change the thread within that same channel. To post at the channel root, use <platform>:<channel-id>: with the final thread segment empty. To target a particular thread, include its ID in that segment. Naming a different channel is refused; it does not redirect the result. This syntax belongs to the scheduled job's delivery configuration, not its prompt or the default-target form.

  5. 05.5

    The local spool retries a scheduled submission under the same run identity until it is acknowledged, receives a permanent client-error refusal, or exhausts its bounded retry budget, approximately one day of transient failures. The implementation uses an attempt limit, not an exact 24-hour expiration. Settled rows are retained for seven days before pruning. Accepted work then follows the server's durable Delivery lifecycle; local submission retries and server provider-delivery retries are different stages.

  6. 05.6

    An unmappable origin is refused before a deliverable completion is captured. A permanent refusal must be diagnosed; it is not automatically replayed when configuration later changes. Changing a default or enabling a Connection does not recover every previously refused completion.

  7. 05.7

    Return a recognized silence marker when there is nothing worth sending: [SILENT], SILENT, NO_REPLY, NO REPLY, or HEARTBEAT_OK. Matching is case-insensitive and ignores surrounding whitespace. An empty response is also suppressed.

  8. 05.8

    The current filter checks the whole response and its first and last non-empty lines. A standalone marker on either of those lines suppresses the entire result, even if other lines contain a substantive report. Do not append a silence marker to a report you want delivered. Text such as Nothing to flag today. is not a silence marker and remains eligible for delivery.

  9. 05.9

    Eligible describes the silence filter only; routing, admission, and provider delivery can still refuse or fail.

  10. 05.10

    NO_REPLY: suppressed.

  11. 05.11

    A report followed by a final standalone NO_REPLY line: entire response suppressed.

  12. 05.12

    A first standalone NO_REPLY line followed by a report: entire response suppressed.

  13. 05.13

    Nothing to flag today.: eligible for delivery.

  14. 05.14

    Report [SILENT] on one line: eligible for delivery.

06

Conversations and Tool Calls have different writers

Conversations and Tool Calls appear together in an Agent's Activity view, but they arrive through different services.

  1. 06.1

    The Communications Gateway persists canonical inbound and outbound Conversation Messages; the Ingest API persists authenticated Runtime Tool Call telemetry. These are separate write paths, not a single stream from the Runtime through Ingest.

  2. 06.2

    The Product API provides authorized reads for the Activity interface. Ingest is not the source of canonical Conversation Messages.

  3. 06.3

    Each conversation location belongs to a particular Communication Connection, and its identity includes both the Connection ID and the provider channel ID. Two Connections can use the same provider channel identifier without combining their histories.

  4. 06.4

    Tool Call identity is scoped by Agent and Runtime invocation identity.

  5. 06.5

    Costs read stored cost records synchronized from LiteLLM spend logs with OpenRouter missing-charge recovery; they are not calculated from Conversations or Tool Calls.

  6. 06.6

    See /guides/activity-conversations-and-telemetry for the complete Conversation, Tool Call, and cost-attribution model.

  7. 06.7

    Incoming chat message: the provider sends it to Communications, which records the accepted message and prepares it for the Agent.

  8. 06.8

    Agent reply: the Runtime submits a reply through Communications, bound to the originating Communication Delivery and its Connection; Communications records the reply and sends it to the provider.

  9. 06.9

    Tool Call or tool result: the Runtime sends telemetry to Ingest, which records the Tool Call and updates its result.

07

When a diagnostic journal entry cannot be saved

Communications records provider observations and admission-policy decisions to help explain what happened to a message.

  1. 07.1

    These diagnostic writes are best-effort. If one fails, Agent Barn logs the failure and continues handling otherwise valid provider traffic. A diagnostic storage problem should not cause an already-consumed provider message to be lost.

  2. 07.2

    This behavior applies to observation and policy-admission journal entries. It does not make message storage or durable Communication Delivery records optional, and it does not bypass a policy decision that rejects a message.

  3. 07.3

    See /guides/observe-and-govern/communication-diagnostics for Connection status, delivery history, and recovery actions.

08

Choose the right troubleshooting path

These starting points route a symptom to the service that owns it. They are not proof of a particular failure.

  1. 08.1

    For example, a chat reply can fail because the Runtime could not complete its work, even when the provider Connection is healthy.

  2. 08.2

    Chat messages or replies are missing: start with the affected Connection and its Communications delivery diagnostics.

  3. 08.3

    Conversations work but Tool Calls are missing: start with the Runtime telemetry path to Ingest.

  4. 08.4

    The Agent Runtime does not become ready: start with Runtime health and logs.

  5. 08.5

    Spend information is missing or unexpected: start with stored cost attribution, synchronization freshness, and the selected reporting period.

09

Deployed services and boundaries

A complete deployment runs more than the Product API, Ingest, worker, and UI.

  1. 09.1

    Product routes are mounted under /api/v1, Runtime Tool Call telemetry under /ingest/v1, and the Runtime Communications protocol and provider ingress under /communications/v1. The Communications service is deployed separately from the Product API process and supervises provider ingress and outbound delivery.

  2. 09.2

    Only the provider-webhook prefix requires public Communications ingress; Runtime protocol traffic remains internal. Communications durable delivery is separate from Domain Event delivery.

  3. 09.3

    Product API on port 8000

  4. 09.4

    Ingest API on port 8001

  5. 09.5

    Communications service on port 8002

  6. 09.6

    Web app, PostgreSQL, Redis, and LiteLLM

  7. 09.7

    Domain Event delivery worker and reconciliation CronJob

  8. 09.8

    Monitoring stack

  9. 09.9

    Hermes and OpenClaw Runtime images

  10. 09.10

    Kubernetes resources for running Agents

10

Deployment topology and releases

The existing k3s main and staging deployments are testing grounds, not hosted public production. Their API and UI images use moving environment tags such as latest and latest-staging.

  1. 10.1

    Hosted public production runs on the separate Talos cluster and deploys only from a vX.Y.Z Git tag, or a manual dispatch of an existing release tag; public API and UI images are pinned to that release tag in registry.agentbarn.dev. A public deployment and the k3s testing deployments are separate workflows.

  2. 10.2

    See /guides/self-hosting/deploy-kubernetes and /guides/self-hosting/upgrades for the operational deployment and upgrade procedures.

11

Workers and monitoring

The API image also runs Dramatiq Domain Event delivery workers and a scheduled reconciliation CronJob; see /guides/domain-events-and-delivery for that model.

  1. 11.1

    The default namespace-scoped Prometheus configuration collects Product API metrics on port 8000, Ingest metrics on port 8001, LiteLLM metrics, Agent Runtime health, and Kubernetes workload state. Alertmanager routes configured alerts to Slack, and Grafana provides dashboards.

  2. 11.2

    Communications exposes metrics on port 8002, but the default monitoring configuration does not collect them. Add a Prometheus scrape target for the Communications Service to collect these metrics. Keep the endpoint internal; see /guides/self-hosting/communications for the collection boundary and configuration guidance.

  3. 11.3

    Communications exposes low-cardinality metrics for Connection status, delivery outcomes, queue depth, oldest queued delivery age, delivery latency, reconnect requests, and policy dispositions.

12

Change checklist

Check both Runtimes, every shipped Platform Plugin, Runtime builders, base images, lifecycle tests, Communications and Ingest contracts, charts, and environment-derived tags. Chart versions change only when chart templates or values change.

  1. 12.1

    See /guides/agents/choose-runtime for Runtime selection guidance and /guides/platforms/compatibility for current Platform behavior.

Documentation