Operational outcome
A separately operated Communications process with private Runtime delivery and narrowly public provider ingress
Use this guide to deploy, scale, monitor, and troubleshoot the service that owns Communication Connections and provider message delivery.
- Provider sessions and webhooks reach Communications on port
8002. - Agent Runtimes use the internal versioned Communications protocol.
- Only the provider-webhook prefix is exposed outside the cluster.
Communications is a separate process
Agent Barn has three HTTP composition roots. Communications is separately served from the same API image as the Product and Ingest applications, using api.communications_main:app.
| Process | Prefix | Default port | Responsibility |
|---|---|---|---|
| Product API | /api/v1 | 8000 | Human-facing Organization, Agent, Connection, and diagnostics operations |
| Ingest API | /ingest/v1 | 8001 | Authenticated Runtime Tool Call telemetry |
| Communications | /communications/v1 | 8002 | Provider webhooks, provider sessions, durable Delivery processing, and Runtime-neutral communication |
Communications owns supervised provider ingress, Slack Socket Mode sessions, Telegram getUpdates polling, Discord Gateway sessions, Microsoft Teams webhook ingress, Platform Plugin normalization and policy admission, durable inbound and outbound Communication Deliveries, Runtime claim/reply/completion routes, outbound provider delivery, Connection health transitions, content-free operational journal writes and pruning, and Communications metrics.
Traffic and ownership boundaries
- ProviderUses a supervised session or polling loop, or sends an authenticated public webhook.
- Communications :8002Platform Plugins normalize events; the ingress supervisor, durable Deliveries, outbound processor, journal, and metrics run here.
- Internal versioned protocolOnly Agent Runtimes claim Deliveries, write replies, and complete processing.
- Agent RuntimesProvider credentials are never sent to the Runtime.
| Product API :8000 | Human-facing responsibility |
|---|---|
| Connection CRUD | Create and manage Connection configuration. |
| Diagnostics | Read Connection health and operational journal data. |
| Reconnect and retry | Request controlled recovery actions after reviewing diagnostics. |
Run Communications locally
Start the development API processes with:
make dev-apiThis starts Product API on 8000, Ingest on 8001, and Communications on 8002. To run only Communications:
make dev-communicationsFor the complete local stack, including the separately served Communications process, use:
./run.shSee Local development and operations for the surrounding development workflow.
Configure the internal Runtime URL
Set COMMUNICATIONS_BASE_URL to the internal Communications base URL, including /communications/v1. The Product API uses this value when building Runtime resources, and every Agent Runtime workload must resolve and reach it.
For local development, the default shape is:
http://host.docker.internal:8002/communications/v1In Kubernetes, the Product API chart configures a Service URL equivalent to:
http://<release-name>-communications:8002/communications/v1Runtime and Platform Driver protocols
Runtime communication requires a per-Agent Communications bearer credential generated during Agent start, the Agent ID in the route, and X-AgentBarn-Communications-Version: 1.
POST /communications/v1/agents/{agent_id}/deliveries/claim
POST /communications/v1/agents/{agent_id}/deliveries/{delivery_id}/replies
POST /communications/v1/agents/{agent_id}/deliveries/{delivery_id}/completeUnsupported protocol versions return 426 Upgrade Required; invalid Runtime credentials return 401 Unauthorized. Keep these routes internal. The Ingest bearer key is not interchangeable with the Communications credential, and provider credentials never reach the Runtime.
Platform Driver events use a separate internal route and require a Connection driver bearer credential with X-AgentBarn-Driver-Version: 1:
POST /communications/v1/connections/{connection_id}/eventsExpose only provider webhook ingress
POST /communications/v1/webhooks/{connection_id}Microsoft Teams currently uses this capability. Each webhook is scoped to a Communication Connection; the Platform Plugin authenticates the provider request. Microsoft Teams verifies the Bot Framework bearer token against the Connection’s App ID and activity service URL. A webhook URL contains no provider credential.
In Kubernetes, publicly route only:
/communications/v1/webhooksDo not expose all of /communications/v1. In particular, do not use the retired /api/v1/webhooks/teams/{agent_id}/messages shape.
Public URL configuration
API_EXTERNAL_URL must be the public Agent Barn origin used to construct Connection webhook URLs:
https://<public-host>/communications/v1/webhooks/<connection-id>- The hostname resolves publicly and TLS is valid.
- Ingress routes the webhook prefix to the Communications Service.
- The Product API, Runtime protocol, metrics, and health endpoints do not need public exposure.
Deploy Communications in Kubernetes
The API Helm chart renders a separate Communications Deployment and ClusterIP Service. Communications is not a sidecar of the Product API.
communications:
enabled: true
replicaCount: 1
service:
port: 8002
resources:
requests:
memory: 256Mi
cpu: 100m
limits:
memory: 512Mi
cpu: 500mThe container runs uvicorn api.communications_main:app --host 0.0.0.0 --port 8002 and uses the component label app.kubernetes.io/component: communications. See Deploy Kubernetes for deployment-wide requirements.
Health, supervision, and reconciliation
The process-level health endpoint is GET /health on port 8002:
{
"status": "ok"
}The Helm chart uses this endpoint for both readiness and liveness probes. It does not prove that every provider Connection is healthy or that end-to-end Delivery processing succeeds; use Communication diagnostics and Communications metrics for those checks.
During its lifespan, Communications starts a Platform ingress supervisor and outbound Delivery worker. The supervisor finds enabled Communication Connections, starts provider-specific sessions where required, reconciles Connection revision changes, stops sessions for disabled or retired Connections, records observed health independently from Agent lifecycle, retries setup or session failures, and prunes expired journal rows. Webhook-based Connections do not require the same persistent provider-session loop.
Scale without competing provider sessions
Supervised provider ingress uses PostgreSQL-backed leases scoped to a Communication Connection. Only one active supervisor owns a Connection’s provider ingress at a time; multiple Communications replicas coordinate through the database. If a replica is lost, another can acquire the expired lease. A Connection revision makes the owning supervisor cancel and recreate the provider session.
Communication Deliveries are durable in PostgreSQL. Outbound Delivery claims preserve ordering within a Conversation, and retries reuse a stable provider idempotency key. Scaling must not create competing provider sessions for the same Connection.
Use the shared database and encryption configuration
Communications requires the same PostgreSQL database as Product API. It reads and writes Communication Connections, encrypted credential envelopes, Communication Deliveries, canonical Conversation Messages, operational journal entries, and ingress leases.
AGENT_TOKEN_ENCRYPTION_KEYmust match the Product API configuration.- Connection credentials remain encrypted and are decrypted only inside the Communications and validation boundaries that need them.
- Provider credentials are never exposed through read APIs or Runtime configuration.
Relevant configuration and journal retention
| Setting | Purpose |
|---|---|
COMMUNICATIONS_BASE_URL | Internal Runtime-facing Communications base URL |
API_EXTERNAL_URL | Public origin used to construct provider webhook URLs |
AGENT_TOKEN_ENCRYPTION_KEY | Encrypts and decrypts Agent and Connection credential material |
COMMUNICATION_JOURNAL_RETENTION_DAYS | Retention for the content-free operational journal |
SLACK_DIRECTORY_CACHE_TTL_SECONDS | Cache duration for Slack directory discovery |
SLACK_REQUEST_TIMEOUT_SECONDS | Timeout for Slack provider API requests |
TEAMS_PUBLISHER_NAME | Publisher name used in generated Teams app packages |
TEAMS_PUBLISHER_WEBSITE_URL | Public HTTPS publisher website |
TEAMS_PRIVACY_URL | Public HTTPS privacy URL |
TEAMS_TERMS_URL | Public HTTPS terms URL |
The provider-validation bypasses SKIP_SLACK_TOKEN_VALIDATION, SKIP_TELEGRAM_TOKEN_VALIDATION, SKIP_DISCORD_TOKEN_VALIDATION, and SKIP_TEAMS_TOKEN_VALIDATION are development or test controls only. Do not enable them in production.
COMMUNICATION_JOURNAL_RETENTION_DAYS controls content-free operational journal retention. Its default is 31 days; valid values range from 1 through 3650. The Communications supervisor runs the pruning sweep, and changing retention is an operational configuration change.
Collect Communications metrics
The Communications service exposes operational metrics at /metrics on port 8002. In the standard deployment, the internal address is:
http://agentbarn-api-communications:8002/metricsIf you changed the API Helm release name, replace agentbarn-api-communications with <your-api-release-name>-communications.
The default Agent Barn monitoring configuration does not collect metrics from this endpoint. To collect them, configure your Prometheus installation to scrape the Communications Service.
Keep the endpoint internal to your cluster. You do not need to publish it through an internet-facing ingress.
Until a scrape target is configured, missing Communications metrics in Prometheus or Grafana do not by themselves indicate that the service has failed. You can still inspect individual connections and deliveries in Agent Barn through Communication Diagnostics.
See Monitor the platform for the monitoring configuration and its current coverage. Dashboards and alert rules for Communications also need to be configured separately.
The endpoint exposes the following metric families:
agentbarn_communication_connection_statusagentbarn_communication_delivery_outcomesagentbarn_communication_queue_depthagentbarn_communication_oldest_queued_age_secondsagentbarn_communication_delivery_latency_secondsagentbarn_communication_reconnectsagentbarn_communication_policy_dispositions
Metric labels are deliberately low-cardinality: status, direction, outcome, and admission disposition. Communications metrics do not use Organization, Agent, Connection, Conversation, or User IDs as labels.
Operational checklist
- Communications Deployment has ready replicas and its ClusterIP Service exposes
8002. /healthresponds internally.- Product API uses the correct
COMMUNICATIONS_BASE_URL, and Agent Runtime pods can resolve and reach the Service. - Only
/communications/v1/webhooksis publicly routed, andAPI_EXTERNAL_URLproduces a valid public webhook URL. - Provider Connections show fresh observed health; queue depth and oldest queued age remain within expected ranges.
- Journal retention matches operational requirements. If you want Communications metrics collected, a Prometheus scrape target for the Communications Service has been configured.
- Provider-validation bypasses are disabled in production.
Troubleshooting boundaries
| Symptom | Check |
|---|---|
| Runtime cannot claim Deliveries | Internal DNS, Service port, COMMUNICATIONS_BASE_URL, Runtime credential, and protocol version |
| Teams webhook fails | Public DNS and TLS, ingress prefix, Connection ID, and Bot Framework authentication |
| Slack, Telegram, or Discord disconnected | Connection credentials, enabled state, observed health, supervisor logs, and ingress lease |
| Duplicate provider consumers | Replica lease behavior and competing external pollers or sessions |
| Replies remain queued | Runtime completion, outbound worker, provider errors, and queue metrics |
| Journal is empty | Connection activity, diagnostics window, database access, and retention |
| /health passes but messages fail | Connection diagnostics and pipeline metrics |
| Provider session changed but Runtime did not restart | This is expected: Connection reconciliation is independent of the Agent Runtime |
Agent Email
Inbound-secret rotation must update the Product API service and the inbound Worker for the same environment. The current single-secret contract has an interruption window; a secret change alone does not select the path-filtered Worker publication. See Configure Agent Email for the coordinated deployment procedure.
Agent Email needs Cloudflare Email Routing, an inbound Worker, sending access, and matching environment-specific inbound secrets. Transactional email alone does not enable it. Follow Configure Agent Email for setup and routing-rule ownership.