Skip to main content
Version: Next

On-Prem / Kubernetes Deployment

Deploy Agent Kernel's queue-execution pipeline to any Kubernetes cluster with the official Helm chart: an io-handler Deployment (REST API + Response Handler), an agent-runner Deployment (the consumers executing your agents), and an optional ws-gateway Deployment for WebSocket delivery. Backing services (Valkey, NATS JetStream) ship as condition-gated dependencies of their official charts, and deployment flavors (dev, baremetal, AWS EKS) are values files over one set of templates.

The chart lives at ak-deployment/ak-k8s and is published as an OCI artifact.

Topology​

This is the same five-component queue pipeline that runs in-process locally and over SQS on AWS, split into the two-process topology: the transport is configuration, not code.

Quick Start​

Build your application images (the examples/k8s/openai-queue-mode example walks this end to end on k3d, microk8s, and k3s), load them into your cluster, then:

helm pull oci://ghcr.io/yaalalabs/charts/agent-kernel --version 0.9.3 --untar # unpacks the flavor values files
helm install ak oci://ghcr.io/yaalalabs/charts/agent-kernel --version 0.9.3 \
-f agent-kernel/values-dev.yaml \
--set ioHandler.image.repository=<io image> \
--set agentRunner.image.repository=<runner image> --set image.tag=<tag>

kubectl port-forward service/ak-agent-kernel-io 8000:80
curl -s -X POST http://localhost:8000/api/v1/chat \
-H 'Content-Type: application/json' \
-d '{"prompt": "Hello", "session_id": "s1", "agent": "triage"}'

The chart is published to GHCR as an OCI artifact with every Agent Kernel release, versioned like the package, and the flavor values files (values-dev.yaml, values-baremetal.yaml, values-eks.yaml) ship inside it, which is what the helm pull --untar unpacks. To work from a repository checkout instead (unreleased chart changes), add the valkey and nats chart repos, run helm dependency build ak-deployment/ak-k8s/chart, and install from that path.

The dev flavor runs single replicas over in-cluster NATS and Valkey with the JetStream objects auto-provisioned at startup. The smallest install is the single-process profile documented in values-dev.yaml: one pod, in-process queues, no backing services at all.

The Application Image Contract​

The chart runs your images: your config.yaml (baked into the image) declares what runs, and the chart injects where it runs as AK_* environment variables, the same app/infra split the ECS Terraform deployment uses.

DeploymentEntry point
io-handlerIOHandler.run(), or IOHandler.run(handlers=[WebhookRESTRequestHandler(...)]) to serve messaging webhooks alongside the chat route
agent-runnerregisters your agent modules, then AgentRunner.run()
ws-gatewayWebSocketGateway.run(auth_validator=...)

The poller tier (Gmail)​

A polled integration has no webhook, so it runs as its own workload calling PollerRunner.run(GmailInboundAdapter()) — at one replica. It serves no HTTP, so it must not ride the io tier's CPU autoscaler: scaling the webhook tier for Slack load would otherwise multiply the poll rate for no reason, and the poller's already-handled record is per process.

The chart does not template this Deployment yet; run it as your own workload (a copy of the agent-runner Deployment with replicas: 1 and your poller entry point is enough). On the in_memory transport there is no separate container at all: pass IOHandler.run(pollers=[PollerRunner(adapter)]) and it runs as a peer thread.

Flavors​

Flavors never fork templates: every difference is a value.

Values filePosture
values-dev.yamlMicro-cluster (k3d, kind, microk8s, k3s): single replicas, auto-provisioned JetStream, TLS off, port-forward entry
values-baremetal.yamlEnvoy Gateway class, cert-manager issuer annotations, NACK-managed JetStream objects, OpenEBS hostpath storage, MetalLB prerequisite
values-eks.yamlAWS Load Balancer Controller gateway classes (ALB), ACM certificates, EBS gp3 storage, Pod Identity; sqs, kafka, and nats transports all valid

Prerequisites (Gateway API CRDs and an implementation, MetalLB, cert-manager, KEDA, the NACK controller, the Strimzi operator) are documented per flavor and deliberately never installed by the chart.

Transports​

transport.type selects the broker; the pipeline semantics (per-session ordering, bounded retry, dedup, permanent-failure replies) are identical over all of them:

  • nats (default, recommended on-prem): JetStream work-queue streams with one durable consumer per partition. Dev clusters auto-provision; production manages the objects declaratively through the chart's NACK CRs and fails loudly on a missing object.
  • kafka: pair with the Strimzi operator; the chart renders the cluster, node pool, and topics as CRs.
  • sqs: no broker to operate; the EKS option via Pod Identity.

WebSocket Modes (async / stream)​

Enabling wsGateway adds the gateway tier for async and stream execution modes: gateway pods own the client sockets, enqueue chat frames directly to the transport, and receive each reply or token chunk from the Response Handler on an authenticated internal push endpoint, addressed through a shared connection store on the session backend. Replies reach all of a user's connections on whichever gateway pod holds them, and io/runner pods roll without dropping a single connection. See the WebSocket delivery section of the Queue Mode Guide for the mechanism and the chart README for the values.

Sandbox Worker Tier​

Enabling sandboxWorker adds the sandbox broker worker: it consumes sandbox execution requests from the sandbox queues (same transport, its own queue names), runs them through a sandbox provider (typically kubernetes pods running as a ServiceAccount the chart binds to nothing, so the RBAC you grant it is the security boundary), and returns completions over the sandbox output queue into the shared response store. The tier ships with its own ServiceAccount and RBAC, KEDA scaling on the sandbox input backlog, and values-gated namespace hardening (Pod Security Admission, default-deny egress, quotas).

The tier also installs standalone: when the rest of Agent Kernel runs outside the cluster (Lambda mode, ECS), disable ioHandler and agentRunner and this chart deploys only the sandbox worker; the agent side and the worker then meet solely on the shared sandbox queues and response store. See the chart README's sandbox worker section for the values and the agent-side mirror configuration.

Autoscaling​

The agent-runner tier scales on queue depth via KEDA (Kafka consumer lag, NATS JetStream pending, or SQS queue length, selected by the transport), because LLM-bound work idles the CPU while requests back up. The io-handler tier scales on plain CPU. Runners drain gracefully on SIGTERM: consumers stop claiming work and finish in-flight turns within terminationGracePeriodSeconds.

Observability and Air-Gap​

Observability ships as documented recipes (kube-prometheus-stack, per-broker exporters, an OpenTelemetry Collector funnel for the Langfuse/OpenLLMetry/Logfire tracing providers), not as chart dependencies. Air-gapped installs set global.imageRegistry (application images and Valkey) plus global.image.registry (the NATS subchart) and mirror the per-release images.txt manifest. Both are covered in the chart README.

Next Steps​

Ready to Ship Your
First Agent?

Free, open-source, Apache 2.0. No licensing costs, no vendor lock-in. Join hundreds of developers building production AI agents with Agent Kernel.

Agent Kernel
Ask Agent Kernel