Automation Platform > Deployment & hosting
Self-hosting troubleshooting
# Self-hosting troubleshooting Use these checks when the `oz-agent-worker` daemon won't start or connect, tasks stay queued, or tasks fail. :::note The steps below apply to the [managed architecture](/platform/self-hosting/#managed-architecture) (`oz-agent-worker` daemon). For [unmanaged](/platform/self-hosting/unmanaged/) deployments, refer to the documentation for the environment running `oz agent run` (e.g., GitHub Actions, Kubernetes). ::: --- ## Worker won't start ### Docker backend **Cause:** Docker isn't running, or the daemon platform isn't supported. **Fix:** 1. Verify Docker is running: `docker info`. 2. Confirm the daemon platform is `linux/amd64` or `linux/arm64`. Windows containers are not supported. 3. If the worker runs inside Docker, confirm the `/var/run/docker.sock` mount is correct and the mounting user has permission to the socket. ### Kubernetes backend **Cause:** The worker Deployment couldn't start, reach the Kubernetes API, or create its preflight Job. **Fix:** 1. Run `kubectl describe pod -n NAMESPACE WORKER_POD`. Replace `NAMESPACE` with the chart namespace and `WORKER_POD` with the worker pod name. For `CreateContainerConfigError`, verify the Secret and key configured by `warp.apiKeySecret`. 2. Check the worker logs for Kubernetes API or preflight diagnostics: `kubectl logs -n NAMESPACE WORKER_POD`. 3. Confirm the worker's namespace has these permissions: `create`, `get`, `list`, `watch`, `delete` on `jobs`; `get`, `list`, `watch` on `pods`; `get` on `pods/log`; `list` on `events`. 4. Confirm the task namespace allows pods with a root init container, unless you enabled native image volumes with `kubernetesBackend.useImageVolumes=true`. 5. If your cluster restricts image sources, set `kubernetesBackend.preflightImage` to an allowlisted image. The default is `busybox:1.36`. 6. To pull the preflight image from a private registry, configure `imagePullSecrets` in `kubernetesBackend.podTemplate`. A successful preflight validates the configured pod shape, not task-specific images, Secrets, setup commands, or network access. ### Direct backend **Cause:** The `oz` CLI isn't installed or isn't on the worker's `PATH`. **Fix:** 1. Install the Oz CLI on the worker host. See [Installing the CLI](/reference/cli/#installing-the-cli). 2. If the CLI isn't on `PATH`, set `oz_path` in the config file to the absolute path of the `oz` binary. --- ## Worker won't connect **Cause:** The API key is invalid, expired, or the host cannot reach the Automation Platform's backend. **Fix:** 1. Confirm your API key is correct, not expired, and has team scope. 2. Regenerate the API key in **Settings** > **Cloud platform** > **API keys** if you suspect it's invalid. 3. Ensure the host has outbound internet access to `oz.warp.dev:443`. 4. Check that no firewall rules are blocking WebSocket connections to `wss://oz.warp.dev`. 5. Increase log verbosity with `--log-level debug` to see connection details. See [Security and networking](/platform/self-hosting/security-and-networking/#network-requirements) for the full list of outbound endpoints the worker needs. --- ## Tasks not being picked up **Cause:** The worker isn't running, the `--host` value doesn't match the worker's `--worker-id`, or the worker and task belong to different teams. **Fix:** 1. Confirm the worker is running and connected. Check the worker logs for `Successfully connected to server`. 2. Verify the `--host` (or `worker_host`) value you passed matches your `--worker-id` exactly. Case-sensitive. 3. Ensure the worker's team matches the team creating the task. --- ## Metrics not appearing **Cause:** The worker is running, but metrics do not reach Prometheus or your collector. **Fix:** 1. For Prometheus, confirm `OTEL_METRICS_EXPORTER=prometheus`, bind to `0.0.0.0` in Docker or Kubernetes, and run `curl -s localhost:9464/metrics`. 2. Confirm no firewall or network policy blocks the metrics port. 3. For OTLP, confirm the collector endpoint is reachable and the protocol is correct. 4. For Helm, confirm `metrics.enabled=true` and check the Kubernetes Service or PodMonitor. A PodMonitor requires the Prometheus Operator CRDs. 5. Restart the worker with `--log-level debug` and check for metrics errors. See [Monitoring](/platform/self-hosting/monitoring/) for the full setup guide. --- ## Task failures **Cause:** The task environment, resources, or dependencies failed. **Fix (all backends):** 1. Review task logs in the <a href=https://oz.warp.dev>cloud agent dashboard</a> or via [session sharing](/agents/local-agents/session-sharing/). 2. Use `--no-cleanup` to keep the container, Job, or workspace indefinitely for inspection. With Kubernetes cleanup enabled, failed Jobs remain for 24 hours by default. 3. Use `--log-level debug` to see detailed execution logs. ### Docker backend (task failures) 1. Verify Docker is running (`docker info`). 2. If using a custom image, confirm it is **glibc-based** (not Alpine/musl) and that its architecture matches the worker's Docker daemon platform. ### Kubernetes backend (task failures) The Helm chart's `worker.resources` configures the worker Deployment, not task Jobs. Task containers have no worker-defined CPU or memory defaults. Configure the `task` container in `kubernetesBackend.podTemplate` or use a workload-specific [runner instance shape](/platform/self-hosting/managed-kubernetes/#pod-template). There is no universal task size; choose values based on the workload's measured needs. Start with the Pod state and events: ```bash kubectl get jobs,pods -n NAMESPACE kubectl describe pod -n NAMESPACE TASK_POD ``` Replace `NAMESPACE` with the task namespace and `TASK_POD` with the task pod name. #### Pod stays `Pending` with `FailedScheduling` **Verify:** In `kubectl describe pod`, look for `PodScheduled=False`, `FailedScheduling`, `Insufficient cpu`, or `Insufficient memory`. Check the Pod's requests and scheduling constraints against eligible node capacity. **Fix:** Right-size requests, lower `worker.maxConcurrentTasks`, or add compatible capacity. Configure autoscaling for nodes that satisfy the Pod's selectors, affinity, and tolerations. Raising only a limit does not make a Pending Pod schedulable. #### Running container terminates with `OOMKilled` **Verify:** Confirm the `task` container's termination reason is `OOMKilled`. Compare peak memory use with its limit, then check Pod events and node pressure. **Fix:** Reduce peak memory use or increase task memory in the runner instance shape or `pod_template`. An instance shape sets the memory request to the same value as the limit, so confirm an eligible node can schedule the larger request. #### Pod is `Evicted` **Verify:** Read the Pod reason and events for memory, disk, or node pressure. **Fix:** Restore node headroom, reduce concurrency, or add capacity before rerunning the task. For voluntary disruption, see [Protect active task pods from disruption](/platform/self-hosting/managed-kubernetes/#protect-active-task-pods-from-disruption). #### Task exits with code `143` Exit code `143` generally indicates `SIGTERM`; it does not prove an out-of-memory failure. **Verify:** Check the container termination reason and Pod events for eviction, preemption, node drain, the Job deadline, or manual deletion. **Fix:** Address the recorded cause. Treat the failure as OOM only when Kubernetes reports `OOMKilled`. #### Other Kubernetes failures * **Image pull errors** - Follow [Image pull failures](#image-pull-failures). Preflight does not pull every task image. * **`CreateContainerConfigError` or `FailedMount`** - Pod events name the missing Secret, ConfigMap, service account, volume, or key. * **Init container failure** - Check each init container's status and logs. The Warp sidecar-loading init container runs as root unless native image volumes are enabled. Custom init containers must finish before `task` starts. * **Network failure** - Test DNS, TLS, and the destination from a task Pod, not the worker Pod. * **Missing task credentials** - Provide repository, registry, and application credentials through your approved Secret integration. The worker API key is not a task credential. ### Direct backend (task failures) 1. Verify the Oz CLI is accessible. 2. Verify the workspace root directory has write permissions for the user running the worker. --- ## Image pull failures ### Docker backend (image pull) 1. If using a private registry, ensure Docker credentials are available to the worker. See [Private Docker registries](/platform/self-hosting/managed-docker/#private-docker-registries). 2. Try pulling the image manually on the worker host: `docker pull <image>`. ### Kubernetes backend (image pull) 1. Configure `imagePullSecrets` in the `pod_template` section of your worker config. 2. Verify the Secret exists in the task namespace and contains valid credentials. ### Both backends (image pull) * Verify the image exists and the tag is correct. * Check network connectivity from the worker/cluster to the registry. --- ## Related pages * [Self-hosting overview](/platform/self-hosting/) — Architecture and decision guide. * [Self-hosted worker reference](/platform/self-hosting/reference/) — CLI flags and config schema, including every flag mentioned here. * [Security and networking](/platform/self-hosting/security-and-networking/) — Outbound endpoints the worker needs. * [Agent Session Sharing](/agents/local-agents/session-sharing/) — Attach to running tasks to debug interactively.Walk me through resolving this issue: https://docs.warp.dev/platform/self-hosting/troubleshooting/Diagnose and fix common problems with self-hosted Automation Platform worker daemons across Docker, Kubernetes, and Direct backends.
Use these checks when the oz-agent-worker daemon won’t start or connect, tasks stay queued, or tasks fail.
Worker won’t start
Section titled “Worker won’t start”Docker backend
Section titled “Docker backend”Cause: Docker isn’t running, or the daemon platform isn’t supported.
Fix:
- Verify Docker is running:
docker info. - Confirm the daemon platform is
linux/amd64orlinux/arm64. Windows containers are not supported. - If the worker runs inside Docker, confirm the
/var/run/docker.sockmount is correct and the mounting user has permission to the socket.
Kubernetes backend
Section titled “Kubernetes backend”Cause: The worker Deployment couldn’t start, reach the Kubernetes API, or create its preflight Job.
Fix:
- Run
kubectl describe pod -n NAMESPACE WORKER_POD. ReplaceNAMESPACEwith the chart namespace andWORKER_PODwith the worker pod name. ForCreateContainerConfigError, verify the Secret and key configured bywarp.apiKeySecret. - Check the worker logs for Kubernetes API or preflight diagnostics:
kubectl logs -n NAMESPACE WORKER_POD. - Confirm the worker’s namespace has these permissions:
create,get,list,watch,deleteonjobs;get,list,watchonpods;getonpods/log;listonevents. - Confirm the task namespace allows pods with a root init container, unless you enabled native image volumes with
kubernetesBackend.useImageVolumes=true. - If your cluster restricts image sources, set
kubernetesBackend.preflightImageto an allowlisted image. The default isbusybox:1.36. - To pull the preflight image from a private registry, configure
imagePullSecretsinkubernetesBackend.podTemplate.
A successful preflight validates the configured pod shape, not task-specific images, Secrets, setup commands, or network access.
Direct backend
Section titled “Direct backend”Cause: The oz CLI isn’t installed or isn’t on the worker’s PATH.
Fix:
- Install the Oz CLI on the worker host. See Installing the CLI.
- If the CLI isn’t on
PATH, setoz_pathin the config file to the absolute path of theozbinary.
Worker won’t connect
Section titled “Worker won’t connect”Cause: The API key is invalid, expired, or the host cannot reach the Automation Platform‘s backend.
Fix:
- Confirm your API key is correct, not expired, and has team scope.
- Regenerate the API key in Settings > Cloud platform > API keys if you suspect it’s invalid.
- Ensure the host has outbound internet access to
oz.warp.dev:443. - Check that no firewall rules are blocking WebSocket connections to
wss://oz.warp.dev. - Increase log verbosity with
--log-level debugto see connection details.
See Security and networking for the full list of outbound endpoints the worker needs.
Tasks not being picked up
Section titled “Tasks not being picked up”Cause: The worker isn’t running, the --host value doesn’t match the worker’s --worker-id, or the worker and task belong to different teams.
Fix:
- Confirm the worker is running and connected. Check the worker logs for
Successfully connected to server. - Verify the
--host(orworker_host) value you passed matches your--worker-idexactly. Case-sensitive. - Ensure the worker’s team matches the team creating the task.
Metrics not appearing
Section titled “Metrics not appearing”Cause: The worker is running, but metrics do not reach Prometheus or your collector.
Fix:
- For Prometheus, confirm
OTEL_METRICS_EXPORTER=prometheus, bind to0.0.0.0in Docker or Kubernetes, and runcurl -s localhost:9464/metrics. - Confirm no firewall or network policy blocks the metrics port.
- For OTLP, confirm the collector endpoint is reachable and the protocol is correct.
- For Helm, confirm
metrics.enabled=trueand check the Kubernetes Service or PodMonitor. A PodMonitor requires the Prometheus Operator CRDs. - Restart the worker with
--log-level debugand check for metrics errors.
See Monitoring for the full setup guide.
Task failures
Section titled “Task failures”Cause: The task environment, resources, or dependencies failed.
Fix (all backends):
- Review task logs in the cloud agent dashboard or via session sharing.
- Use
--no-cleanupto keep the container, Job, or workspace indefinitely for inspection. With Kubernetes cleanup enabled, failed Jobs remain for 24 hours by default. - Use
--log-level debugto see detailed execution logs.
Docker backend (task failures)
Section titled “Docker backend (task failures)”- Verify Docker is running (
docker info). - If using a custom image, confirm it is glibc-based (not Alpine/musl) and that its architecture matches the worker’s Docker daemon platform.
Kubernetes backend (task failures)
Section titled “Kubernetes backend (task failures)”The Helm chart’s worker.resources configures the worker Deployment, not task Jobs. Task containers have no worker-defined CPU or memory defaults. Configure the task container in kubernetesBackend.podTemplate or use a workload-specific runner instance shape. There is no universal task size; choose values based on the workload’s measured needs.
Start with the Pod state and events:
kubectl get jobs,pods -n NAMESPACEkubectl describe pod -n NAMESPACE TASK_PODReplace NAMESPACE with the task namespace and TASK_POD with the task pod name.
Pod stays Pending with FailedScheduling
Section titled “Pod stays Pending with FailedScheduling”Verify: In kubectl describe pod, look for PodScheduled=False, FailedScheduling, Insufficient cpu, or Insufficient memory. Check the Pod’s requests and scheduling constraints against eligible node capacity.
Fix: Right-size requests, lower worker.maxConcurrentTasks, or add compatible capacity. Configure autoscaling for nodes that satisfy the Pod’s selectors, affinity, and tolerations. Raising only a limit does not make a Pending Pod schedulable.
Running container terminates with OOMKilled
Section titled “Running container terminates with OOMKilled”Verify: Confirm the task container’s termination reason is OOMKilled. Compare peak memory use with its limit, then check Pod events and node pressure.
Fix: Reduce peak memory use or increase task memory in the runner instance shape or pod_template. An instance shape sets the memory request to the same value as the limit, so confirm an eligible node can schedule the larger request.
Pod is Evicted
Section titled “Pod is Evicted”Verify: Read the Pod reason and events for memory, disk, or node pressure.
Fix: Restore node headroom, reduce concurrency, or add capacity before rerunning the task. For voluntary disruption, see Protect active task pods from disruption.
Task exits with code 143
Section titled “Task exits with code 143”Exit code 143 generally indicates SIGTERM; it does not prove an out-of-memory failure.
Verify: Check the container termination reason and Pod events for eviction, preemption, node drain, the Job deadline, or manual deletion.
Fix: Address the recorded cause. Treat the failure as OOM only when Kubernetes reports OOMKilled.
Other Kubernetes failures
Section titled “Other Kubernetes failures”- Image pull errors - Follow Image pull failures. Preflight does not pull every task image.
CreateContainerConfigErrororFailedMount- Pod events name the missing Secret, ConfigMap, service account, volume, or key.- Init container failure - Check each init container’s status and logs. The Warp sidecar-loading init container runs as root unless native image volumes are enabled. Custom init containers must finish before
taskstarts. - Network failure - Test DNS, TLS, and the destination from a task Pod, not the worker Pod.
- Missing task credentials - Provide repository, registry, and application credentials through your approved Secret integration. The worker API key is not a task credential.
Direct backend (task failures)
Section titled “Direct backend (task failures)”- Verify the Oz CLI is accessible.
- Verify the workspace root directory has write permissions for the user running the worker.
Image pull failures
Section titled “Image pull failures”Docker backend (image pull)
Section titled “Docker backend (image pull)”- If using a private registry, ensure Docker credentials are available to the worker. See Private Docker registries.
- Try pulling the image manually on the worker host:
docker pull <image>.
Kubernetes backend (image pull)
Section titled “Kubernetes backend (image pull)”- Configure
imagePullSecretsin thepod_templatesection of your worker config. - Verify the Secret exists in the task namespace and contains valid credentials.
Both backends (image pull)
Section titled “Both backends (image pull)”- Verify the image exists and the tag is correct.
- Check network connectivity from the worker/cluster to the registry.
Related pages
Section titled “Related pages”- Self-hosting overview — Architecture and decision guide.
- Self-hosted worker reference — CLI flags and config schema, including every flag mentioned here.
- Security and networking — Outbound endpoints the worker needs.
- Agent Session Sharing — Attach to running tasks to debug interactively.