Automation Platform > Deployment & hosting
Self-hosting troubleshooting
# Self-hosting troubleshooting Diagnostic guides for the `oz-agent-worker` daemon and its task execution. Use this page when a worker won't start, won't connect, tasks stay queued, or tasks fail. :::note The steps below apply to the [managed architecture](/platform/self-hosting/#managed-architecture) (`oz-agent-worker` daemon). For [unmanaged](/platform/self-hosting/unmanaged/) deployments, refer to the documentation for the environment running `oz agent run` (e.g., GitHub Actions, Kubernetes). ::: --- ## Worker won't start ### Docker backend **Cause:** Docker isn't running, or the daemon platform isn't supported. **Fix:** 1. Verify Docker is running: `docker info`. 2. Confirm the daemon platform is `linux/amd64` or `linux/arm64`. Windows containers are not supported. 3. If the worker runs inside Docker, confirm the `/var/run/docker.sock` mount is correct and the mounting user has permission to the socket. ### Kubernetes backend **Cause:** The startup preflight Job failed. Common reasons include insufficient RBAC, restrictive Pod Security policies, or an unreachable Kubernetes API server. **Fix:** 1. Check the worker logs for the preflight diagnostic message. 2. Confirm the worker's namespace has these permissions: `create`, `get`, `list`, `watch`, `delete` on `jobs`; `get`, `list`, `watch` on `pods`; `get` on `pods/log`; `list` on `events`. 3. Confirm the task namespace allows pods with a **root init container** (required for sidecar materialization). 4. If your cluster restricts image sources, set `preflight_image` in the worker config to an allowlisted image (default is `busybox:1.36`). 5. To pull the preflight image from a private registry, configure `imagePullSecrets` in `pod_template` — these secrets also apply to the preflight Job. ### Direct backend **Cause:** The `oz` CLI isn't installed or isn't on the worker's `PATH`. **Fix:** 1. Install the Oz CLI on the worker host. See [Installing the CLI](/reference/cli/#installing-the-cli). 2. If the CLI isn't on `PATH`, set `oz_path` in the config file to the absolute path of the `oz` binary. --- ## Worker won't connect **Cause:** The API key is invalid, expired, or the host cannot reach the Automation Platform's backend. **Fix:** 1. Confirm your API key is correct, not expired, and has team scope. 2. Regenerate the API key in **Settings** > **Cloud platform** > **API keys** if you suspect it's invalid. 3. Ensure the host has outbound internet access to `oz.warp.dev:443`. 4. Check that no firewall rules are blocking WebSocket connections to `wss://oz.warp.dev`. 5. Increase log verbosity with `--log-level debug` to see connection details. See [Security and networking](/platform/self-hosting/security-and-networking/#network-requirements) for the full list of outbound endpoints the worker needs. --- ## Tasks not being picked up **Cause:** The worker isn't running, the `--host` value doesn't match the worker's `--worker-id`, or the worker and task belong to different teams. **Fix:** 1. Confirm the worker is running and connected. Check the worker logs for `Listening for tasks` or similar. 2. Verify the `--host` (or `worker_host`) value you passed matches your `--worker-id` exactly. Case-sensitive. 3. Ensure the worker's team matches the team creating the task. --- ## Metrics not appearing **Cause:** The worker is running but metrics aren't showing up in Prometheus or your collector. **Fix:** 1. Verify `OTEL_METRICS_EXPORTER` is set correctly on the worker process. Run `curl -s localhost:9464/metrics` from the worker host (for `prometheus` mode) to confirm the endpoint is serving. 2. For Prometheus scrape mode, confirm the bind address is `0.0.0.0` (not `localhost`) when running in Docker or Kubernetes. `localhost` is only reachable from inside the container. 3. Confirm no firewall or network policy blocks the metrics port (default `9464`). 4. For OTLP push mode, verify `OTEL_EXPORTER_OTLP_ENDPOINT` points to a reachable collector and that the protocol matches (`http/protobuf` vs `grpc`). 5. When using the Helm chart, confirm `metrics.enabled=true` is set. Check that the `Service` and (optionally) `PodMonitor` were created: `kubectl get svc,podmonitor -n <namespace>`. 6. If using `metrics.podMonitor.create=true`, verify the `monitoring.coreos.com` CRDs are installed in the cluster. The `PodMonitor` resource requires the Prometheus Operator. 7. Restart the worker with `--log-level debug` and look for metrics-related error messages at startup. See [Monitoring](/platform/self-hosting/monitoring/) for the full setup guide. --- ## Task failures **Cause:** A variety of reasons depending on backend. Start with the diagnostic steps common to all backends, then follow the backend-specific checks. **Fix (all backends):** 1. Review task logs in the <a href=https://oz.warp.dev>cloud agent dashboard</a> or via [session sharing](/agents/local-agents/session-sharing/). 2. Use `--no-cleanup` to keep the container, Job, or workspace around for inspection after failure. 3. Use `--log-level debug` to see detailed execution logs. 4. Ensure the worker machine or cluster has sufficient resources (CPU, memory, disk). ### Docker backend (task failures) 1. Verify Docker is running (`docker info`). 2. If using a custom image, confirm it is **glibc-based** (not Alpine/musl) and that its architecture matches the worker's Docker daemon platform. ### Kubernetes backend (task failures) First determine whether the task Pod started. A running container that exceeds its memory limit and a Pending Pod that cannot fit on a node require different fixes. The Helm chart's `worker.resources` configures the long-running worker Deployment only. Task containers have no worker-defined CPU or memory defaults unless you set resources in `pod_template` or assign an explicit [runner instance shape](/platform/self-hosting/managed-kubernetes/#size-task-workloads). #### Task container was terminated with `OOMKilled` `OOMKilled` means Kubernetes reports that a container started and then encountered an out-of-memory condition. Confirm the termination reason before changing resources. 1. Find the failed task Pod: ```bash kubectl get jobs,pods -n NAMESPACE ``` Replace `NAMESPACE` with the task namespace. Use the returned task Pod name as `POD_NAME` below. 2. Inspect the `task` container's terminated state, resource settings, and Pod events: ```bash kubectl describe pod POD_NAME -n NAMESPACE kubectl get pod POD_NAME -n NAMESPACE -o yaml ``` 3. Compare peak memory usage with the configured limit using your cluster metrics. Check the node for `MemoryPressure` and eviction events. 4. Reduce the task's peak memory use or increase its memory limit. For a workload-specific [runner](/platform/runners/), increase the instance shape. For a baseline shared by all tasks on the worker, change the `task` container's resources in `pod_template`. Increasing a runner's memory also increases the task container's memory request to the same value. Confirm that a compatible node has enough allocatable memory, or the replacement Pod can remain Pending. If the Pod reason is `Evicted` instead, diagnose node pressure rather than a container limit. Restore node headroom, add compatible capacity, or adjust scheduling and concurrency before rerunning the task. #### Task remains `Pending` with `FailedScheduling` A Pending Pod with a `PodScheduled=False` condition and `FailedScheduling` events has not started. Messages such as `Insufficient cpu` or `Insufficient memory` mean no eligible node has enough allocatable capacity for the Pod's requests. 1. Read the scheduler message and recent events: ```bash kubectl describe pod POD_NAME -n NAMESPACE kubectl get events -n NAMESPACE --sort-by=.lastTimestamp ``` 2. Compare the Pod's requests with node allocatable capacity and, when resource metrics are available, current usage: ```bash kubectl describe nodes kubectl top nodes kubectl top pods -n NAMESPACE --containers ``` 3. Review the Pod's `nodeSelector`, affinity, tolerations, and taints. A node with free resources is not eligible if another scheduling constraint excludes it. 4. Check how many task Jobs run concurrently. Set `max_concurrent_tasks` to keep aggregate requests within cluster capacity when needed. 5. Right-size requests only if the task can run reliably at the lower values. Otherwise, add compatible node capacity or configure cluster autoscaling for nodes that satisfy the Pod's scheduling constraints. Raising only a memory limit does not help an unschedulable Pod because the scheduler places Pods from requests. The worker stops waiting after the configured `unschedulable_timeout`; fix the scheduling constraint rather than extending the timeout when the cluster lacks capacity. #### Task exits with code `143` Exit code `143` generally indicates `SIGTERM`; it does not prove that a container ran out of memory. Check the container termination reason, Pod conditions, and events before choosing a remediation. ```bash kubectl describe pod POD_NAME -n NAMESPACE kubectl get events -n NAMESPACE --sort-by=.lastTimestamp ``` Look for eviction, preemption, node drain, `activeDeadlineSeconds`, or manual deletion. Correlate the event timeline with node pressure and resource metrics. Treat the failure as OOM only when Kubernetes reports `OOMKilled`. #### Other Kubernetes task failures * **Image pull failures** - Inspect `imagePullSecrets` in `pod_template`. * **Admission policy rejections** - Review Pod Security Standards, OPA Gatekeeper, Kyverno, or similar admission controllers. With cleanup enabled, failed Jobs and Pods remain temporarily available for diagnosis before Kubernetes TTL cleanup. Use `--no-cleanup` when you need to retain them longer. ### Direct backend (task failures) 1. Verify the Oz CLI is accessible. 2. Verify the workspace root directory has write permissions for the user running the worker. --- ## Image pull failures ### Docker backend (image pull) 1. If using a private registry, ensure Docker credentials are available to the worker. See [Private Docker registries](/platform/self-hosting/managed-docker/#private-docker-registries). 2. Try pulling the image manually on the worker host: `docker pull <image>`. ### Kubernetes backend (image pull) 1. Configure `imagePullSecrets` in the `pod_template` section of your worker config. 2. Verify the Secret exists in the task namespace and contains valid credentials. ### Both backends (image pull) * Verify the image exists and the tag is correct. * Check network connectivity from the worker/cluster to the registry. --- ## Related pages * [Self-hosting overview](/platform/self-hosting/) — Architecture and decision guide. * [Self-hosted worker reference](/platform/self-hosting/reference/) — CLI flags and config schema, including every flag mentioned here. * [Security and networking](/platform/self-hosting/security-and-networking/) — Outbound endpoints the worker needs. * [Agent Session Sharing](/agents/local-agents/session-sharing/) — Attach to running tasks to debug interactively.Walk me through resolving this issue: https://docs.warp.dev/platform/self-hosting/troubleshooting/Diagnose and fix common problems with self-hosted Automation Platform worker daemons across Docker, Kubernetes, and Direct backends.
Diagnostic guides for the oz-agent-worker daemon and its task execution. Use this page when a worker won’t start, won’t connect, tasks stay queued, or tasks fail.
Worker won’t start
Section titled “Worker won’t start”Docker backend
Section titled “Docker backend”Cause: Docker isn’t running, or the daemon platform isn’t supported.
Fix:
- Verify Docker is running:
docker info. - Confirm the daemon platform is
linux/amd64orlinux/arm64. Windows containers are not supported. - If the worker runs inside Docker, confirm the
/var/run/docker.sockmount is correct and the mounting user has permission to the socket.
Kubernetes backend
Section titled “Kubernetes backend”Cause: The startup preflight Job failed. Common reasons include insufficient RBAC, restrictive Pod Security policies, or an unreachable Kubernetes API server.
Fix:
- Check the worker logs for the preflight diagnostic message.
- Confirm the worker’s namespace has these permissions:
create,get,list,watch,deleteonjobs;get,list,watchonpods;getonpods/log;listonevents. - Confirm the task namespace allows pods with a root init container (required for sidecar materialization).
- If your cluster restricts image sources, set
preflight_imagein the worker config to an allowlisted image (default isbusybox:1.36). - To pull the preflight image from a private registry, configure
imagePullSecretsinpod_template— these secrets also apply to the preflight Job.
Direct backend
Section titled “Direct backend”Cause: The oz CLI isn’t installed or isn’t on the worker’s PATH.
Fix:
- Install the Oz CLI on the worker host. See Installing the CLI.
- If the CLI isn’t on
PATH, setoz_pathin the config file to the absolute path of theozbinary.
Worker won’t connect
Section titled “Worker won’t connect”Cause: The API key is invalid, expired, or the host cannot reach the Automation Platform‘s backend.
Fix:
- Confirm your API key is correct, not expired, and has team scope.
- Regenerate the API key in Settings > Cloud platform > API keys if you suspect it’s invalid.
- Ensure the host has outbound internet access to
oz.warp.dev:443. - Check that no firewall rules are blocking WebSocket connections to
wss://oz.warp.dev. - Increase log verbosity with
--log-level debugto see connection details.
See Security and networking for the full list of outbound endpoints the worker needs.
Tasks not being picked up
Section titled “Tasks not being picked up”Cause: The worker isn’t running, the --host value doesn’t match the worker’s --worker-id, or the worker and task belong to different teams.
Fix:
- Confirm the worker is running and connected. Check the worker logs for
Listening for tasksor similar. - Verify the
--host(orworker_host) value you passed matches your--worker-idexactly. Case-sensitive. - Ensure the worker’s team matches the team creating the task.
Metrics not appearing
Section titled “Metrics not appearing”Cause: The worker is running but metrics aren’t showing up in Prometheus or your collector.
Fix:
- Verify
OTEL_METRICS_EXPORTERis set correctly on the worker process. Runcurl -s localhost:9464/metricsfrom the worker host (forprometheusmode) to confirm the endpoint is serving. - For Prometheus scrape mode, confirm the bind address is
0.0.0.0(notlocalhost) when running in Docker or Kubernetes.localhostis only reachable from inside the container. - Confirm no firewall or network policy blocks the metrics port (default
9464). - For OTLP push mode, verify
OTEL_EXPORTER_OTLP_ENDPOINTpoints to a reachable collector and that the protocol matches (http/protobufvsgrpc). - When using the Helm chart, confirm
metrics.enabled=trueis set. Check that theServiceand (optionally)PodMonitorwere created:kubectl get svc,podmonitor -n <namespace>. - If using
metrics.podMonitor.create=true, verify themonitoring.coreos.comCRDs are installed in the cluster. ThePodMonitorresource requires the Prometheus Operator. - Restart the worker with
--log-level debugand look for metrics-related error messages at startup.
See Monitoring for the full setup guide.
Task failures
Section titled “Task failures”Cause: A variety of reasons depending on backend. Start with the diagnostic steps common to all backends, then follow the backend-specific checks.
Fix (all backends):
- Review task logs in the cloud agent dashboard or via session sharing.
- Use
--no-cleanupto keep the container, Job, or workspace around for inspection after failure. - Use
--log-level debugto see detailed execution logs. - Ensure the worker machine or cluster has sufficient resources (CPU, memory, disk).
Docker backend (task failures)
Section titled “Docker backend (task failures)”- Verify Docker is running (
docker info). - If using a custom image, confirm it is glibc-based (not Alpine/musl) and that its architecture matches the worker’s Docker daemon platform.
Kubernetes backend (task failures)
Section titled “Kubernetes backend (task failures)”First determine whether the task Pod started. A running container that exceeds its memory limit and a Pending Pod that cannot fit on a node require different fixes.
The Helm chart’s worker.resources configures the long-running worker Deployment only. Task containers have no worker-defined CPU or memory defaults unless you set resources in pod_template or assign an explicit runner instance shape.
Task container was terminated with OOMKilled
Section titled “Task container was terminated with OOMKilled”OOMKilled means Kubernetes reports that a container started and then encountered an out-of-memory condition. Confirm the termination reason before changing resources.
-
Find the failed task Pod:
Terminal window kubectl get jobs,pods -n NAMESPACEReplace
NAMESPACEwith the task namespace. Use the returned task Pod name asPOD_NAMEbelow. -
Inspect the
taskcontainer’s terminated state, resource settings, and Pod events:Terminal window kubectl describe pod POD_NAME -n NAMESPACEkubectl get pod POD_NAME -n NAMESPACE -o yaml -
Compare peak memory usage with the configured limit using your cluster metrics. Check the node for
MemoryPressureand eviction events. -
Reduce the task’s peak memory use or increase its memory limit. For a workload-specific runner, increase the instance shape. For a baseline shared by all tasks on the worker, change the
taskcontainer’s resources inpod_template.
Increasing a runner’s memory also increases the task container’s memory request to the same value. Confirm that a compatible node has enough allocatable memory, or the replacement Pod can remain Pending.
If the Pod reason is Evicted instead, diagnose node pressure rather than a container limit. Restore node headroom, add compatible capacity, or adjust scheduling and concurrency before rerunning the task.
Task remains Pending with FailedScheduling
Section titled “Task remains Pending with FailedScheduling”A Pending Pod with a PodScheduled=False condition and FailedScheduling events has not started. Messages such as Insufficient cpu or Insufficient memory mean no eligible node has enough allocatable capacity for the Pod’s requests.
-
Read the scheduler message and recent events:
Terminal window kubectl describe pod POD_NAME -n NAMESPACEkubectl get events -n NAMESPACE --sort-by=.lastTimestamp -
Compare the Pod’s requests with node allocatable capacity and, when resource metrics are available, current usage:
Terminal window kubectl describe nodeskubectl top nodeskubectl top pods -n NAMESPACE --containers -
Review the Pod’s
nodeSelector, affinity, tolerations, and taints. A node with free resources is not eligible if another scheduling constraint excludes it. -
Check how many task Jobs run concurrently. Set
max_concurrent_tasksto keep aggregate requests within cluster capacity when needed. -
Right-size requests only if the task can run reliably at the lower values. Otherwise, add compatible node capacity or configure cluster autoscaling for nodes that satisfy the Pod’s scheduling constraints.
Raising only a memory limit does not help an unschedulable Pod because the scheduler places Pods from requests. The worker stops waiting after the configured unschedulable_timeout; fix the scheduling constraint rather than extending the timeout when the cluster lacks capacity.
Task exits with code 143
Section titled “Task exits with code 143”Exit code 143 generally indicates SIGTERM; it does not prove that a container ran out of memory. Check the container termination reason, Pod conditions, and events before choosing a remediation.
kubectl describe pod POD_NAME -n NAMESPACEkubectl get events -n NAMESPACE --sort-by=.lastTimestampLook for eviction, preemption, node drain, activeDeadlineSeconds, or manual deletion. Correlate the event timeline with node pressure and resource metrics. Treat the failure as OOM only when Kubernetes reports OOMKilled.
Other Kubernetes task failures
Section titled “Other Kubernetes task failures”- Image pull failures - Inspect
imagePullSecretsinpod_template. - Admission policy rejections - Review Pod Security Standards, OPA Gatekeeper, Kyverno, or similar admission controllers.
With cleanup enabled, failed Jobs and Pods remain temporarily available for diagnosis before Kubernetes TTL cleanup. Use --no-cleanup when you need to retain them longer.
Direct backend (task failures)
Section titled “Direct backend (task failures)”- Verify the Oz CLI is accessible.
- Verify the workspace root directory has write permissions for the user running the worker.
Image pull failures
Section titled “Image pull failures”Docker backend (image pull)
Section titled “Docker backend (image pull)”- If using a private registry, ensure Docker credentials are available to the worker. See Private Docker registries.
- Try pulling the image manually on the worker host:
docker pull <image>.
Kubernetes backend (image pull)
Section titled “Kubernetes backend (image pull)”- Configure
imagePullSecretsin thepod_templatesection of your worker config. - Verify the Secret exists in the task namespace and contains valid credentials.
Both backends (image pull)
Section titled “Both backends (image pull)”- Verify the image exists and the tag is correct.
- Check network connectivity from the worker/cluster to the registry.
Related pages
Section titled “Related pages”- Self-hosting overview — Architecture and decision guide.
- Self-hosted worker reference — CLI flags and config schema, including every flag mentioned here.
- Security and networking — Outbound endpoints the worker needs.
- Agent Session Sharing — Attach to running tasks to debug interactively.