The Runner management page shows Idle, but the target Job remains queued.

Fastest fix: do not add another Mac first. Check, in order, whether no Runner matches the labels and Runner Group, whether all eligible Runners are busy, or whether the macOS service is online but cannot claim work. Add remote Mac capacity only after routing and service recovery still leave a stable queue.

This guide is for:

  • Mobile developers using GitHub Actions self-hosted Mac Runners for Xcode builds.
  • DevOps engineers responsible for labels, Runner Groups, repository access, and persistent node services.
  • Platform owners deciding whether to repair an existing Runner, register an isolated replacement, or expand remote Mac capacity.
01

Start with the queue state, not the machine count

A queued Job does not always mean that GitHub Actions is waiting for an available Mac. The workflow may still be waiting for a prerequisite, an approval, a concurrency slot, or a Runner that satisfies every routing condition. The first diagnostic task is to identify the first blocking state visible in the workflow run.

Use the run page, not only the Runner management page. Open the affected Job and record:

  • The exact runs-on value.
  • Any required environment, approval, or protected-resource status.
  • The Job annotation shown beside the queued state.
  • The repository, organization, or enterprise scope of the Runner.
  • The labels currently shown on the Runner.
  • The Runner Group assigned to that Runner.
  • Whether other Jobs are consuming the same eligible pool.

GitHub routes a self-hosted Job through its requested labels, Runner Group access, and Runner availability. The official Runner selection documentation should be used as the reference when the interface wording differs from the examples here.

Freeze the evidence before making changes. Save the workflow revision, the current labels, the Runner Group membership, the repository access setting, and the service status. Changing YAML, labels, permissions, and launch services at the same time can remove the evidence needed to distinguish a routing fault from a node fault.

A useful first classification is:

Observed condition Most likely fault layer First evidence to inspect Low-risk next action
Job is queued before a Runner assignment appears Workflow condition, approval, concurrency, or routing Job annotation, workflow run details, concurrency, runs-on Resolve the first visible blocker without changing the production pool
Matching labels exist, but the repository cannot use the Runner Runner Group access Group policy and repository authorization Test access with a no-secret diagnostic Job
Eligible Runners are visible but all are busy Capacity or stuck work Active Jobs, processes, workspace locks Confirm real occupancy before cleaning or adding nodes
Runner is listed as online or idle but does not claim work Runner process, service account, network, or work directory Diagnostic log, service state, recent restart evidence Repair the service only after preserving logs and recovery access
A minimal Job runs but the production Job queues Production-specific label, environment, concurrency, or toolchain rule Differences between the two workflows Correct the narrow condition rather than widening labels

Do not treat an idle display as proof that the Runner can execute a Job. “Idle” describes the state reported by the Runner control plane. It does not prove that the service can maintain its connection, access its working directory, or start the requested process.

02

Online self-hosted Runners can still refuse work

An online self-hosted Runner can remain unused when the local process and the GitHub routing decision disagree. SSH access proves that the host is reachable for administration. It does not prove that the Runner service is registered correctly, connected to the right scope, or able to launch a Job.

Check the Runner’s diagnostic log from the host. The exact path depends on the Runner installation and version, so the current macOS monitoring and troubleshooting documentation should take precedence over copied commands. Look for evidence of:

  • Repeated connection loss or reconnect attempts.
  • Failure to read or write the working directory.
  • A service account that differs from the account used during setup.
  • Permission failures in the Runner directory or temporary workspace.
  • A process that starts manually but does not survive a logout or reboot.
  • Registration errors after a host restore or service migration.
  • An update or restart that left the service without a usable session.

On macOS, inspect the launchd service state from the same administrative context used by the Runner. A basic evidence capture can look like this:

uname -a
id
pwd
ls -ld .
launchctl list | grep -i runner

Example output:

Darwin build-node  macOS
uid=... 
/Users/ci/actions-runner
drwxr-xr-x  ...  /Users/ci/actions-runner
...  com.example.actions.runner

The output is evidence, not a pass condition. A service entry can exist while the service is failing repeatedly. Compare the service account, working directory, and recent diagnostic log entries. If the Runner works only after an interactive SSH session starts, the persistent service configuration is not yet reliable.

Avoid immediately removing and re-registering the Runner. Removal changes the control-plane identity and can discard the easiest path to compare the old node with a replacement. If removal becomes necessary, preserve the logs, workflow evidence, configuration notes, and access route first. GitHub’s Runner removal guidance explains the effects of removing a Runner and should be followed before revoking credentials or deleting the local installation.

A restart is a recovery action, not a diagnosis. Record the failing state first, then restart only the service or node that the evidence identifies.

If the service registration is damaged, repair or re-register it only after confirming that the failure is local to registration or service startup. Keep an emergency administrative path available. A service reinstall without a preserved recovery route can turn a recoverable Runner fault into an unavailable build node.

03

Do runs-on labels match, or is the Job waiting for an impossible Runner?

The runs-on value is a routing requirement, not a suggestion. When a Job specifies multiple self-hosted labels, the same Runner must satisfy the full combination. A Runner with the correct operating system label but without the required architecture or toolchain label is not eligible.

For example:

jobs:
  build:
    runs-on: [self-hosted, macOS, arm64, xcode]
    steps:
      - uses: actions/checkout@v4
      - run: xcodebuild -version

The exact labels in production may differ. The important point is to compare the workflow declaration with the labels shown on the actual Runner. Check for:

  • A label removed during node rebuild.
  • A spelling or case difference introduced in YAML.
  • An architecture label that no longer describes the host.
  • A custom Xcode or SDK label that was never added to the replacement node.
  • A label copied from an old Runner while the new node uses a different naming convention.
  • A runs-on expression that resolves differently from the value expected by the maintainer.

GitHub’s self-hosted Runner label documentation explains how labels identify eligible machines. The self-hosted Runner routing reference should be used to confirm how the current interface represents online, offline, and busy states.

Create a temporary diagnostic Job with no signing secrets, deployment credentials, or production artifacts:

name: runner-routing-check

on:
  workflow_dispatch:

jobs:
  route-check:
    runs-on: [self-hosted, macOS, arm64]
    steps:
      - name: Show execution identity
        run: |
          sw_vers
          uname -m
          whoami
          pwd

A successful run proves that the selected label combination can route to a usable Runner. It does not prove that the production label set, signing environment, simulator state, or build workload is healthy.

Do not permanently replace a precise production selector with self-hosted simply to make the Job start. That can send an Xcode build to a node without the required architecture, SDK, certificates, or isolation policy. Use a narrow diagnostic selector, compare the result, and then restore the intended production constraints.

04

Runner Group access can block a matching label

A label match does not override Runner Group authorization. A Runner can show the expected labels and still be unavailable to a repository if the repository is not permitted to use the Group containing that Runner.

Inspect the Runner’s scope and Group assignment, then compare them with the repository that owns the queued workflow. Check:

  • Whether the Runner is assigned at repository, organization, or enterprise scope.
  • Whether the target repository is included in the Group’s allowed repository list.
  • Whether a policy change moved the Runner into a more restricted Group.
  • Whether the workflow is being tested from a fork or a repository outside the authorized boundary.
  • Whether the Job requests a Group or label combination that cannot be used by that repository.

Use the official Runner Group access documentation for the current permission model. Do not widen organization access as a first response. A broad permission change can make a powerful Mac node available to repositories that were never intended to run code on it.

A low-risk access test should contain no secrets and no release action:

name: runner-group-access-check

on:
  workflow_dispatch:

jobs:
  access-check:
    runs-on: [self-hosted, macOS]
    steps:
      - run: |
          echo "Runner Group routing test"
          sw_vers

Run the test from the exact repository that is experiencing the queue. Record whether the Runner becomes selectable, whether a Job is assigned, and whether the visible Runner list changes after the policy adjustment. If access is corrected, retain the narrowest repository authorization that satisfies the CI design.

This is also where a confusing “idle” state often becomes explainable. The Runner may be idle because it has no authorized work, while the Job remains queued because the repository cannot use that Runner. The machine is healthy; the routing boundary is not.

05

How should you separate real occupancy from a routing failure?

When the eligible Mac Runners are genuinely busy, the queue is a capacity symptom. Before adding a node, confirm that the occupancy is real and attributable to active work.

Inspect the Actions run list and the local host at the same time. Classify each apparent occupant as:

  • An active Xcode build.
  • A Simulator test process that is still running.
  • A signing or packaging task waiting on a child process.
  • A long-running maintenance or scheduled Job.
  • A completed Job whose Runner process failed to return cleanly.
  • A stale workspace or lock left after interruption.
  • A process unrelated to Actions that is consuming the build environment.

The local process list can help identify whether the host is doing work:

ps aux | grep -E 'Runner|xcodebuild|simctl|swift|clang' | grep -v grep

Treat this as an investigation command, not a universal cleanup command. Killing a process can corrupt a build, interrupt signing, or remove evidence of the failure. First correlate the process with the Job shown in the Actions interface.

Xcode builds, Simulator tests, signing, and release packaging can have different isolation needs. A single Runner may appear available between steps while its workspace, simulator state, or keychain remains unsuitable for the next Job. That is not automatically a routing defect. It may indicate that the workflow needs clearer cleanup, separate labels, or a dedicated task pool.

The distinction matters:

  • If no eligible Runner is ever assigned, investigate labels and access.
  • If an eligible Runner is assigned and stays busy with a valid Job, investigate concurrency and workload design.
  • If the Job finished but the Runner remains occupied, investigate cleanup, child processes, and service recovery.
  • If the Runner returns to idle but new Jobs remain queued, return to routing and workflow-level conditions.

GitHub Actions can intentionally serialize work through concurrency. Review the workflow and repository settings against the official concurrency documentation. A concurrency group can make a Job wait even when a Mac Runner is idle. Do not remove a concurrency rule merely to reduce the queue; it may protect release branches, shared signing state, or deployment ordering.

06

The repair and retest sequence

Use this sequence when the evidence is incomplete. Each step changes only one fault layer at a time.

  • [ ] Save the queued Job URL, workflow revision, runs-on value, Runner labels, Runner Group, and repository access state.
  • [ ] Confirm whether the Job is waiting for approval, a prerequisite, concurrency, routing, or an available Runner.
  • [ ] Compare every requested label with the labels on the intended Mac, including architecture and custom toolchain labels.
  • [ ] Confirm that the target repository is authorized to use the Runner Group.
  • [ ] Run a no-secret diagnostic Job with the narrowest label set that should match the node.
  • [ ] Compare the Runner management state with the local Runner log and macOS service state.
  • [ ] Check active Jobs and local processes before stopping anything or clearing a workspace.
  • [ ] Repair the smallest confirmed fault: YAML routing, Group access, service startup, or stale local work.
  • [ ] Run a minimal command Job and confirm that the Runner claims and completes it.
  • [ ] Run a production-like Xcode build with the intended labels and artifact path.
  • [ ] Confirm that the artifact returns correctly and that the Runner becomes available for the next eligible Job.
  • [ ] Record whether the result supports repair, isolated replacement, task-pool separation, or capacity expansion.

The retest should use three levels:

  1. Routing test: a minimal command Job reaches the intended Runner.
  2. Execution test: the node can run the required macOS and Xcode commands.
  3. Production-like test: the workflow uses the real label set, checkout behavior, signing boundary, simulator requirements, and artifact upload path.

A routing test that passes while the production Job remains queued points back to production-specific conditions. A production-like Job that starts but fails during execution is not a queue-routing success; it is a separate build or environment failure.

07

When should the team add another remote Mac Runner?

Add capacity only after the evidence shows a persistent eligible queue caused by real workload occupancy. The decision should not be based on an idle label in the management page, a single slow build, or a temporary service restart.

Repair the existing node when:

  • The Job does not match because of an incorrect label.
  • The repository lacks Runner Group access.
  • The Runner service is offline, unstable, or unable to read its workspace.
  • A completed Job left behind a recoverable process or lock.
  • A workflow-level concurrency rule is serializing work unintentionally.

Rebuild or isolate a node when:

  • Registration state is damaged and the logs support a clean replacement.
  • The current host has mixed workloads that cannot be safely cleaned between Jobs.
  • The service account or working directory cannot be restored without weakening security.
  • A replacement is needed to preserve the existing node for forensic comparison.

Expand the remote Mac pool when:

  • The intended labels and Group permissions are correct.
  • The Runner service consistently claims and completes diagnostic work.
  • Production-like Jobs complete and release the node.
  • Eligible nodes remain occupied by valid, concurrent workloads.
  • The queue persists after workflow concurrency and task partitioning have been reviewed.

For teams that need a controlled isolation test, a separate remote Mac can be useful before changing the production pool. NodeMini’s remote Mac options can provide a distinct host for reproducing the same workflow, checking label and Group behavior, and validating service recovery without immediately altering the existing build node.

The acceptance record should state the actual conclusion: routing fixed, service restored, node rebuilt, task pool separated, or capacity expanded. Avoid the vague conclusion “added another Mac and the queue cleared.” That result does not prove the original fault was capacity-related.

08

Keep the repaired Mac Runner observable

After the queue clears, retain the evidence that proves the fix. Store the final workflow selector, Runner labels, Group authorization, service configuration, and the diagnostic Job result with the incident record.

For a persistent remote build node, also document:

  • Which account owns the Runner process.
  • Where diagnostic logs are collected.
  • How the service is checked after a restart.
  • Which Jobs may share the node.
  • Which Jobs require exclusive access.
  • How to preserve access before removing or re-registering the Runner.
  • Which workflow proves routing, execution, and artifact delivery.

A remote Mac used for CI should be treated as a build service, not merely as a host that accepts SSH connections. NodeMini’s remote Mac development environment can be used when the team needs an isolated environment to reproduce a queued workflow or validate a backup build path. The existing platform should still be repaired or expanded according to evidence, not replaced by an untested endpoint.

If the current setup is a Mac mini hosted in a location that makes service recovery, network access, or physical replacement difficult, compare it with a remote Mac using the same workflow and acceptance checks. The local Mac mini may remain appropriate for stable long-running workloads with direct hardware access, but it can be a poor fit when the team needs a disposable test node, rapid isolation, or a second build environment. Remote rental is the more flexible option when the requirement is temporary CI capacity or a controlled reproduction host rather than permanent ownership.

The practical sequence is therefore straightforward: identify the first queue reason, verify the complete label and Group route, prove that the macOS service can claim work, and only then measure real capacity pressure. That order prevents an expensive node expansion from hiding a permissions, routing, or service failure.