Turn loop + MCP

How the harness wakes up, what it asks claude to do, and what tools claude has access to in return.

The loop

Each agent harness (hive-agent — one serve-loop binary for all agents) runs:

  1. Check the pause marker (<harness>/paused). While it exists the loop does nothing but re-stat it every 5 s — no broker poll, no claude process. Because step 1 is never reached, messages stay queued and unacked, so a resume drains the backlog instead of losing it; reminders and todo wakes buffer in their channels. Set it with hivectl agent <name> pause or the dashboard toggle; see persistence.
  2. Long-poll Recv on its socket. The host-side broker (broker.rs::recv_blocking_batch) returns immediately if there's a pending message, otherwise waits up to 30 s for a broker Sent event for this recipient.
  3. Pop one message. Peek the remaining inbox depth with Status.
  4. Emit LiveEvent::TurnStart { from, body, unread } onto the SSE bus.
  5. Spawn claude (one process per turn) and pipe the wake prompt over stdin.
  6. Stream stdout (JSON lines) into the bus as LiveEvent::Stream(value). Pump stderr as Note.
  7. Wait for claude to exit and classify the turn's outcome from the stream + exit — success, compaction, rate-limit, auth-failure, or hard failure. The outcome drives the post-turn action (see Turn outcomes); compaction is handled inside the session (see Compaction). Rate-limit and auth-failure detection is described below.
  8. Emit LiveEvent::TurnEnd { ok, note }. Sleep poll_ms to avoid tight loops on transient failures.

Failure detection and login

Harness binary shape

Two sibling crates, both role-agnostic (there is one role: agent — the privilege boundary lives server-side at the broker socket (/run/hive/mcp.sock), which refuses privileged Request variants regardless of who sends them):

hive-agent's wire types (hive_core_agent_sock::{Request, Response} — one unified enum shared by the agent and manager sockets) and its turn loop are factored through a small Surface trait with one zero-sized impl, so the loop itself has no per-role branches. See hive-agent/src/main.rs's module doc for the trait shape.

Boot wiring

serve_main reads HIVE_PORT (default DEFAULT_WEB_PORT) + HIVE_LABEL (default "hive" for standalone runs; the meta flake sets it unconditionally for any container-deployed agent; see docs/conventions.md::Hive identity for the env stack), opens turn-stats sqlite, prepares the on-boot files (see claude-invocation), installs claude plugins, spawns web_ui::serve + vacuum::run, and either drops into serve_loop directly (Online) or parks on the login flow first (NeedsLogin). Forge notifications are polled by their own process, not this loop — see hive-forge-notify in forge.md.

Boot also opens the todos store and the socket in-container producers dial. Matrix / bash / forge-notify daemons and the in-process disk_watch todo producer (low state-disk space) are the built-in producers, but the socket accepts any subsystem marker — a user-configured MCP server can push its own todos the same way. See docs/persistence.md for what each built-in todo producer watches and how the store + get_loose_ends merge work.

Plugin install failures are not fatal: each entry comes back as a human-readable failure string that gets routed via Surface::send_to_parent to the agent's topology parent (the broker resolves <parent> per topology::resolve_recipient; root agents and the manager fall through to operator).

Turn outcomes

turn::TurnOutcome (Result<bool, TurnError>Ok(compacted) on success, else a TurnError) drives the post-claude branch:

Outcome Action
Ok(_) (false normal / true compacted) ack_turn
Err(PromptTooLong) drive_turn archived the session (the lib already compacted + retried and it still overflowed); requeue inflight so the message redelivers into a fresh session that fits — no status park
Err(RateLimited) sleep HIVE_RATE_LIMIT_SLEEP_SECS (default 300), requeue inflight, status back to online
Err(AuthFailed) emit needs_login_idle sentinel, requeue inflight, park in wait_for_login
Err(SessionNotFound) resume + create self-heal both missed ("shouldn't happen"); requeue inflight so the next turn creates fresh — no status park, message not dropped
Err(ApiStall) idle watchdog killed claude after HIVE_TURN_IDLE_SECS (default 600) of output silence; sleep HIVE_STALL_SLEEP_SECS (default 60), requeue inflight, status back to online
Err(Failed(err)) route [system] \` claude turn failed:\ntoviasend_to_parent`

ApiStall catches an Anthropic API stall — a multi-retry connection storm where the stream goes silent for minutes. The idle watchdog lives in hive-claude's driver (Config::idle_timeout, enforced around child.wait()): the timer resets on every stdout line, so a large but still-streaming turn is never cut — only complete output silence for the window trips it. The harness sets the window from HIVE_TURN_IDLE_SECS (0 disables) and maps the driver's Error::IdleTimeout onto TurnError::ApiStall.

After the outcome handler, the stats sink records a row. handle_turn reports the result to serve_loop via TurnControl { auth_failed } — on auth failure the loop parks in wait_for_login; otherwise it loops straight back to the idle wait (step 1). There is no same-turn self-continue mechanism: every multi-step continuation rides an external wake instead — a new inbox message, a remind, or an in-container todo wake (bash-task completion, forge notification, matrix activity). Ending the turn and letting one of those drive the next one is strictly better than parking in-process: it checkpoints the session and observes wakes that only reach the harness between turns.

Sub-pages

The rest lives alongside this page, in three topic files:

Per-subsystem impl detail lives in each module's //! doc-comment; these pages describe present-state behaviour + wiring, not line-level mechanics.