10 KiB
Concurrent runtime: multiple controllers, channels, supervision, reconnect
Date: 2026-08-20 Status: Approved (pending spec review)
Goal
Make the HASS runtime production-quality:
- Support multiple controllers.
- A single websocket read loop; no per-controller connections.
- The read loop broadcasts messages to controllers over channels (
TChan). - Each controller runs on its own thread (
async). - Controllers know nothing of the websocket connection: service calls go through an outbound channel; a writer thread does the actual sends.
- Crashed worker threads are restarted with backoff (
annotated-exceptionfor context). - The websocket reconnects on network failure.
Implemented as two slices, each independently buildable and committed:
- Concurrency architecture: bus, reader/writer threads, controller threads, channel interpreter, wiring. Any worker crash exits the process (today's behavior).
- Robustness: supervisor with backoff,
annotated-exception, fatal vs retryable failures, reconnect-as-restart.
Current state
HomeAssistant.Runtime.app runs everything on one thread inside
WS.runClient: auth handshake, a get_states debug dump, one
subscribe_events, then a loop that steps a single lightController
via dryRunHassEval. The get_states dump is dead debug code and is
removed (state pre-seeding is a separate future concern).
Architecture
HA websocket
▲ │
sends │ ▼ receives
┌─────────────┴──┐ ┌──────────────┐
│ writer thread │ │ reader thread │
└───────▲────────┘ └───────┬──────┘
│ │ decode → Value
│ TChan Service │ broadcast (single write)
│ ▼
┌───────┴───────────────────────────┐
│ busInbound (broadcast TChan) │
└──┬──────────────┬──────────────┬──┘
▼ dupTChan ▼ ▼
controller 1 controller 2 controller N (1 thread each)
Bus
data Bus = Bus
{ busInbound :: TChan Value -- newBroadcastTChanIO; reader writes only
, busOutbound :: TChan Service -- controllers → writer; fire-and-forget
, busConn :: TVar (Maybe WS.Connection)
, busGen :: CallIdGen -- existing IORef-based, thread-safe
}
-
busInboundcomes fromnewBroadcastTChanIO: a write-only broadcast channel. A plain never-readTChanwould pin its entire history; the broadcast variant does not. Controllers receive viadupTChanIO busInbound; the reader writes each message once. -
busConnis a plain current-value cell (noMVarblocking semantics).Maybe+ STMretrylets the writer block until a connection exists and letsdefaultMainspawn all workers up front: the initial connect is just the reader's first attempt, so first-connect failures and reconnect failures take the same backoff path. -
The controller list is static. Each entry is an existential:
data Controller = forall b. Controller T.Text (HASS (Event Value) b)No
Showconstraint: machine outputs are discarded; observability will come from a logging effect later. The name tags restart logs.
Send-safety without locks
No two threads ever send on the same live connection:
- The reader performs all setup sends (auth, subscribe) before
swapping the connection into
busConn; afterwards it only receives. - The writer only sends on connections read from
busConn.
On reconnect the reader builds a new connection, handshakes, then
atomically swaps busConn. A writer mid-send on the dead connection
throws, its supervisor restarts it, and it picks up the new connection.
Outbound backpressure
Fire-and-forget: services queue in busOutbound during outages and are
sent after reconnect. The queue is naturally bounded in practice: no
inbound events means controllers produce no calls.
Module layout
| Module | Concern |
|---|---|
HomeAssistant.Runtime.Supervisor |
Generic restart-with-backoff combinator, Backoff, Fatal; no HA knowledge |
HomeAssistant.Runtime.Bus |
Bus, newBus, channel interpreter for HASSEff |
HomeAssistant.Runtime.Connection |
readerAction, writerAction; connect/auth/subscribe, receive-decode-broadcast loop, send loop |
HomeAssistant.Runtime |
Glue: defaultMain, controllers, Controller, runController, step, dryRunHassEval |
Exports removed from HomeAssistant.Runtime: app, hassEval,
wsCallService, receiveJSON (internal or superseded). Kept:
defaultMain, step, dryRunHassEval, CallIdGen, mkCallIdGen.
app/Main.hs unchanged.
Connection lifecycle
Reader action (restartable unit; restart = reconnect):
connect → expect auth_required → send token → expect auth_ok
→ send subscribe_events state_changed (id from busGen)
→ swap busConn
→ forever: receive → decode → broadcast to busInbound
auth_invalid(and undecodable handshake messages) is fatal: a bad token cannot be fixed by retrying. The reader throwsFatal; the supervisor rethrows it (with annotations) and the process exits.- Undecodable messages in the receive loop are not fatal: log a warning and skip. Reconnecting cannot fix a decode problem, so crash-restarting would just be a hot loop.
- Controller Mealy state survives reconnects (explicit decision). State
may be stale until each watched entity's next
state_changedevent. Re-seeding viaget_statesis a separate future concern.
Writer action:
forever: readTChan busOutbound
→ readTVar busConn (retry until Just)
→ encode with fresh id from busGen → send
Encoding fixes a latent bug: today's wsCallService drops
serviceData; the writer encodes it as "service_data" when present.
Supervision
supervised :: Text -> Backoff -> IO a -> IO Void -- never returns normally
-
Catches synchronous exceptions; rethrows
SomeAsyncException(no restarting on cancellation). -
On
Fatal: rethrow with annotations; process exits. -
Otherwise: log (component name, attempt, annotated exception), sleep per backoff, restart the action.
-
Backoff: exponential from base, capped; resets after a quiet period.
data Backoff = Backoff { backoffBase :: NominalDiffTime -- first restart delay, e.g. 100ms , backoffCap :: NominalDiffTime -- max delay, e.g. 30s , backoffQuiet :: NominalDiffTime -- uptime that resets delay, e.g. 30s }Delay after the n-th consecutive crash:
min cap (base * 2^(n-1)). -
annotated-exceptionadds context at catch sites (component, phase such as "authenticating" / "receiving") so logs read like[reader] attempt 3: ConnectionClosed while receiving.
Wiring
- Slice 1:
defaultMain=newBus→mapConcurrently_over reader, writer, and controller actions. First worker crash cancels the rest and exits the process. - Slice 2: each action wrapped in
supervised; main waits forever. The workerIOactions are identical in both slices; only the spawning changes.
Controller runner
runController :: Bus -> Controller -> IO Void
-- dupTChanIO busInbound
-- loop: readTChan → step (channelHassEval bus) → discard output
On a supervised restart the action re-dups: a fresh port sees only messages written after the dup, consistent with the machine also restarting from its initial state. Messages broadcast during a restart window are lost (accepted, documented).
Interpreter
channelHassEval :: Bus -> HASSEff a -> IO a
channelHassEval bus (CallService svc) = atomically (writeTChan (busOutbound bus) svc)
channelHassEval _ (Pure a) = pure a
dryRunHassEval stays exported for experiments/tests.
Logging
Tagged plain lines ([reader] …) via putStrLn. No framework.
Dependencies
Added: stm, async (slice 1); annotated-exception (slice 2).
After editing the cabal file, regenerate default.nix via
nix run nixpkgs#cabal2nix -- ./. > default.nix (AGENTS.md).
Testing
hspec for unit tests, hedgehog for property tests (AGENTS.md):
- hspec:
- Broadcast semantics: two
dupTChanports see all writes, in order. channelHassEvalwritesServiceto the outbound chan.runControllerend-to-end on a realBuswith a test controller: feed values intobusInbound, assert services land inbusOutbound.- Writer encoding: golden JSON for
call_serviceincl.service_data. - Supervisor (tiny backoff values): crash-twice-then-succeed restarts;
Fatalrethrown, not restarted; async exceptions not restarted.
- Broadcast semantics: two
- hedgehog:
- Backoff property: delays double from base, clamp at cap, reset after quiet period.
Slices
- Concurrency architecture:
Bus(already reconnect-ready:TVar (Maybe Connection)),channelHassEval,readerAction,writerAction,runController, module split, cabal deps, wiring viamapConcurrently_, tests for bus/runner/encoding. Decode log-and-skip included. - Robustness:
Supervisormodule,annotated-exceptiondep, backoff,Fatalauth failures, rewiredefaultMainto supervised workers, supervisor tests and backoff property.
In slice 1 an auth_invalid response exits the process like any other
crash; the fatal/retryable distinction only exists once the supervisor
does (slice 2).
Each slice: builds clean, tests green, committed separately.
Non-goals
- Request/response correlation for service calls (approach B, deferred; outbound message type can grow later).
- Dynamic controller registration (static list).
- Structured logging / logging effect (future).
- Host/port/token configuration via env (beyond current
HA_TOKEN). - State pre-seeding via
get_states(user's future concern). - Multiple websocket connections.