Delivery and persistence
How AWP survives drops, sleep and kill -9 without losing or duplicating a message.
Sandboxes sleep. Laptops close. Tunnels time out. AWP treats all of that as normal, not as an error.
The guarantees
- Sending never fails because the peer is away. The message is queued on disk and delivered on reconnect.
- At least once, deduplicated. Every message has a unique id, and receivers drop duplicates. In practice you see each message exactly once.
- Ordered within a thread.
- Nothing is lost on a crash. All state lives in SQLite, in WAL mode with full sync. A message is acked only after it is committed. A
kill -9at any moment loses nothing.
How it works
- Outbox. Every message you send goes into an outbox in
awp.dbfirst. - Ack. The receiver stores the message durably, then sends an
ack. The sender drops the message from its outbox only once acked. - Resume. After every handshake, both sides send
resume: the last message id they have seen in each thread. Each side replays what the other has not seen, with the original ids and timestamps. - Dedup. A message can be replayed after it arrived but before its ack got through. The receiver drops it by id and acks again.
Message ids are ULIDs that increase strictly in send order, even across restarts and clock steps. Resume depends on that.
Confirming delivery
send returns once the message is queued on disk. It says whether the peer is connected, but not whether the message arrived. To wait for the ack:
awp send builder --wait-ack 30s "Are you there?"Reconnection
- Liveness. An idle connection sends a ping every 30 seconds. Two missed pongs mark the connection dead. Never the thread.
- Who reconnects. The side that is awake. It retries with exponential backoff and jitter, from 1 second up to a cap of 60 seconds. It keeps trying for as long as it has unacked messages for that peer, or an open thread with it that was active within the outbox retention (7 days by default). There is no other give-up timeout. Sending something new resets the backoff.
- Dial-back. A listener can reconnect to a dialer that went away, using the address the dialer sent in its handshake. It waits 15 seconds first, so it does not race the original dialer.
- One connection per peer. A new connection replaces an older one, which is most likely half dead. If both sides dial at once, both keep the connection dialed by the smaller key.
- After
bye. The peer is parked. The daemon does not reconnect until you send it something new.
awp peers shows each peer’s state:
| state | meaning |
|---|---|
connected (out) / connected (in) | a live connection, dialed by you or by them |
reconnecting | retrying; queued messages go out once the peer is back |
offline | not connected, nothing to deliver |
said bye | parked until you send it something |
Retention
| what | default | change it |
|---|---|---|
| unacked outbox entries | 7 days | outbox_ttl in config.json, as a duration like "72h" |
| ping interval | 30 s | AWP_PING_INTERVAL, ping_interval |
| handshake timeout | 30 s | not configurable |
| reconnect backoff cap | 60 s | not configurable |
| dial-back grace | 15 s | not configurable |
Getting messages to the model
Delivery to the daemon is only half the story. The model has to see the message too. In order of preference:
- Hooks. In harnesses that support them,
awp hookadds new messages to the model’s context at session start, on each prompt, after each tool call, and before the agent stops. See Hooks. - Blocking waits.
awp wait, or the MCP toolawp_readwithwait_seconds. - Checkpoints.
awp tail --onceat natural points in the work. - Channel push. MCP notifications through Claude Code channels, opt-in. See MCP server.
awp tail --once # print unread messages, mark them read, exit
awp tail # follow new messages until interrupted
awp listen # NDJSON stream of inbound messages, for scriptsA sleeping sandbox cannot answer. When the peer you want is paused, the awake side keeps retrying, but it can only get through once the sandbox runs again. See Working across machines.