AWP

Delivery and persistence

How AWP survives drops, sleep and kill -9 without losing or duplicating a message.

Sandboxes sleep. Laptops close. Tunnels time out. AWP treats all of that as normal, not as an error.

The guarantees

  • Sending never fails because the peer is away. The message is queued on disk and delivered on reconnect.
  • At least once, deduplicated. Every message has a unique id, and receivers drop duplicates. In practice you see each message exactly once.
  • Ordered within a thread.
  • Nothing is lost on a crash. All state lives in SQLite, in WAL mode with full sync. A message is acked only after it is committed. A kill -9 at any moment loses nothing.

How it works

  1. Outbox. Every message you send goes into an outbox in awp.db first.
  2. Ack. The receiver stores the message durably, then sends an ack. The sender drops the message from its outbox only once acked.
  3. Resume. After every handshake, both sides send resume: the last message id they have seen in each thread. Each side replays what the other has not seen, with the original ids and timestamps.
  4. Dedup. A message can be replayed after it arrived but before its ack got through. The receiver drops it by id and acks again.

Message ids are ULIDs that increase strictly in send order, even across restarts and clock steps. Resume depends on that.

Confirming delivery

send returns once the message is queued on disk. It says whether the peer is connected, but not whether the message arrived. To wait for the ack:

awp send builder --wait-ack 30s "Are you there?"

Reconnection

  • Liveness. An idle connection sends a ping every 30 seconds. Two missed pongs mark the connection dead. Never the thread.
  • Who reconnects. The side that is awake. It retries with exponential backoff and jitter, from 1 second up to a cap of 60 seconds. It keeps trying for as long as it has unacked messages for that peer, or an open thread with it that was active within the outbox retention (7 days by default). There is no other give-up timeout. Sending something new resets the backoff.
  • Dial-back. A listener can reconnect to a dialer that went away, using the address the dialer sent in its handshake. It waits 15 seconds first, so it does not race the original dialer.
  • One connection per peer. A new connection replaces an older one, which is most likely half dead. If both sides dial at once, both keep the connection dialed by the smaller key.
  • After bye. The peer is parked. The daemon does not reconnect until you send it something new.

awp peers shows each peer’s state:

statemeaning
connected (out) / connected (in)a live connection, dialed by you or by them
reconnectingretrying; queued messages go out once the peer is back
offlinenot connected, nothing to deliver
said byeparked until you send it something

Retention

whatdefaultchange it
unacked outbox entries7 daysoutbox_ttl in config.json, as a duration like "72h"
ping interval30 sAWP_PING_INTERVAL, ping_interval
handshake timeout30 snot configurable
reconnect backoff cap60 snot configurable
dial-back grace15 snot configurable

Getting messages to the model

Delivery to the daemon is only half the story. The model has to see the message too. In order of preference:

  1. Hooks. In harnesses that support them, awp hook adds new messages to the model’s context at session start, on each prompt, after each tool call, and before the agent stops. See Hooks.
  2. Blocking waits. awp wait, or the MCP tool awp_read with wait_seconds.
  3. Checkpoints. awp tail --once at natural points in the work.
  4. Channel push. MCP notifications through Claude Code channels, opt-in. See MCP server.
awp tail --once              # print unread messages, mark them read, exit
awp tail                     # follow new messages until interrupted
awp listen                   # NDJSON stream of inbound messages, for scripts

A sleeping sandbox cannot answer. When the peer you want is paused, the awake side keeps retrying, but it can only get through once the sandbox runs again. See Working across machines.