Working across machines
Connect agents on laptops, cloud sandboxes and private networks, and deal with sleep.
AWP’s main case is two agents on different machines, often behind NAT, often in sandboxes that sleep. This guide covers what to expect and the lessons from running it that way.
Pick a transport
| between | use | why |
|---|---|---|
| any two machines | tailcat (default) | works through NAT, encrypted, no setup |
| machines on one private network (Fly 6PN, a VPC) | tcp:HOST:PORT | lower latency, no relay |
| two homes on one machine | unix:/path | no network at all |
Tailcat is on by default. To add a private-network listener as well, list both in the home’s config, so a daemon started on demand picks them up too:
{ "listen": ["tailcat", "tcp:[fdaa::3]:7000"] }Then restart the daemon with awp down and awp up. AWP_LISTEN and awp daemon --listen do the same for one start. See Environment variables and config.
Peers on the private network connect with awp connect tcp:[fdaa::3]:7000. Remember that TCP and Unix bindings have no transport encryption. The handshake still authenticates both keys, and awp refuses plain TCP to public addresses.
Latency
Once a connection is up, round trips are about 1 ms on one host and about 10 ms between cloud sandboxes in one region. awp web shows each link’s latency.
- A fresh tailcat dial takes seconds, not milliseconds.
- The first round-trip measurement is taken about 3 seconds after a connection resumes, then with every ping (every 30 seconds when idle).
- Each CLI call costs about 9 ms to start.
Sleep and pause
AWP is built for sleep. Messages queue on disk and are kept for 7 days (outbox_ttl). While there is unfinished business, the awake side keeps retrying, at most 60 seconds apart. Nothing is lost or duplicated when the connection resumes. But there is one hard limit.
A paused sandbox cannot be dialed over tailcat. Its tailcat listener waits on a DERP relay connection, and a frozen process cannot answer, so a connection attempt cannot wake it. This is tracked in issue #4.
Until there is a wake-on-connect binding (for example, a WebSocket through the sandbox’s HTTP URL), work around it:
- Let the sleepy side dial out. If B may pause and A is always on, have B run
awp connect <A's address>, not the other way round. When B wakes, it resumes the connection itself. B’s dial-back address also lets A reconnect while B is awake. - Keep the sandbox awake while it has work. A long-running command keeps most sandboxes from pausing. For example, run a background loop through your platform’s exec API for the duration of the task.
- Don’t block on a sleeping peer. Use
awp waitwith a timeout, and treat exit code2as “nothing yet”, not as failure.
Reconnecting after a restart
Your tailcat address survives restarts. The keys and relay region are in tailcat.json in the home. A sandbox restored from a snapshot is back at the same address, and peers with unfinished business reconnect on their own.
If you copy a home to a new machine, you copy its identity too. Don’t run two daemons with the same home at once: both would claim the same key.
Credentials between machines
A lesson from running agents across sandboxes: some harness logins do not survive being copied.
- Codex, ChatGPT sign-in. It uses rotating refresh tokens. Copying
~/.codex/auth.jsonto a second machine makes the two copies conflict, and one of them gets logged out. Log in on each machine, or use an API key.
Checklist for a new machine
# install and wire into harnesses
curl -fsSL https://raw.githubusercontent.com/agentwireprotocol/awp/main/install.sh | sh
# bring it up with a clear name, and publish presence for awp web
awp up --name codex@build-box --about "Release builds" --presence
# connect to the always-on side
awp connect <address>
awp peers--name and --about are remembered. --presence is not, and like the others it applies only when awp up starts the daemon. To publish presence on every start, set "presence": true in config.json. See Presence.
Troubleshooting
| symptom | check |
|---|---|
awp up shows no address | tailcat is still starting; wait a few seconds, then awp status |
peer stuck in reconnecting | is the other side paused? awp web shows “can’t reach” and the dial error |
| messages not reaching the model | are hooks installed? awp bootstrap --list; otherwise awp tail --once |
| an old version keeps coming back | restart harness sessions; an old MCP server restarts an old daemon |
| anything else | ~/.awp/daemon.log; restart with awp down then awp daemon --trace to log every protocol line |