User Tools

Site Tools


Writing /app/www/public/data/meta/the-lab/runbooks.meta failed
the-lab:runbooks

# Runbooks

Step-by-step for common problems.

## “All the sites are down”

1. Check pi4 (`.2`) — is it powered and on the network? (`ping 10.0.0.2`) 2. SSH in (`root`) and run `systemctl status caddy`. If it's not running: `systemctl start caddy`. 3. The health-check cron should already be restarting it; if not, check `/usr/local/bin/caddy-health-check.sh`. 4. If pi4 is up but sites still 502, check the individual service host (see “How it's wired”).

## “One site is down (502)”

1. Find which host serves it (see “Portals & sites” or `infra/inventory.yaml`). 2. SSH to that host and check its container/service:

  1. Docker: `docker ps -a | grep ` → `docker start ` if exited.
  2. systemd: `systemctl status ` → `systemctl restart `.

3. Check the edge log for the exact error: `journalctl -u caddy -n 50` on pi4.

## Edge failover (pi4 dead, LAN only)

- LAN clients should auto-fail to mini1 (`.4`) via dual-A DNS after the TTL expires. - Public traffic has no failover yet — it needs pi4 back.

## Reboot order (after a full power outage)

1. Router → wait for it to come up. 2. pi4 (`.2`) — DNS + edge first. 3. mini1 (`.4`) — DHCP + standby edge. 4. mini2 (`.5`), macpro (`.3`), windows (`.7`), pi5 (`.13`). 5. Family devices last.

## “A container won't start”

1. `docker logs ` — read the actual error. 2. Common causes: permission on a bind mount (chown to the container uid), a missing tmpfs dir (recreate + chown), or a config typo. 3. Log the fix to the Gotchas page so it's not re-solved twice.

the-lab/runbooks.txt · Last modified: by 127.0.0.1