Operations Runbooks
Step-by-step guides for common operational scenarios. Each runbook covers diagnosis, resolution, and prevention.
Agent Not Connecting
When a mezd fails to register or loses its tunnel:
- Check agent logs for connection errors:Check agent logsbash
journalctl -u mezd -n 100 --no-pager - Verify the join token has not expired:List active tokensbash
mezctl tokens ls - Check certificate expiry on the agent. The agent's X.509 certificate lives under
$MEZITE_DATA_DIR/agent/x509.pem(default/var/lib/mezite/agent/x509.pem):Inspect agent certificatebashopenssl x509 -in /var/lib/mezite/agent/x509.pem -noout -dates - Check the reverse tunnel — verify the agent can reach port 3024 on the proxy:Test tunnel connectivitybash
curl -v telnet://mezite.example.com:3024 - Check network / firewall rules — ensure ports 3024 and 3025 are open from the agent to the Mezite server.
Backup and Restore
All Mezite state — users, roles, tokens, CA private keys (encrypted at rest with ca_key_passphrase), audit events, session recordings metadata — lives in a single database. Regular backups of that database, plus the recording storage backend, are sufficient to recover the cluster.
- PostgreSQL backup with pg_dump:Backup PostgreSQLbash
pg_dump -h localhost -U mezite -d mezite -F c -f mezite-backup-$(date +%Y%m%d).dump - SQLite backup — the database file lives at
$data_dir/mezhub.db(default/var/lib/mezite/mezhub.db). Use SQLite's online backup command so writes in flight don't corrupt the copy:Backup SQLitebashsqlite3 /var/lib/mezite/mezhub.db ".backup '/backups/mezhub-$(date +%Y%m%d).db'" - Restore (PostgreSQL):Restore PostgreSQLbash
pg_restore -h localhost -U mezite -d mezite --clean --if-exists mezite-backup-20260324.dump - Restore (SQLite): stop
mezhub, copy the backup file into place, and startmezhubagain.Restore SQLitebashsudo systemctl stop mezhub sudo cp /backups/mezhub-20260324.db /var/lib/mezite/mezhub.db # Match the ownership mezhub runs as. The shipped unit file sets no # User=, so mezhub runs as root unless you have added one; if you did, # chown to that account instead. sudo chown root:root /var/lib/mezite/mezhub.db sudo chmod 600 /var/lib/mezite/mezhub.db sudo systemctl start mezhub - Don't forget the recording bucket / dir — if you use S3 recording, snapshot the bucket. If you use local recording, back up
$data_dir/recordings. - Verify the backup by restoring to a test instance before relying on it for disaster recovery. You will also need the original
MEZITE_CA_KEY_PASSPHRASEon the restored instance to decrypt CA private keys.
CA Certificate Expiry
CA certificates and private keys are stored in the database (encrypted at rest), not on disk. Monitor expiry with mezctl ca status and start a rotation when the remaining lifetime drops below 90 days. Rotation is a multi-phase state machine with four phases — init → update_clients → update_servers → complete. mezctl ca status reports both the phase (rotation_phase) and the rotation state (rotation_state: standby / in_progress/ rollback). A fresh rotation lands in init, so three advance calls are needed to reach complete. The state flips back to standby only when the final advance intocomplete runs.
- Check rotation state and expiry for every CA type (host, user, spiffe):Check CA statusbash
mezctl ca status - Start rotation of the host CA — this mints a new key, makes it the active signing key, keeps the old key in the trust bundle, and leaves rotation in the
initphase. Agents re-download the two-key bundle on their next reconnect.Rotate host CAbashmezctl ca rotate --type=host - Advance through the rotation phases. Rotation starts at
init, so reachingcompletetakesthree advances, not two. Checkmezctl ca statusbetween each one and only advance when the step's condition actually holds.Advance host CA rotationbash# 1. init -> update_clients (once clients hold the new trust bundle) mezctl ca advance --type=host # 2. update_clients -> update_servers (once servers have reissued their certs) mezctl ca advance --type=host # 3. update_servers -> complete (finalize, drop old CA — irreversible) mezctl ca advance --type=host - Rotate the user CA the same way. Repeat for
--type=spiffeif you use workload identity.Rotate user CAbashmezctl ca rotate --type=user mezctl ca advance --type=user mezctl ca advance --type=user mezctl ca advance --type=user - Rollback is only valid while
rotation_state = in_progress. If something looks wrong before you run the finaladvance, abort withmezctl ca rollback --type=<host|user|spiffe>to restore the previous CA. Once the finaladvanceintocompleteruns, the previous CA's keys are deleted and the rotation cannot be rolled back.
Database Performance
Applies to PostgreSQL deployments. SQLite is single-process and tuning is limited to disk performance.
- Identify slow queries:Find slow queriessql
SELECT pid, now() - pg_stat_activity.query_start AS duration, query FROM pg_stat_activity WHERE state != 'idle' ORDER BY duration DESC LIMIT 10; - Run VACUUM:Vacuum and analyzesql
VACUUM ANALYZE; - Check connection pool usage — ensure
max_connectionsin PostgreSQL is set higher than the sum of all Mezite instances' pool sizes. - Lock contention — check for blocked queries:Check for lock contentionsql
SELECT blocked.pid, blocked.query, blocking.pid AS blocking_pid, blocking.query AS blocking_query FROM pg_stat_activity blocked JOIN pg_locks bl ON bl.pid = blocked.pid JOIN pg_locks bk ON bk.locktype = bl.locktype AND bk.relation = bl.relation AND bk.pid != bl.pid JOIN pg_stat_activity blocking ON blocking.pid = bk.pid WHERE NOT bl.granted;
Security Hardening Checklist
- TLS: Ensure all Mezite ports (3025, 3080, 3023, 3024) use TLS. Terminate TLS at the proxy where possible to preserve mutual TLS; when an upstream load balancer must terminate TLS, set
auth.grpc_allow_http: true(orMEZITE_AUTH_H2C=true) and pin a trusted-IP header / PROXY-protocol source viaproxy.trusted_ip_headerorproxy.proxy_protocol_trusted_cidrs. - Authentication: Require MFA for all human users. Keep session lifetimes short by setting
max_session_ttlon every role — a role that omits it contributes no cap of its own, and the shipped roles set 12h (admin), 8h (editor,ssh-access) and 4h (viewer). Add a cluster-wide ceiling withproxy.max_session_durationso a role with no TTL cannot run unbounded. Prefer SSO connectors (OIDC, SAML, LDAP, GitHub) over local passwords. - Authorization: Apply least-privilege roles. Restrict SSH logins by node label. Require access requests for privileged roles.
- Network: Restrict the auth port (3025) to internal networks. Expose only the proxy HTTPS port (3080) and the SSH port (3023) publicly. Use firewall rules to limit agent reverse-tunnel access (3024) to known agent subnets where possible.
- Audit: Enable session recording (
recording.backend). Forward audit events to an external SIEM with the webhook or file sink (MEZITE_AUDIT_SINK_WEBHOOK_URL/MEZITE_AUDIT_SINK_FILE_PATH). Audit events are notautomatically pruned — operators must manage retention on the database and external sinks themselves. - CA private keys: Always set
ca_key_passphrase(viaMEZITE_CA_KEY_PASSPHRASE) in production. Without it, CA signing keys and other sensitive fields are stored in plaintext in the database.
Upgrading
Upgrading mezhub means a short planned outage: a cluster runs one instance, so there is no second replica to carry traffic during the swap. Established SSH sessions are cut and agents reconnect on their own. Schedule accordingly.
- Pre-flight checks: back up the database (seeBackup and Restore) and confirm the running cluster is healthy.Pre-flightbash
# Verify the proxy is up curl -sf https://mezite.example.com:3080/healthz # Back up PostgreSQL (or SQLite — see Backup and Restore) pg_dump -h localhost -U mezite -d mezite -F c -f /backups/pre-upgrade-$(date +%Y%m%d).dump - Database migrations run automatically when
mezhubstarts — there is no separatemezhub migratesubcommand. Migrations are forward-only and idempotent; the firstmezhubinstance to start with the new binary applies them before serving traffic. - Restart mezhub: stop the old binary, put the new one in place, start it, and verify
/healthzand/readyzbefore declaring the upgrade done. A cluster runs a singlemezhubinstance — each agent's reverse tunnel lives in one process and there is no proxy-to-proxy forwarding — so this step is a brief outage, not a rolling restart. Agents reconnect automatically once the tunnel port is listening again. - Upgrade agents: Agents are backward-compatible with newer servers. Upgrade them after the server rollout is complete.Upgrade agent binarybash
sudo systemctl stop mezd sudo cp mezd-new /usr/local/bin/mezd sudo systemctl start mezd - Rollback: If issues arise, stop the new instances, restore the database backup, and restart with the previous binary version. Because migrations are forward-only, a binary that pre-dates a migration cannot run against the upgraded schema — pin a tested previous version before relying on it for rollback.