Production Monitoring & Alerting
Database Management
This is the detailed technical record in the same style as the replication project's writeup: every command and query actually run, every issue actually hit, and how each was actually resolved — including a real cascading connection-exhaustion incident during the final verification step, kept in full because it's stronger evidence than a clean scripted test would have been.
1. Why this, and why now
Distinct from the existing Metabase setup: Metabase answers business
questions about the data in the tables (open vs closed jobs); this
answers operational questions about the database server itself — the
question a DBA/SRE gets paged for. Job-market research behind
db-management-track-planning.md found "production monitoring/alerting"
named explicitly in DBA/SRE postings as distinct from BI work. Uptime
Kuma already covers "is the container running"; it can't say why
something is slow or warn before it goes down — that gap is what this
closes.
2. Architecture
node-exporter (host CPU/mem/disk) ──┐
├─► Prometheus ──► Grafana ──► Discord
mysqld-exporter (MySQL internals) ──┘ │
▼
alert rules (visible on
Prometheus's own page,
but don't notify by
themselves — see §5)
Two exporters expose metrics over HTTP; Prometheus pulls (scrapes) both every 15 seconds and evaluates alert thresholds continuously; Grafana queries Prometheus for dashboards and separately runs its own alert-evaluation engine, which is the thing that actually notifies Discord — not Prometheus itself.
3. Part 1 — Standing up the exporters
3.1 Least-privilege MySQL user for the exporter
-- sql/create_exporter_user.sql, run as root on swe-2
CREATE USER IF NOT EXISTS 'mysqld_exporter'@'%' IDENTIFIED BY '<generated>';
GRANT PROCESS, REPLICATION CLIENT ON *.* TO 'mysqld_exporter'@'%';
GRANT SELECT ON performance_schema.* TO 'mysqld_exporter'@'%';
FLUSH PRIVILEGES;
mysql -uroot -p < sql/create_exporter_user.sql
Deliberately minimal: PROCESS (for SHOW PROCESSLIST/InnoDB status
metrics), REPLICATION CLIENT (for SHOW SLAVE STATUS — a no-op at the
time, since replication didn't exist yet, but collected in advance for
Project A), SELECT on performance_schema only — it cannot read or
write any actual application data, so a leaked exporter credential is a
non-event rather than an incident.
cp mysqld-exporter/.my.cnf.example mysqld-exporter/.my.cnf
# edit the password to match
3.2 Bring up the overlay
# add TAILSCALE_IP and GRAFANA_ADMIN_PASSWORD to /home/swe/stack/.env first
docker compose -f docker-compose.yml -f docker-compose.monitoring.yml up -d
Using -f twice (an overlay, not editing the base compose file directly)
keeps the diff of "what did monitoring add" contained to one file, while
still sharing the same Compose project/network so mysqld-exporter can
reach mysql by service name like every other container.
3.3 Issue: node-exporter unreachable under network_mode: host — hairpin NAT
Planned to run node-exporter with network_mode: host for accurate
host-NIC visibility (tailscale0, etc.), bound to
--web.listen-address=${TAILSCALE_IP}:9100. It came up and served metrics
locally, but Prometheus (on the regular bridge network) couldn't reach it:
context deadline exceeded
against 100.124.72.42:9100. Root cause: the same hairpin-NAT
limitation already hit and documented for Uptime Kuma during the original
server build — a container on the bridge network can't reliably reach
back through its own host's Tailscale IP. Pointing the scrape target at
the literal IP didn't help, same reason.
Fix: dropped network_mode: host entirely, ran node-exporter with
standard bridge networking plus bind-mounted /proc, /sys, /
(--path.procfs, --path.sysfs, --path.rootfs), so Prometheus reaches
it by Compose service name like every other exporter.
Tradeoff accepted: node_network_* metrics now reflect the
container's own virtual interface, not the host's real NICs — CPU/mem/disk
stay fully accurate since those come straight from the mounted /proc/
/sys. Per-interface network monitoring wasn't a stated goal, so this was
an acceptable trade rather than something worth chasing further at the
time.
3.4 Verify each layer bottom-up before touching Grafana
curl http://100.124.72.42:9104/metrics | grep mysql_up
# mysql_up 1 -> exporter's auth against MySQL worked
curl http://100.124.72.42:9100/metrics | grep node_uname_info
curl http://100.124.72.42:9090/-/healthy
Then Prometheus's own targets page
(http://100.124.72.42:9090/targets) — confirms node and mysql show
UP from inside Prometheus's container, not just reachable by curl
from outside (a target can be curl-reachable but still fail from inside
Prometheus if the service name or network is wrong).
4. Part 2 — Grafana: dashboards and the alerting trap
4.1 Dashboards
Imported two community dashboards by ID rather than building from scratch: 1860 (Node Exporter Full) and 7362 (MySQL Overview) — Dashboards → New → Import, Prometheus datasource picked explicitly during import (a panel stuck on "No data" is almost always the datasource variable not being pointed at your one datasource during import).
4.2 Issue: Grafana's alert list lies by omission
Once a Discord contact point existed and one webhook test succeeded, the
first end-to-end trip of MysqlTooManyConnections showed the rule as
Firing right there on Grafana's Alert rules page. No Discord message
arrived.
Root cause: Grafana's Alert rules page also lists Prometheus's own
prometheus/alerts.yml rules, grouped under a separate "Mimir/Cortex/
Loki" section rather than under "Grafana." They show real Firing/Pending
state — because Prometheus itself evaluates and fires them — but that's
read-only visibility into Prometheus's own alerting, with zero
notification channel attached. Every visible signal (state, color,
timestamp) looks identical to a real, wired-up alert right up until you
check why Discord never got pinged.
Fix: rebuilt all six thresholds as actual Grafana-managed rules
(Alerting → Alert rules → New alert rule → Grafana-managed, not
Data-source-managed), each with its contact point selected directly.
prometheus/alerts.yml stays in the repo as the portable,
version-controlled definition of what the thresholds are — just not the
thing that notifies anyone by itself.
4.3 Issue: /-/reload refused
curl -X POST http://100.124.72.42:9090/-/reload
# "Lifecycle API is not enabled"
Prometheus's lifecycle API is off by default. Added
--web.enable-lifecycle to the compose command; until redeployed with
that flag, a config change needed docker restart prometheus instead.
4.4 Issue: fresh dashboard showed "No data" with a healthy target
After fixing §3.3, the Node Exporter Full dashboard's Job/Nodename/
Instance template variables at the top were still cached as None from
before Prometheus had any node_uname_info series to populate them from.
A full page reload (not waiting for auto-refresh) re-resolved them once
real data existed. Worth checking those dropdowns first any time a
freshly-imported community dashboard shows "No data" everywhere.
5. Alert rules, as deployed
-- example threshold logic, MysqlTooManyConnections
mysql_global_status_threads_connected / mysql_global_variables_max_connections > 0.8
| Rule | Condition | For | Fires on |
|---|---|---|---|
HostHighCpu | CPU busy > 85% | 10m | host |
HostLowDisk | / free space < 10% | 5m | host |
HostLowMemory | available memory < 10% | 10m | host |
MysqlDown | mysql_up < 1 | 1m | mysql |
MysqlTooManyConnections | connections > 80% of max_connections | 5m | mysql — verified firing end-to-end |
MysqlSlowQueriesRising | slow-query rate > 1/s (5m avg) | 10m | mysql |
All six live under Grafana folder swe-2 → evaluation group
swe-2-alerts, routed to a Discord contact point reusing Uptime Kuma's
existing webhook (Alerting → Contact points → New → Discord integration,
same webhook URL, no new notification channel created).
6. Part 3 — Verifying it for real: a genuine incident, not a simulation
A monitoring stack that's never actually fired an alert hasn't proven
anything — same principle as "a backup you've never test-restored is a
backup you don't actually have." Deliberately tripped
MysqlTooManyConnections rather than trusting the config.
6.1 Attempt 1 — raced its own test window, no fire
-- 110 throwaway sessions, run concurrently
DO SLEEP(300);
Baseline was ~20 connections; 110 more clears the 80% threshold
comfortably. It didn't fire. The sessions' 5-minute sleep and the
rule's 5-minute for: window were sized almost identically — the
connections expired and the count dropped back down right as the window
would have closed. Lost the race by seconds.
6.2 Attempt 2 — fired, but caused a real cascading incident
Raised sleep duration to 10 minutes:
DO SLEEP(600);
This time the rule fired — but pushed connections to 152 against a
max_connections of 151, one over the hard cap, and something unplanned
happened:
- The
mysqld_exporter's own scrape connection started getting rejected too, surfacing asmysql_up 0— which tripped the separateMysqlDownrule as a genuine, if accidental, side effect. - Three real services sharing this MySQL instance (
budgetapp,taskapp,metabase_app) began failing to connect for real, confirmed via:
ERROR 1040 (HY000): Too many connections
SELECT * FROM information_schema.processlist;
-- visible reconnect storm from the affected app services
6.3 Recovery
MySQL reserves one connection slot for an admin/superuser login even at the hard connection cap:
docker exec -it mysql mysql -uroot -p
This still connected when nothing else could. From there:
SELECT id FROM information_schema.processlist WHERE user = 'mysqld_exporter' AND command = 'Sleep';
-- KILL <id>; for each throwaway session
Connection count dropped back to a normal ~40 within seconds. No data was lost; the outage was self-inflicted, self-diagnosed, and self-recovered in under a few minutes.
6.4 Attempt 3 — sized properly, clean fire, no side effects
-- 105 sessions against the same ~20 baseline, targeting ~125 total
-- (comfortably above the 121/80% line, comfortably below the 151 hard cap)
DO SLEEP(600);
Completed cleanly: Pending → Firing → Discord notification arrived, real services untouched throughout.
Kept as the headline evidence for the eventual portfolio entry, deliberately. A real cascading connection-exhaustion incident — including the moment it revealed that the monitoring system itself is a consumer of the resource it monitors, and can be starved by the same failure it's supposed to detect — is more honest, harder-to-fake evidence of understanding than a test that went perfectly on the first try.
7. Other real issues, quick reference
Grafana port collision. task-frontend already owned 3003. Checked
via docker compose ps and ss -tlnp (the latter catches host-level
listeners the former alone won't show) and moved Grafana to 3006, the
next free slot after bookstack=3004, budget-frontend=3005.
8. Issues encountered — summary table
| # | Issue | Root cause | Fix |
|---|---|---|---|
| 1 | Grafana port collision | task-frontend already on 3003 | Moved Grafana to 3006 after checking docker compose ps + ss -tlnp |
| 2 | node-exporter unreachable from Prometheus | Hairpin NAT — a container can't reach back through its own host's Tailscale IP | Dropped network_mode: host, used bridge networking + bind-mounted /proc//sys// |
| 3 | Grafana alert showed "Firing" but Discord never pinged | Grafana's Alert page also lists Prometheus's own rules (read-only, no notification channel) — looks identical to a real one | Rebuilt all 6 thresholds as actual Grafana-managed rules with a contact point attached |
| 4 | /-/reload refused | Prometheus's lifecycle API off by default | Added --web.enable-lifecycle; used docker restart until redeployed |
| 5 | Fresh dashboard showed "No data" | Template variables cached as None from before real data existed | Full page reload, not auto-refresh |
| 6 | Connection-spike test 1 didn't fire | Sleep duration matched the alert's for: window almost exactly — lost the race | Raised sleep duration comfortably past the window |
| 7 | Connection-spike test 2 caused a real outage | Pushed connections 1 over the hard max_connections cap, starving the exporter's own connection and 3 real apps | Recovered via MySQL's reserved admin connection, KILLed the throwaway sessions |
| 8 | Test 3 needed proper sizing | Confirmed the actual fix: size a connection-spike test against real headroom under the hard cap, not just above the alert threshold | 105 sessions targeting ~125 total, comfortably under 151 |
9. Lessons learned
- A monitoring stack that's never fired an alert hasn't proven anything. The verification step here wasn't optional polish — it's what actually surfaced the Grafana-vs-Prometheus alert-list trap (§4.2), which would otherwise have looked fully working while being completely silent.
- Size a load test against the hard resource cap, not just the alert threshold. Sizing "enough to clear 80%" without checking distance to 100% is how a verification test becomes a real incident.
- The thing doing the monitoring is itself a consumer of the resource
being monitored.
mysqld_exporter's own connection got starved by the same connection exhaustion it exists to detect — worth designing around (e.g., a reserved connection class) if this were a production system rather than a single-host demo. - A visible "Firing" state is not the same as "someone gets notified." Any alerting stack with more than one evaluation engine in the path (here: Prometheus and Grafana both capable of evaluating the same rule file) needs an explicit check that the one actually wired to notifications is the one being trusted.
- MySQL's reserved admin connection is the real emergency escape hatch during a connection-exhaustion incident — worth knowing it exists and confirming it works before it's needed for real.
- A real, self-inflicted incident recovered live is better portfolio material than a clean simulated test — it demonstrates diagnosis and recovery under actual (if self-caused) pressure, not just that a config file was written correctly.