Done

Production Monitoring & Alerting

Database Management

This is the detailed technical record in the same style as the replication project's writeup: every command and query actually run, every issue actually hit, and how each was actually resolved — including a real cascading connection-exhaustion incident during the final verification step, kept in full because it's stronger evidence than a clean scripted test would have been.

DocsLast updated August 23, 2026

1. Why this, and why now

Distinct from the existing Metabase setup: Metabase answers business questions about the data in the tables (open vs closed jobs); this answers operational questions about the database server itself — the question a DBA/SRE gets paged for. Job-market research behind db-management-track-planning.md found "production monitoring/alerting" named explicitly in DBA/SRE postings as distinct from BI work. Uptime Kuma already covers "is the container running"; it can't say why something is slow or warn before it goes down — that gap is what this closes.

2. Architecture

node-exporter (host CPU/mem/disk)  ──┐
                                      ├─► Prometheus ──► Grafana ──► Discord
mysqld-exporter (MySQL internals)  ──┘        │
                                               ▼
                                        alert rules (visible on
                                        Prometheus's own page,
                                        but don't notify by
                                        themselves — see §5)

Two exporters expose metrics over HTTP; Prometheus pulls (scrapes) both every 15 seconds and evaluates alert thresholds continuously; Grafana queries Prometheus for dashboards and separately runs its own alert-evaluation engine, which is the thing that actually notifies Discord — not Prometheus itself.


3. Part 1 — Standing up the exporters

3.1 Least-privilege MySQL user for the exporter

-- sql/create_exporter_user.sql, run as root on swe-2
CREATE USER IF NOT EXISTS 'mysqld_exporter'@'%' IDENTIFIED BY '<generated>';
GRANT PROCESS, REPLICATION CLIENT ON *.* TO 'mysqld_exporter'@'%';
GRANT SELECT ON performance_schema.* TO 'mysqld_exporter'@'%';
FLUSH PRIVILEGES;
mysql -uroot -p < sql/create_exporter_user.sql

Deliberately minimal: PROCESS (for SHOW PROCESSLIST/InnoDB status metrics), REPLICATION CLIENT (for SHOW SLAVE STATUS — a no-op at the time, since replication didn't exist yet, but collected in advance for Project A), SELECT on performance_schema only — it cannot read or write any actual application data, so a leaked exporter credential is a non-event rather than an incident.

cp mysqld-exporter/.my.cnf.example mysqld-exporter/.my.cnf
# edit the password to match

3.2 Bring up the overlay

# add TAILSCALE_IP and GRAFANA_ADMIN_PASSWORD to /home/swe/stack/.env first
docker compose -f docker-compose.yml -f docker-compose.monitoring.yml up -d

Using -f twice (an overlay, not editing the base compose file directly) keeps the diff of "what did monitoring add" contained to one file, while still sharing the same Compose project/network so mysqld-exporter can reach mysql by service name like every other container.

3.3 Issue: node-exporter unreachable under network_mode: host — hairpin NAT

Planned to run node-exporter with network_mode: host for accurate host-NIC visibility (tailscale0, etc.), bound to --web.listen-address=${TAILSCALE_IP}:9100. It came up and served metrics locally, but Prometheus (on the regular bridge network) couldn't reach it:

context deadline exceeded

against 100.124.72.42:9100. Root cause: the same hairpin-NAT limitation already hit and documented for Uptime Kuma during the original server build — a container on the bridge network can't reliably reach back through its own host's Tailscale IP. Pointing the scrape target at the literal IP didn't help, same reason.

Fix: dropped network_mode: host entirely, ran node-exporter with standard bridge networking plus bind-mounted /proc, /sys, / (--path.procfs, --path.sysfs, --path.rootfs), so Prometheus reaches it by Compose service name like every other exporter.

Tradeoff accepted: node_network_* metrics now reflect the container's own virtual interface, not the host's real NICs — CPU/mem/disk stay fully accurate since those come straight from the mounted /proc/ /sys. Per-interface network monitoring wasn't a stated goal, so this was an acceptable trade rather than something worth chasing further at the time.

3.4 Verify each layer bottom-up before touching Grafana

curl http://100.124.72.42:9104/metrics | grep mysql_up
# mysql_up 1   -> exporter's auth against MySQL worked

curl http://100.124.72.42:9100/metrics | grep node_uname_info
curl http://100.124.72.42:9090/-/healthy

Then Prometheus's own targets page (http://100.124.72.42:9090/targets) — confirms node and mysql show UP from inside Prometheus's container, not just reachable by curl from outside (a target can be curl-reachable but still fail from inside Prometheus if the service name or network is wrong).


4. Part 2 — Grafana: dashboards and the alerting trap

4.1 Dashboards

Imported two community dashboards by ID rather than building from scratch: 1860 (Node Exporter Full) and 7362 (MySQL Overview) — Dashboards → New → Import, Prometheus datasource picked explicitly during import (a panel stuck on "No data" is almost always the datasource variable not being pointed at your one datasource during import).

4.2 Issue: Grafana's alert list lies by omission

Once a Discord contact point existed and one webhook test succeeded, the first end-to-end trip of MysqlTooManyConnections showed the rule as Firing right there on Grafana's Alert rules page. No Discord message arrived.

Root cause: Grafana's Alert rules page also lists Prometheus's own prometheus/alerts.yml rules, grouped under a separate "Mimir/Cortex/ Loki" section rather than under "Grafana." They show real Firing/Pending state — because Prometheus itself evaluates and fires them — but that's read-only visibility into Prometheus's own alerting, with zero notification channel attached. Every visible signal (state, color, timestamp) looks identical to a real, wired-up alert right up until you check why Discord never got pinged.

Fix: rebuilt all six thresholds as actual Grafana-managed rules (Alerting → Alert rules → New alert rule → Grafana-managed, not Data-source-managed), each with its contact point selected directly. prometheus/alerts.yml stays in the repo as the portable, version-controlled definition of what the thresholds are — just not the thing that notifies anyone by itself.

4.3 Issue: /-/reload refused

curl -X POST http://100.124.72.42:9090/-/reload
# "Lifecycle API is not enabled"

Prometheus's lifecycle API is off by default. Added --web.enable-lifecycle to the compose command; until redeployed with that flag, a config change needed docker restart prometheus instead.

4.4 Issue: fresh dashboard showed "No data" with a healthy target

After fixing §3.3, the Node Exporter Full dashboard's Job/Nodename/ Instance template variables at the top were still cached as None from before Prometheus had any node_uname_info series to populate them from. A full page reload (not waiting for auto-refresh) re-resolved them once real data existed. Worth checking those dropdowns first any time a freshly-imported community dashboard shows "No data" everywhere.


5. Alert rules, as deployed

-- example threshold logic, MysqlTooManyConnections
mysql_global_status_threads_connected / mysql_global_variables_max_connections > 0.8
RuleConditionForFires on
HostHighCpuCPU busy > 85%10mhost
HostLowDisk/ free space < 10%5mhost
HostLowMemoryavailable memory < 10%10mhost
MysqlDownmysql_up < 11mmysql
MysqlTooManyConnectionsconnections > 80% of max_connections5mmysql — verified firing end-to-end
MysqlSlowQueriesRisingslow-query rate > 1/s (5m avg)10mmysql

All six live under Grafana folder swe-2 → evaluation group swe-2-alerts, routed to a Discord contact point reusing Uptime Kuma's existing webhook (Alerting → Contact points → New → Discord integration, same webhook URL, no new notification channel created).


6. Part 3 — Verifying it for real: a genuine incident, not a simulation

A monitoring stack that's never actually fired an alert hasn't proven anything — same principle as "a backup you've never test-restored is a backup you don't actually have." Deliberately tripped MysqlTooManyConnections rather than trusting the config.

6.1 Attempt 1 — raced its own test window, no fire

-- 110 throwaway sessions, run concurrently
DO SLEEP(300);

Baseline was ~20 connections; 110 more clears the 80% threshold comfortably. It didn't fire. The sessions' 5-minute sleep and the rule's 5-minute for: window were sized almost identically — the connections expired and the count dropped back down right as the window would have closed. Lost the race by seconds.

6.2 Attempt 2 — fired, but caused a real cascading incident

Raised sleep duration to 10 minutes:

DO SLEEP(600);

This time the rule fired — but pushed connections to 152 against a max_connections of 151, one over the hard cap, and something unplanned happened:

  • The mysqld_exporter's own scrape connection started getting rejected   too, surfacing as mysql_up 0 — which tripped the separate MysqlDown   rule as a genuine, if accidental, side effect.
  • Three real services sharing this MySQL instance (budgetapp,   taskapp, metabase_app) began failing to connect for real, confirmed   via:
ERROR 1040 (HY000): Too many connections
SELECT * FROM information_schema.processlist;
-- visible reconnect storm from the affected app services

6.3 Recovery

MySQL reserves one connection slot for an admin/superuser login even at the hard connection cap:

docker exec -it mysql mysql -uroot -p

This still connected when nothing else could. From there:

SELECT id FROM information_schema.processlist WHERE user = 'mysqld_exporter' AND command = 'Sleep';
-- KILL <id>; for each throwaway session

Connection count dropped back to a normal ~40 within seconds. No data was lost; the outage was self-inflicted, self-diagnosed, and self-recovered in under a few minutes.

6.4 Attempt 3 — sized properly, clean fire, no side effects

-- 105 sessions against the same ~20 baseline, targeting ~125 total
-- (comfortably above the 121/80% line, comfortably below the 151 hard cap)
DO SLEEP(600);

Completed cleanly: Pending → Firing → Discord notification arrived, real services untouched throughout.

Kept as the headline evidence for the eventual portfolio entry, deliberately. A real cascading connection-exhaustion incident — including the moment it revealed that the monitoring system itself is a consumer of the resource it monitors, and can be starved by the same failure it's supposed to detect — is more honest, harder-to-fake evidence of understanding than a test that went perfectly on the first try.


7. Other real issues, quick reference

Grafana port collision. task-frontend already owned 3003. Checked via docker compose ps and ss -tlnp (the latter catches host-level listeners the former alone won't show) and moved Grafana to 3006, the next free slot after bookstack=3004, budget-frontend=3005.


8. Issues encountered — summary table

#IssueRoot causeFix
1Grafana port collisiontask-frontend already on 3003Moved Grafana to 3006 after checking docker compose ps + ss -tlnp
2node-exporter unreachable from PrometheusHairpin NAT — a container can't reach back through its own host's Tailscale IPDropped network_mode: host, used bridge networking + bind-mounted /proc//sys//
3Grafana alert showed "Firing" but Discord never pingedGrafana's Alert page also lists Prometheus's own rules (read-only, no notification channel) — looks identical to a real oneRebuilt all 6 thresholds as actual Grafana-managed rules with a contact point attached
4/-/reload refusedPrometheus's lifecycle API off by defaultAdded --web.enable-lifecycle; used docker restart until redeployed
5Fresh dashboard showed "No data"Template variables cached as None from before real data existedFull page reload, not auto-refresh
6Connection-spike test 1 didn't fireSleep duration matched the alert's for: window almost exactly — lost the raceRaised sleep duration comfortably past the window
7Connection-spike test 2 caused a real outagePushed connections 1 over the hard max_connections cap, starving the exporter's own connection and 3 real appsRecovered via MySQL's reserved admin connection, KILLed the throwaway sessions
8Test 3 needed proper sizingConfirmed the actual fix: size a connection-spike test against real headroom under the hard cap, not just above the alert threshold105 sessions targeting ~125 total, comfortably under 151

9. Lessons learned

  1. A monitoring stack that's never fired an alert hasn't proven    anything. The verification step here wasn't optional polish — it's    what actually surfaced the Grafana-vs-Prometheus alert-list trap    (§4.2), which would otherwise have looked fully working while being    completely silent.
  2. Size a load test against the hard resource cap, not just the alert    threshold. Sizing "enough to clear 80%" without checking distance to    100% is how a verification test becomes a real incident.
  3. The thing doing the monitoring is itself a consumer of the resource    being monitored. mysqld_exporter's own connection got starved by    the same connection exhaustion it exists to detect — worth designing    around (e.g., a reserved connection class) if this were a    production system rather than a single-host demo.
  4. A visible "Firing" state is not the same as "someone gets notified."    Any alerting stack with more than one evaluation engine in the path    (here: Prometheus and Grafana both capable of evaluating the same    rule file) needs an explicit check that the one actually wired to    notifications is the one being trusted.
  5. MySQL's reserved admin connection is the real emergency escape    hatch during a connection-exhaustion incident — worth knowing it    exists and confirming it works before it's needed for real.
  6. A real, self-inflicted incident recovered live is better portfolio    material than a clean simulated test — it demonstrates diagnosis and    recovery under actual (if self-caused) pressure, not just that a    config file was written correctly.