Issue No. 01 · Data EngineeringAvailable for Work
Data Analyst & Database Engineer

JOSUE

GARCIA

Building data systems end-to-end — normalized schemas, automated ingestion pipelines, slow-query diagnosis, and index tuning. 6 chained projects. One connected dataset. Real results.

Stack

MySQL · Apache Airflow · MinIO · Python · Docker

Live at

swe-2.tail174d56.ts.net

Zero managed cloud databases • 1M+ synthetic rows benchmarked • Tailscale-only admin • self-hosted on a home Ubuntu server • Zero managed cloud databases • 1M+ synthetic rows benchmarked • Tailscale-only admin • self-hosted on a home Ubuntu server • Zero managed cloud databases • 1M+ synthetic rows benchmarked • Tailscale-only admin • self-hosted on a home Ubuntu server • Zero managed cloud databases • 1M+ synthetic rows benchmarked • Tailscale-only admin • self-hosted on a home Ubuntu server •
Projects

Each project feeds the next. Chained output, one connected dataset.

Three Projects.
One Dataset.

01
01Done

The complete technical record, in the order it actually happened: GTID-based replication from scratch, a first failover attempt that surfaced real gaps and was deliberately paused, then a completed drill and a genuine multi-app production incident. Every command below actually ran.

MySQLGTID ReplicationHigh AvailabilityFailover EngineeringDockerDocker Compose
Read Full Case Study →
03
03Not started

Add role-based access control, PII column masking, and audit logging to an existing schema — and specifically, instrument the `portfolio_admin` write-path built earlier this session (2026-08-18) with real audit logging, so the project is partly about auditing this portfolio's own system rather than only a synthetic exercise.

Read Full Case Study →
04
04Not started

Turn Project 3's manual `mysqldump --single-transaction | gzip` runbook into a scheduled, automated pipeline with retention policy and — critically — automated, *tested* restores, not just backups that are assumed to work.

Read Full Case Study →
05
05Done

This is the detailed technical record in the same style as the replication project's writeup: every command and query actually run, every issue actually hit, and how each was actually resolved — including a real cascading connection-exhaustion incident during the final verification step, kept in full because it's stronger evidence than a clean scripted test would have been.

Read Full Case Study →
06
26.8ms → 1.6msQuery time after composite index
06Not started

Load the schema with a million synthetic rows and run a full slow-query diagnosis and tuning case study.

EXPLAINIndexingBenchmarkingBackup/restore
Read Full Case Study →
Skills & Tools
Technologies used across every project
Languages
BashSQL
Databases
MySQLEXPLAINIndexing
Infra
DockerDocker ComposeLinux AdministrationBackup/restore
Analysis
GTID ReplicationHigh AvailabilityFailover EngineeringPrometheusGrafanaIncident ResponseBenchmarking
Writing

Latest Posts

All Posts →
ContactOpen to Work · 2026

Let’s Talk
Data.

Open to data analyst, database engineer, and data engineering roles. I prefer async-first communication — GitHub issues, email, or LinkedIn DMs.

Send Email →