Remix.run Logo
humpf 12 hours ago

Location: Kansas City, KS (US Central)

Remote: Yes

Willing to relocate: No

Technologies: Python, web scraping and data extraction, ETL pipelines, Playwright, BeautifulSoup, Tesseract OCR, PostgreSQL, SQLite, pytest, GitHub Actions, Astro, Cloudflare Pages, C / ESP-IDF, NumPy

Résumé/CV: https://braedenkeena.pages.dev

Email: braeden@thekeenas.com

I build data pipelines that know when they're lying to you.

GAScout (https://gascout.pages.dev) — Georgia public-records pipeline. Four counties live, each a different ingestion problem resolved into one schema: a 9,915-page fixed-width mainframe PDF (409,142 records, zero dropped rows), a weekly XLSX roll, monthly text PDFs, and scanned sheriff's levy documents via OCR. Adding a county is a config file, not code. The other 155 are labeled planned rather than filled with invented numbers.

FindStorage (https://findstorage.pages.dev) — national self-storage pricing tracker. 4,639 locations, 245,000+ price changes logged, running unattended on a daily schedule. Aborts rather than publishing when the store count moves more than it should, because the failure that costs you isn't a crash — it's a scraper quietly returning less after a site changes and nobody noticing for three weeks.

Also built a WiFi CSI device-fingerprinting array on ESP32 — custom ESP-IDF firmware and a Python DSP pipeline, running as my apartment's security system. Cross-manufacturer separation holds at 11-15 sigma; same-model discrimination tops out near 77% and true clock twins are the open problem.

Before this I ran a 500-acre motocross park in Northern California: acquired dormant land, re-entitled it, directed a 120-person race-day crew, ran the medical plan through a mid-event helicopter evacuation, and sold my stake.

Self-taught, ~18 months in. Interested in data engineering, pipeline work, and anything where messy real-world sources have to become trustworthy structured data.