| ▲ | humpf 12 hours ago | |
Location: Kansas City, KS (US Central) Remote: Yes Willing to relocate: No Technologies: Python, web scraping and data extraction, ETL pipelines, Playwright, BeautifulSoup, Tesseract OCR, PostgreSQL, SQLite, pytest, GitHub Actions, Astro, Cloudflare Pages, C / ESP-IDF, NumPy Résumé/CV: https://braedenkeena.pages.dev Email: braeden@thekeenas.com I build data pipelines that know when they're lying to you. GAScout (https://gascout.pages.dev) — Georgia public-records pipeline. Four counties live, each a different ingestion problem resolved into one schema: a 9,915-page fixed-width mainframe PDF (409,142 records, zero dropped rows), a weekly XLSX roll, monthly text PDFs, and scanned sheriff's levy documents via OCR. Adding a county is a config file, not code. The other 155 are labeled planned rather than filled with invented numbers. FindStorage (https://findstorage.pages.dev) — national self-storage pricing tracker. 4,639 locations, 245,000+ price changes logged, running unattended on a daily schedule. Aborts rather than publishing when the store count moves more than it should, because the failure that costs you isn't a crash — it's a scraper quietly returning less after a site changes and nobody noticing for three weeks. Also built a WiFi CSI device-fingerprinting array on ESP32 — custom ESP-IDF firmware and a Python DSP pipeline, running as my apartment's security system. Cross-manufacturer separation holds at 11-15 sigma; same-model discrimination tops out near 77% and true clock twins are the open problem. Before this I ran a 500-acre motocross park in Northern California: acquired dormant land, re-entitled it, directed a 120-person race-day crew, ran the medical plan through a mid-event helicopter evacuation, and sold my stake. Self-taught, ~18 months in. Interested in data engineering, pipeline work, and anything where messy real-world sources have to become trustworthy structured data. | ||