Writing
Notes from production: data engineering, generative AI, machine learning, and engineering in regulated finance. Subscribe via RSS.
2026
Financial machine learning from zero, part 2: tensors, time-ordered batching, and why leakage hides in .shift()
The second post in the financial machine learning series, and the first with code. It turns a Freddie Mac-style mortgage panel into PyTorch tensors one named step at a time, then breaks the pipeline on purpose with eight one-line mistakes and measures each one. Some push validation PR-AUC from 0.31 to 0.96. Others barely move it, which is worse, because nothing looks wrong. It ends with the checks that catch every one of them before a reviewer does.
machine-learning27 min read
AWS security for data engineers: seventeen recipes, worked through
Most data engineers meet AWS security as a wall of AccessDenied errors and a bucket someone made public. This post is the walkthrough I wish I had been given early: seventeen short recipes in the order I would apply them to a new account. Account-wide guardrails first (public access block, CloudTrail, default encryption), then identity (roles instead of keys, MFA, stale-key hunting), then permissions you generate and simulate instead of guess, permissions boundaries and SCPs so a team can self-serve safely, Session Manager and Secrets Manager for the machines, and finally the bucket policies, cross-account roles, and CloudFront setup that keep data buckets private. Every step is an AWS CLI command you can run in a sandbox account and clean up afterwards.
data-engineering37 min read
How to read the Spark UI
The Spark UI is the most underused tool in Spark. Most engineers learn to write DataFrame code long before they learn to read what Spark did with it. This is the walkthrough I give people joining my team: the five screens that matter — Jobs, the DAG, SQL, Stages, Storage — what each one is telling you, and the order I check them in when a job misbehaves. Every screenshot comes from a small retail join you can reproduce on a laptop.
data-engineering15 min read
Financial machine learning from zero, part 1: the vocabulary, the traps, and the metrics that matter
The first post in a beginner-friendly series on machine learning in finance, built around the problems banks actually run models on: credit scoring, fraud detection, and anti-money-laundering analytics, with market prediction as a supporting track. No code yet. Instead, the things I wish every newcomer knew before training a model on financial data: why it breaks the textbook assumptions, the vocabulary in every model notebook, the mistakes that make models look brilliant in development and fail in production, and a reference table of the metrics that matter, with which direction is good for each.
machine-learning22 min read
How granular should a pipeline task be?
A pipeline with one 300-line task is undebuggable; a pipeline with a task per pandas call is unreadable. The rule I use to cut between them: put a task boundary where you would stop and inspect the data, or where you would want a restart to begin — and nowhere else.
data-engineering11 min read
Adaptive AI agents: three layers of memory, from a Neo4j graph to Qwen's weights
An agent that solves a problem on Monday and rediscovers the same fix on Tuesday is burning money to stand still. This post works through what adaptation actually means — the agent loop, the four kinds of agent memory, and the three places an agent can improve — and a companion mini-project that builds all of it with Python, Neo4j, and a QLoRA fine-tune of Qwen.
genai10 min read
The Verhoeff algorithm: validating Aadhaar numbers at the ingestion edge
Every Aadhaar number carries its own tamper-evidence: the 12th digit is a Verhoeff check digit, computed from group theory rather than modular arithmetic. That makes it something a data engineer can exploit — a free, offline data-quality gate that catches typos, OCR misreads, and truncated numbers before they poison a customer dimension.
regulated-finance5 min read
TRY…CATCH, ROLLBACK, THROW: the case for explicit error handling in T-SQL
Plenty of warehouse procs rely on SET XACT_ABORT ON to clean up after failures — a one-liner most of the team has cargo-culted without knowing what it does. My position: make the error path explicit. TRY…CATCH, a visible ROLLBACK, and a THROW form a contract any reviewer can read — with XACT_ABORT kept underneath as the safety net it actually is.
data-engineering8 min read
Metadata-driven ETL: how far should configuration go?
Spec-driven pipelines are the right instinct — until the spec language grows conditionals and becomes a worse programming language. Where the line sits between metadata and code, why load patterns should be a closed set, and the smells that say your platform has crossed it.
data-engineering7 min read
Explaining the extra 2%: performance attribution as an Airflow pipeline
Fund managers live and die by whether they beat their benchmark — and by being able to say why. The math behind that answer is a clean three-factor decomposition, and it turns out to be a near-perfect fit for a pandas + Airflow pipeline with dynamic task mapping across dimensions.
regulated-finance7 min read
One DAG, every environment: the environment-aware DAG factory
How a small decorator around Airflow's @dag plus a tiered config object lets the same DAG definition run in dev, sit, uat, and prod — with no per-environment copies, no config drift, and promotion as a one-line gate.
data-engineering5 min read
The scheduler inside the scheduler: an Airflow anti-pattern and its paved-road fix
Why the controller-DAG-plus-generic-executor pattern hurts more than it helps, and how spec-driven DAG factories keep jobs as config without giving up per-job scheduling, backfills, and observability.
data-engineering7 min read
Migrating from enterprise schedulers to Apache Airflow: patterns that hold up
A pattern guide for moving batch workloads off Control-M or Autosys onto Airflow: inventory first, standardize the DAG template, and reconcile before every cutover.
data-engineering5 min read