Notes from building ETL
for messy agency data.
Why schema drift breaks more pipelines than bad data ever does
Most ETL failures aren't caused by messy values — they're caused by a column that quietly changed shape upstream. Here's how NNIPA detects drift before it reaches your canonical schema.
From the field
Engineering, governance, and data
Practical notes from building and deploying NNIPA across complex state and federal data environments.
PII masking at ingestion, not as an afterthought
What it actually takes to detect and mask sensitive fields the moment data lands — before anyone downstream can query it.
Semantic field mapping: how 'ssn_num' becomes SocialSecurityNumber
A look at the similarity scoring behind NNIPA's field mapping engine, and why fallback rules matter more than raw accuracy.
Unifying UI wage data from 50 state systems
How a DOL OIG pilot went from Oracle, SQL Server, and CSV exports to one governed schema in PostgreSQL.
OCR is easy. Structuring what it finds is the hard part
Turning a scanned PDF into rows and columns takes more than text extraction — notes from building NNIPA's unstructured data pipeline.
Prompt-driven ETL: what happens after you hit 'run'
A walk-through of how a plain-language prompt gets interpreted into schema detection, mapping, and load steps.
What we log for every field mapping, and why
Confidence scores, overrides, and full mapping history — the audit trail agencies actually ask for.
Stay informed
Get new posts by email
Engineering and governance notes, sent occasionally — no spam.
NNIPA
Building better data systems
See how NNIPA can help your organization turn messy agency data into governed, usable intelligence.