Topic
Big data testing & data quality
Reconciliation, validation and monitoring that make large-scale data pipelines trustworthy.
Case studies
What if most manual data checks could validate themselves?
An LLM-assisted validation framework for big data pipelines that automated more than 70% of manual ETL checks.
How do you prove that every row that left the source arrived in the lakehouse, correctly transformed?
A reference design for reconciling and validating data across CDC, Kafka, PySpark and a Trino-queried lakehouse, from row counts to keyed diffs and anomaly detection.
Articles
Validating CDC pipelines
How to test incremental and change-data-capture pipelines for ordering, late and duplicate events, deletes, idempotent MERGE, watermarks and reconciliation windows.
From expectations to a production data-quality system
Rule catalogues, severity and ownership, where checks run, simple anomaly baselines for volume, freshness and distribution, and alerting that people still read.
Great Expectations: A Data Quality Framework, Part 3 (opens Medium in a new tab)
A practical introduction to Great Expectations, setting up a project and writing expectations to validate data automatically.
Unlocking Insights: The Best Practices and Frameworks for Data Quality, Part 2 (opens Medium in a new tab)
Ten practices for keeping data quality high, and an introduction to the Great Expectations and PyDeequ frameworks for automating checks.
Unlocking Insights: The Art of Big Data Testing and Data Quality, Part 1 (opens Medium in a new tab)
Why big data testing matters, what data quality assurance covers, the main challenges, and the practices that make big data testing effective.
Work with me
Working on a problem like the ones above? I am glad to compare notes.