
How We Automated ETL Regression Testing at Scale for an Enterprise Bank - 200M Records Validated in Under 2 Hours
At a glance -
the essentials
One of the Top 5 Canadian Banks
- The bank's Equity Markets & Prime Brokerage Technology business ran roughly 800 batches in production. Whenever batch logic changed, the team needed to regression-test every affected ETL job to confirm production functionality wasn't disrupted - but with 800 batches and constant time pressure, there was no practical way to do this by hand. ETL batch configurations lived in database tables that weren't consistently promoted across environments, making structure and data comparison tedious, and schema or stored-procedure changes that weren't properly promoted to other environments caused recurring production errors.
- Use DQV to compare ETL job outputs after running them in both SIT-A (the development environment) and SIT-B (the production-mirror environment) - covering Data Load jobs (table-to-table), Data Extract jobs (table-to-file), and File Delivery jobs (file metadata: size, checksum, last-modified)
- Run periodic structure comparisons between SIT-A and SIT-B database schemas - tables, views, and stored procedure definitions - to catch definition drift in schema before it reached production
- Compare database metadata across 9 schemas (83,725 passed / 350 failed objects) in under 2 minutes
- Scale to a full 200-million-record data comparison across 79 tables in 1 hour 45 minutes that includes few larger BLOB/CLOB columns to compare
- Highlight differences at the field level, and integrated DQV into the DevOps pipeline to remove manual steps from the release process
- 20 ETL batches now covered by repeatable, automated regression testing instead of ad hoc manual checks
- Structure comparison across 9 schemas and tens of thousands of metadata objects (83,725 metadata objects passed, 350 failed) completes in under 2 minutes
- A full 200-million-record data comparison across 79 tables completes in 1 hour 45 minutes
- 200,178,235 records compared in total across all 79 tables
- 187,079 mismatched records identified - a 0.09% mismatch rate
- Field-level difference highlighting cuts the time needed to triage and fix issues
- DevOps pipeline integration removed manual intervention from routine regression runs
- Delivered on a fixed-bid, low-budget basis, with ongoing support from the DQV product team
From manual, time-crunched regression checks to a 200M-record validation in under 2 hours.
Kumaran partnered with the Equity Markets & Prime Brokerage Technology team at one of the top 5 Canadian banks to bring automated regression testing to roughly 800 production batches. A POC was executed for 20 ETL batches. Using DQV, the team compares ETL job outputs - table loads, file extracts, and file deliveries - between the SIT-A (development) and SIT-B (production-mirror) environments, and periodically checks the underlying database structures for drift. What once had no practical way to be tested at this scale now runs as a routine, DevOps-integrated check: a full 200-million-record comparison across 79 tables completes in 1 hour 45 minutes, with structure comparisons across 9 schemas finishing in under 2 minutes.
The Background
Equity Markets & Prime Brokerage Technology's roughly 800 production batches relied on ETL batch configurations stored in database tables that weren't consistently promoted across environments, and on schema and stored-procedure changes that weren't always propagated correctly either. Whenever batch logic changed, the team needed to regression-test the affected jobs to protect production functionality, but time pressure and a lack of tooling meant this rarely happened with any rigor - leaving a real risk that structure drift or bad data would reach production unnoticed.
Want to know how we achieved these results?
Contact us today to learn more about our DQV-driven regression testing approach and the success story behind this engagement with the Equity Markets & Prime Brokerage Technology team at one of the top 5 Canadian banks.
More proof

RECORDS MIGRATED, MASKED & VALIDATED PER MINUTE
238K
How We Secured a 3.3M-Record Dev/UAT Migration for a Leading Bank - Masked, Encrypted & Validated in 14 Minutes
The bank's Portfolio Project Management business needed to migrate 3.3 million production records across 61 tables into the Dev/UAT environment for realistic application performance testing...

RECORDS MIGRATED & VALIDATED PER MINUTE
920K
23 Million Records, Zero Errors - How a Leading Canadian Bank Moved Off DB2 with DQV
The bank's Retail Wealth Management Technology business wanted to migrate roughly 23 million records across 15 tables from its IBM DB2 database to a new MS SQL Server database on Azure cloud...
