Benchmark methodology

DataMatch Enterprise 10 Million Record Benchmark Methodology

Data Ladder ran a controlled internal benchmark comparing the DataMatch Enterprise web application (build 1.0.16) with the legacy application (build 3.9.3) on a 10,000,000-record dataset: same source file, same match definition, same machine.

Elapsed time · 10M records
Data import4m vs 5m
Data profiling6m vs 10m
Data matching41–43m vs crashed ~10h
New web appLegacy appDid not complete
10M
records in one 12-column, 1.2 GB CSV file
2 runs
controlled runs with fresh imports and no cache
< 45 min
web app matching time in both runs (~41 and ~43 min)
~40%
less profiling elapsed time than the legacy app
Benchmark scope

What this benchmark measured

The test measured elapsed time for data import, data profiling and data matching. It was designed to compare the two DataMatch Enterprise application architectures under a consistent configuration.

It did not test matching accuracy or compare DataMatch Enterprise with another vendor's product.

Measured In scope

  • Data import elapsed time
  • Data profiling elapsed time
  • Data matching elapsed time and completion

Not measured Out of scope

  • Matching accuracy
  • Comparison with other vendors' products
Test configuration

Same data. Same rules. Same machine.

Both applications were tested one at a time under identical conditions.

Test data and software

Test typeControlled internal comparison
Dataset10,000,000 records in a 12-column, 1.2 GB CSV file
Data typesText and numeric fields
Web applicationBuild 1.0.16
Legacy applicationBuild 3.9.3

Hardware and test controls

Processor11th-generation Intel Core i5
Memory32 GB RAM
StorageNVMe SSD
RunsTwo
Cache stateFresh imports with no cache
Timing sourceEach application's Job Execution Centre
Test procedure

How the test was run

Use the same 10-million-record source file and match definition for both applications.

Run one application at a time on the same machine.

Clear resources between sessions and begin with fresh imports without cached data.

Capture elapsed time from the Job Execution Centre in each application.

Repeat the controlled test to check whether the completion outcome is consistent.

Observed results

Execution time, side by side

Elapsed times are rounded and reported as observed in each application's Job Execution Centre.

Data import

New web app~4 min
Legacy app~5 min
Comparable

Data profiling

New web app~6 min
Legacy app~10 min
~40% less elapsed time

Data matching

New web app~41 & ~43 min
Legacy appCrashed ~10 hr
Completed both runs
OperationNew web appLegacy appObserved result
Data importAbout 4 minAbout 5 minComparable  Web app finished about one minute sooner
Data profilingAbout 6 minAbout 10 min~40% less time  for the web app
Data matchingAbout 41 and 43 minAbout 120+ minsCompleted  Web app completed both runs; legacy app did not complete
Interpretation

How to read these results

Import performance was comparable in this test. The bigger differences appear in profiling and matching.

~40%

Less profiling time

Profiling elapsed time fell from approximately 10 minutes to approximately 6 minutes, a reduction of about 40%.

Completion, not a speed ratio

The matching result should be read as a completion test rather than a direct speed ratio because the legacy application did not complete the workload. Under this configuration, the web application completed both recorded matching runs in less than 45 minutes.

Benchmark limitations

What these results don't tell you

This was an internal Data Ladder benchmark, not an independent third-party test.

The comparison used one dataset, one match definition and one hardware environment.

Actual elapsed time will vary with dataset structure, data quality, matching rules, thresholds, infrastructure and allocated resources.

The test measured elapsed time and completion. It did not measure or compare matching accuracy.

Separate Data Ladder figures for 2-million-record processing, matches per second and matching accuracy are not part of this benchmark and should not be combined with these results.

Evaluation guidance

Validate on your own data

Organizations should validate performance with representative records, matching rules and infrastructure before using these results for capacity planning. A useful proof of concept should measure:

01Completion time
02Precision
03Recall
04False positives
05False negatives
06Manual review required

Run your own proof of concept

Test DataMatch Enterprise with your records, your matching rules and your infrastructure.