The problem with synthetic tests
A screening engine has one job it must not fail: find the designated party, under whatever name the customer gave. Vendors usually test this by taking a listed name and perturbing it. Drop a letter, swap two tokens, transliterate a vowel. The engine is then scored on how many perturbed names it finds.
That measures something real, but not the thing that matters most. A designated party's aliases are often entirely different names. OFAC lists “AVIA IMPORT” as an alias of an organisation whose primary name shares no word with it. No perturbation of the primary name will ever produce it. A synthetic suite cannot measure whether an engine finds such names, and it cannot find the defects that stop it.
Study design
- 2,000
- test cases built from OFAC's published strong aliases
- 19,329
- OFAC entries in the corpus, loaded by the production parser
- 600
- near-miss person names in the false-positive probe
Each case is a published alias; the right answer is the designated party it belongs to. Cases were graded by difficulty, from aliases that share most tokens with the primary name to aliases that share none. We ran the suite three times on the same corpus: once as found, then after each of two fixes. Before and after every change we also ran a false-positive probe: 600 person names recombined from tokens of different designated entries, with any name equal to a real listed name or alias removed. They look like listed names and are not.
The first run
Figure 1
Recall on real OFAC aliases, three runs, against the synthetic suite
Show the data
| Measure | Estimate | 95% CI low | 95% CI high |
|---|---|---|---|
| Synthetic suite | 92.4% | 86.6% | 95.8% |
| First run | 83.9% | 82.2% | 85.4% |
| After fix 1: normaliser | 89.3% | 87.8% | 90.5% |
| After fix 2: name order | 98.1% | 97.4% | 98.6% |
The first run found 1,677 of 2,000 aliases: 83.9%. The synthetic suite had reported 92.4%. The intervals barely touch. The synthetic figure was not wrong about the cases it measured; it was measuring easier cases. Every alias the first run missed was in the index. The engine had the right record and did not return it.
Defect 1: the Arabic article
143 of the 323 misses (44.3%) involved a name with a particle such as al-, el- or ul-, although such names were only 29.5% of the cases. The cause was an asymmetry. When a list entry was stored, the article was kept. When a query arrived, it was removed. The two sides of one comparison were normalised by two different rules, so an exact match on any name containing the particle could never fire.
Figure 2
Names with an Arabic particle were over-represented among misses
Show the data
| Category | Value (%) |
|---|---|
| Share of misses | 44.3 |
| Share of all cases | 29.5 |
The most uncomfortable example was a character-identical published alias of a designated organisation that returned no match at all. The fix compares both normalised forms. It is strictly additive: it can add a match, never remove one. It recovered 108 cases and completed organisation aliases.
Defect 2: name order
Person aliases still lagged. Of the 215 person misses left after the first fix, 179 were never even retrieved as candidates. Listing conventions disagree about which part of a name comes first: the list stored an alias surname first, while a person typing it puts the given name first. The scoring layer already knew these were the same name. The retrieval layer never handed it the candidate. Exact retrieval for person names now tries the possible orderings, within a bound; organisation names are left alone because their word order carries meaning.
After both fixes
Figure 3
Recall by subject and by difficulty, run by run
Dashed line: the first run, for comparison.
Show the data
| Group | Cases | First run | After fix 1: normaliser | After fix 2: name order |
|---|---|---|---|---|
| Organisations | 1,120 | 1,013 (90.4%) | 1,120 (100.0%) | 1,120 (100.0%) |
| Persons | 880 | 664 (75.5%) | 665 (75.6%) | 842 (95.7%) |
| Easy | 179 | 170 (95.0%) | 170 (95.0%) | 175 (97.8%) |
| Medium | 401 | 384 (95.8%) | 388 (96.8%) | 397 (99.0%) |
| Hard | 522 | 499 (95.6%) | 504 (96.6%) | 517 (99.0%) |
| Adversarial | 898 | 624 (69.5%) | 723 (80.5%) | 873 (97.2%) |
Together the two fixes recovered 285 aliases: 14.2 points of recall. Nothing that matched before stopped matching at either step, and the false-positive probe did not move: 26 of 600 names were rated severe before the fixes and after them.
False positives depend on the test set
Vendors quote a false-positive rate as if it were a property of the engine. It is a property of the engine and the names it was tested on. Here is one engine, the same year, on three sets of names that are not on any list.
Figure 4
Severe false-positive rate on three different sets of non-matches
Show the data
| Measure | Estimate | 95% CI low | 95% CI high |
|---|---|---|---|
| Near-misses | 0.000% | 0.000% | 0.077% |
| Synthetic names | 0.004% | 0.001% | 0.023% |
| Recombined tokens | 4.330% | 2.970% | 6.270% |
The rates differ by three orders of magnitude, and none of them is wrong. A two-word name built from the real given name of one designated person and the real surname of another, beside a listed three-word name that contains both, is exactly the case an analyst must look at. Rating it severe is defensible; the probe exists to make sure a recall improvement does not quietly make such ratings more common. That is why it is the number we hold fixed while recall changes.
What buyers should ask
- Was recall measured on real published aliases, or only on generated variants of primary names?
- What is the denominator, and what is the interval?
- Which non-matches was the false-positive rate measured on, and how were they built?
- When recall improved, what happened to false positives on the same test?
- Which misses remain, and what causes them?
Our current production benchmark and its failure cases are on the benchmark page. The companion report examines what a Wikidata-derived PEP list really contains.
This report describes measurements. It is not legal advice; confirm your obligations with a qualified adviser.
Limitations
- Self-administered: Verifex built the suite and ran it. It is evidence about these cases, not independent validation.
- OFAC only, and only the aliases OFAC classes as strong. Other lists, weak aliases and names in non-Latin scripts are not measured here.
- The fixes were found and measured on the same suite. A fresh, held-out alias sample is needed to rule out fitting to this suite.
- Runs used a local database built through the documented fresh-database path and populated by the production OFAC parser (19,329 entries, 0 failures), not the live production index.
- The 25,000-name negative figure comes from a synthetic corpus measured on September 6, 2026; the 5,000-name figure from the August 12, 2026 benchmark. They describe different test sets, which is the point of the comparison, not a like-for-like trend.
- The false-positive probe contains person names only.
Data and code
Aggregate counts for every figure are published below as JSON. The test suite, the probe generator and the per-case results (which name real designated parties) are kept in the Verifex repository and are available to customers under NDA.