Verifex
Verifex Research

Report VR-2026-01Screening accuracy··12 min read

Real aliases, not synthetic variants: measuring name screening on 2,000 OFAC aliases

Verifex Research · Self-administered study · Not independent validation

Abstract

Most screening benchmarks build their test names by perturbing listed names: a dropped letter, a swapped token, a transliteration. Real aliases are different names, not variants. We built 2,000 test cases from OFAC's own published aliases and screened each against the full 19,329-entry OFAC list. The first run found 1,677 (83.9%, 95% CI 82.17 to 85.4%), while our synthetic suite had reported 92.4%. The gap exposed two defects that no synthetic test could see: query and index normalised the Arabic definite article differently, and exact retrieval depended on name order. Fixing both raised recall to 98.1% (95% CI 97.4 to 98.61%), with organisation aliases complete and no change in false positives on a 600-name probe. We also show that the false-positive rate of one engine spans three orders of magnitude depending on which non-matches it is tested on, and argue that a false-positive rate without its test set is not a measurement.

Key findings

  1. 1Synthetic tests overstated recall by 8.5 points: 92.4% on generated variants against 83.9% on 2,000 real published aliases.
  2. 2143 of the 323 first-run misses (44.3%) carried an Arabic particle such as al-, against 29.5% of all cases. A character-identical published alias could return no match.
  3. 3179 of 215 remaining person misses were never retrieved at all: the list stored the alias in one name order and the query used another.
  4. 4After both fixes, 1,962 of 2,000 aliases were found (98.1%): every organisation alias (1,120 of 1,120) and 95.7% of person aliases. Nothing that matched before stopped matching.
  5. 5The same engine's severe false-positive rate was 0 of 5,000, 1 of 25,000 or 26 of 600, depending on the test set. The hardest set did not move when recall rose by 14.2 points.

The problem with synthetic tests

A screening engine has one job it must not fail: find the designated party, under whatever name the customer gave. Vendors usually test this by taking a listed name and perturbing it. Drop a letter, swap two tokens, transliterate a vowel. The engine is then scored on how many perturbed names it finds.

That measures something real, but not the thing that matters most. A designated party's aliases are often entirely different names. OFAC lists “AVIA IMPORT” as an alias of an organisation whose primary name shares no word with it. No perturbation of the primary name will ever produce it. A synthetic suite cannot measure whether an engine finds such names, and it cannot find the defects that stop it.

Study design

2,000
test cases built from OFAC's published strong aliases
19,329
OFAC entries in the corpus, loaded by the production parser
600
near-miss person names in the false-positive probe

Each case is a published alias; the right answer is the designated party it belongs to. Cases were graded by difficulty, from aliases that share most tokens with the primary name to aliases that share none. We ran the suite three times on the same corpus: once as found, then after each of two fixes. Before and after every change we also ran a false-positive probe: 600 person names recombined from tokens of different designated entries, with any name equal to a real listed name or alias removed. They look like listed names and are not.

The first run

Figure 1

Recall on real OFAC aliases, three runs, against the synthetic suite

Show the data
MeasureEstimate95% CI low95% CI high
Synthetic suite92.4%86.6%95.8%
First run83.9%82.2%85.4%
After fix 1: normaliser89.3%87.8%90.5%
After fix 2: name order98.1%97.4%98.6%
Point estimates with Wilson 95% intervals. The synthetic suite (n=132) measured perturbed names; the three runs measured 2,000 real aliases on the same 19,329-entry corpus.

The first run found 1,677 of 2,000 aliases: 83.9%. The synthetic suite had reported 92.4%. The intervals barely touch. The synthetic figure was not wrong about the cases it measured; it was measuring easier cases. Every alias the first run missed was in the index. The engine had the right record and did not return it.

Defect 1: the Arabic article

143 of the 323 misses (44.3%) involved a name with a particle such as al-, el- or ul-, although such names were only 29.5% of the cases. The cause was an asymmetry. When a list entry was stored, the article was kept. When a query arrived, it was removed. The two sides of one comparison were normalised by two different rules, so an exact match on any name containing the particle could never fire.

Figure 2

Names with an Arabic particle were over-represented among misses

Show the data
CategoryValue (%)
Share of misses44.3
Share of all cases29.5
First run. Share of misses that carried a particle (al-, el-, ul-) against the share of all cases that did.

The most uncomfortable example was a character-identical published alias of a designated organisation that returned no match at all. The fix compares both normalised forms. It is strictly additive: it can add a match, never remove one. It recovered 108 cases and completed organisation aliases.

Defect 2: name order

Person aliases still lagged. Of the 215 person misses left after the first fix, 179 were never even retrieved as candidates. Listing conventions disagree about which part of a name comes first: the list stored an alias surname first, while a person typing it puts the given name first. The scoring layer already knew these were the same name. The retrieval layer never handed it the candidate. Exact retrieval for person names now tries the possible orderings, within a bound; organisation names are left alone because their word order carries meaning.

After both fixes

Figure 3

Recall by subject and by difficulty, run by run

Dashed line: the first run, for comparison.

Show the data
GroupCasesFirst runAfter fix 1: normaliserAfter fix 2: name order
Organisations1,1201,013 (90.4%)1,120 (100.0%)1,120 (100.0%)
Persons880664 (75.5%)665 (75.6%)842 (95.7%)
Easy179170 (95.0%)170 (95.0%)175 (97.8%)
Medium401384 (95.8%)388 (96.8%)397 (99.0%)
Hard522499 (95.6%)504 (96.6%)517 (99.0%)
Adversarial898624 (69.5%)723 (80.5%)873 (97.2%)
Switch between runs. The dashed line marks the first run. Difficulty grades how far an alias is from the primary name; adversarial aliases share no token with it.

Together the two fixes recovered 285 aliases: 14.2 points of recall. Nothing that matched before stopped matching at either step, and the false-positive probe did not move: 26 of 600 names were rated severe before the fixes and after them.

False positives depend on the test set

Vendors quote a false-positive rate as if it were a property of the engine. It is a property of the engine and the names it was tested on. Here is one engine, the same year, on three sets of names that are not on any list.

Figure 4

Severe false-positive rate on three different sets of non-matches

Show the data
MeasureEstimate95% CI low95% CI high
Near-misses0.000%0.000%0.077%
Synthetic names0.004%0.001%0.023%
Recombined tokens4.330%2.970%6.270%
Severe ratings with Wilson 95% intervals. Near-misses: adversarial names verified not to be listed (August 12, 2026). Synthetic names: a generated negative corpus (September 6, 2026). Recombined tokens: person names built from name parts of different designated persons (September 7, 2026).

The rates differ by three orders of magnitude, and none of them is wrong. A two-word name built from the real given name of one designated person and the real surname of another, beside a listed three-word name that contains both, is exactly the case an analyst must look at. Rating it severe is defensible; the probe exists to make sure a recall improvement does not quietly make such ratings more common. That is why it is the number we hold fixed while recall changes.

What buyers should ask

  1. Was recall measured on real published aliases, or only on generated variants of primary names?
  2. What is the denominator, and what is the interval?
  3. Which non-matches was the false-positive rate measured on, and how were they built?
  4. When recall improved, what happened to false positives on the same test?
  5. Which misses remain, and what causes them?

Our current production benchmark and its failure cases are on the benchmark page. The companion report examines what a Wikidata-derived PEP list really contains.

This report describes measurements. It is not legal advice; confirm your obligations with a qualified adviser.

Limitations

  • Self-administered: Verifex built the suite and ran it. It is evidence about these cases, not independent validation.
  • OFAC only, and only the aliases OFAC classes as strong. Other lists, weak aliases and names in non-Latin scripts are not measured here.
  • The fixes were found and measured on the same suite. A fresh, held-out alias sample is needed to rule out fitting to this suite.
  • Runs used a local database built through the documented fresh-database path and populated by the production OFAC parser (19,329 entries, 0 failures), not the live production index.
  • The 25,000-name negative figure comes from a synthetic corpus measured on September 6, 2026; the 5,000-name figure from the August 12, 2026 benchmark. They describe different test sets, which is the point of the comparison, not a like-for-like trend.
  • The false-positive probe contains person names only.

Data and code

Aggregate counts for every figure are published below as JSON. The test suite, the probe generator and the per-case results (which name real designated parties) are kept in the Verifex repository and are available to customers under NDA.

How to cite

Verifex Research (2026). Real aliases, not synthetic variants: measuring name screening on 2,000 OFAC aliases. Verifex Research Report VR-2026-01. https://verifex.dev/research/ofac-alias-screening-benchmark

References

  1. U.S. Department of the Treasury, Office of Foreign Assets Control. Specially Designated Nationals and Blocked Persons List (SDN), with aliases.
  2. Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer.
  3. Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212.
  4. Hanley, J. A., & Lippman-Hand, A. (1983). If nothing goes wrong, is everything all right? Interpreting zero numerators. JAMA, 249(13), 1743–1745.
  5. Verifex. Production benchmark and matching methodology.

More research