Similarity Methodology
This page is the technical description of how Watchman scores a query against a watchlist entity. For a program-level briefing written for BSA/AML, sanctions, and model-risk staff, see For compliance and risk. Empirical results on analyst-labeled OpenSanctions pairs are in OpenSanctions Pairs.
Name scoring defaults to token-pairwise Jaro–Winkler. Optional ?algorithm= scorers are compared in Algorithm comparison.
Why multi-field scoring
Watchman scores identifiers, names, dates, addresses, contact, and supporting fields together. Name-only queries produce large review queues. OFAC’s public search is a useful illustration: the query Khamis Al returns many 100% hits on the portal; Watchman’s name scorer spreads those same SDN names from about 0.26 to 0.87. See Comparison with the OFAC portal.
That design follows classical record linkage: Fellegi and Sunter treat matching as a decision on agreement patterns across fields, not a single string compare (Fellegi & Sunter, 1969). Jaro’s comparator, later extended by Winkler, is the usual way to turn typographical variation into a partial-agreement weight (Jaro, 1989; Winkler, 1990).
Score pieces and weights
After Normalize() (case, punctuation, stopwords, phones, addresses), Similarity builds these pieces:
| Piece | Weight | Compared fields |
|---|---|---|
| Exact identifiers | 50 | Type + country + identifier (passport, tax ID, IMO/MMSI, serial, …) |
| Crypto addresses | 50 | Currency + address |
| Government IDs (loose) | 50 | Identifier with optional country (0.7–1.0) |
| Contact | 50 | Email, phone, fax |
| Name | 35 | Best pairwise token alignment, including alt and former names |
| Titles | 35 | Person titles / positions |
| Addresses | 25 | Structured address fields |
| Dates | 15 | Birth, death, incorporation, dissolution |
| Supporting | 15 | Sanctions programs, historical names |
Empty query fields are not compared. The blended score is a coverage-aware weighted average (FINAL_SCORE_* multipliers). Name-only queries are down-ranked (FINAL_SCORE_NAME_ONLY_MULTIPLIER, default 0.95).
When a match is 1.0
A piece forces 1.0 only when it is exact (the identifier and country match) and the field is unique to one party: passport, national ID, IMO, MMSI, aircraft serial, or crypto address.
A matching tax number, company registration, email, or phone raises the blended score. It does not force 1.0. Related companies often share a tax ID; two people can share an office email. A vessel call sign alone also does not force 1.0.
Identifier conflict
If both records have the same ID type (and country, when both set) with different values, the blended score is multiplied by ID_CONFLICT_PENALTY_MULTIPLIER (default 0.70). A missing ID on one side is not a conflict. A matching passport or IMO still scores 1.0 before this penalty.
Entity type
A person query is compared to people, a business query to companies. If the query type is person, business, or organization and the list record is one of those three, Watchman still scores the pair (some source files store a person under a company-like type). A person query is not scored against a vessel or aircraft. GET /v2/search?type=person still only searches the person partition of the index.
Name matching
Jaro–Winkler
sim_jw(s1, s2) = sim_j(s1, s2) + p * l * (1 - sim_j(s1, s2))
sim_jis Jaro similarity (common characters and transpositions)pis the prefix scale (default 0.1)lis the common prefix length, capped at 4
Watchman applies this token-wise, not on the raw full name:
- Names are tokenized (
John Michael Smith→john,michael,smith). - Tokens are paired with positional preference and a first-letter phonetic filter (first letters are rarely mistranscribed).
- Length-difference penalties reduce scores when one token is much shorter.
- Alternate and historical names are tried; scoring can skip remaining aliases once a high-confidence name hit is found.
Optional algorithms
?algorithm= (or MCP algorithm) selects the token-pair metric without a restart:
| Value | Role |
|---|---|
jaro-winkler (default) |
Token pairwise Jaro–Winkler |
soundex, double-metaphone, beider-morse |
Same alignment, phonetic boost when encodings overlap |
soft-bidist, soft-bisim, editex, nsim, nsim-3 |
Replace the token-pair metric; BestPairs / length / first-letter filters stay |
Process-wide USE_SOUNDEX_MATCHING still applies when algorithm is omitted.
TF-IDF and embeddings
Optional TF-IDF down-weights common tokens (Limited, ООО, GmbH) using an index built at list refresh.
Optional embeddings (Ollama, OpenAI, OpenRouter, Azure) add a neural name vector. Default EMBEDDINGS_CROSS_SCRIPT_ONLY=true: Latin queries stay on Jaro–Winkler; non-Latin queries also search the vector index. On OpenSanctions Pairs, this hybrid is the only config that moves transliteration recall in a material way. See Cross-script matching and OpenSanctions Pairs.
Entity-type matching
Person. Government IDs, given/family names and aliases, birth and death dates, titles, gender.
Business / organization. Tax and registration numbers as evidence (not identity), names and aliases, incorporation/dissolution, addresses.
Vessel / aircraft. IMO, MMSI, call sign, serial, ICAO, flag. IMO/MMSI/serial remain unique-identity keys.
Thresholds as policy
The API returns a score in [0, 1]. minMatch is a policy cutoff, not a hidden model parameter.
| Band | Typical minMatch |
Use |
|---|---|---|
| High confidence | 0.95+ | Auto-block / auto-alert after identifier confirmation |
| Screening default | 0.80 | Production hit generation with high precision |
| High recall | ~0.59 | When missing a designation is costlier than extra review |
| Enhanced due diligence | 0.70–0.84 | Broader queues |
On 472,477 analyst-judged OpenSanctions subject pairs, Jaro–Winkler at 0.80 had precision 0.986 and recall 0.689. Dropping the cutoff to 0.59 raised recall to 0.920 at precision 0.945. With cross-script embeddings (hybrid) at 0.80, precision was 0.946 and recall 0.815.
Candidate selection (before scoring)
Scoring runs only on candidates. See Indexing and Performance.
- Partition by source and type.
- Exact crypto / government-ID hits; prefix and QWERTY-near indexes on IMO, MMSI, serial, email, phone.
- Name-token inverted index: intersect distinctive tokens; fall back to the union or the full partition.
- Address prefix blocks for address-only queries.
- Searches with ≤100 candidates skip the admission queue; larger searches take
SEARCH_MAX_IN_FLIGHT.
Empty type under a known source with no entities of that type returns nothing. Empty type= selects the all-types partition for that source (slower). type is not required; send it so GET person/business fields are read and the search stays in one partition. A wrong type can miss the hit.
Measured results
A public labeled set of 755,540 sanctions/OSINT pairs is described in OpenSanctions Pairs. The evaluator in this repo is go run ./research/opensanctions-pairs. Full tables: RESULTS.md.