Data preparation pipeline
Watchman prepares list records and search queries the same way so a query like nicolas maduro can match a list name stored as MADURO MOROS, Nicolas. This page is that preparation. Scoring is Similarity methodology. Candidate indexes are Indexing.
There are two layers:
- List ingest — extra steps some sources run when a file is loaded. OFAC is the main one.
Normalize()— runs on every list entity after ingest and on every search query.
List ingest
When Watchman loads lists (like the OFAC SDN) it organizes the names and values into a common ordering.
Name reordering (people only, SDNType=individual):
MADURO MOROS, Nicolas→Nicolas MADURO MOROS
Company suffix stripping (companies):
SAI ADVISORS INC.→SAI ADVISORS- Suffixes include
INC.,LLC,LTD.,GMBH,CO., and similar.
A query is not reordered or suffix-stripped. You can search Nicolas Maduro or MADURO, Nicolas; after Normalize() both become comparable to the OFAC person name that was already reordered at ingest.
Normalize (every list record and every query)
Entity.Normalize() fills PreparedFields used by scoring and indexes.
Names
- Drop a leading trade-document role marker (
SHIPPER:,BENEFICIARY:,ORDERING CUSTOMER:,MESSRS.,FIELD 59:). - Treat former-name markers as
HistoricalInfoand drop them from the primary name:(ex-Cape Diamond),(f.k.a. …),(formerly …),(formerly known as …), inlinef.k.a./f/k/a/formerly, and anextoken with a separator (Kalliopi Dawn ex Orion Tanager). Chained(ex A, ex B)becomes two former names. A trailing year on the former name (, 2020oruntil 2019) is dropped.(example)stays on the primary name. - Trim, lowercase, turn
.,-and other punctuation/symbols into spaces. - Unicode NFD → strip combining marks → NFC, so
Raúlandraulmatch. - For vessels, drop a trailing port or place (
MV SIAM ORCHID 7, Bangkok→MV SIAM ORCHID 7; a final(Laem Chabang)is dropped), then drop leading ship-type markers (MV,M/V,M/T,SS,N/M,T/B, and the same after punctuation becomes spaces).MV Solenne HarbourandSolenne Harbourcompare as the same prepared name. - For businesses and organizations, rewrite English legal-form phrases to the short token (
incorporated→inc,limited→ltd,limited liability company→llc,public limited company→plc,corporation→corp,company→co).Harrowfield Bearings LimitedandHARROWFIELD BEARINGS LTD.compare as the same prepared name. Jurisdiction-specific forms (GmbH,Sdn. Bhd.) are left as they are. - Split on whitespace into tokens.
- Drop stopwords (
of,the,and, and the same idea in other languages). Language is guessed from the name. Numeric tokens (11420) are kept. SetKEEP_STOPWORDS=trueto skip this step.
The same steps run on the primary name, aliases, and former names (HistoricalInfo of type Former Name).
Example: BANK OF AMERICA → tokens bank, america.
Phones and fax
Digits are kept and, when possible, normalized with a country guess (+1 (555) 010-0100 and 555-010-0100 can compare).
Addresses
- Free-text
address=on the query is parsed into line, city, state, postal code, country. Docker images and Linux/macOS GitHub releases use libpostal. Windows.exeandgo runuse usaddress unless built with-tags libpostal. Any build can enable deepparse. See Addresses. - Parsed fields are lowercased. Commas are stripped from street lines. Country codes (
US,USA) fold to a common name (United States). - Street, city, and similar fields are split into tokens for scoring.
Gender (query parameter)
gender= is mapped to male, female, or unknown (m / male / f / female, and a few synonyms).
Government IDs
Query IDs are gov_<type>=COUNTRY:IDENTIFIER (for example gov_passport=IR:Y53914915). Country on the ID is normalized the same way as address country when scoring.
What this means for searches
- Case, accents, and punctuation do not need to match the list file:
Joséandjosecompare as the same letters. - Common words can be omitted:
Bank of AmericaandBank America. - For OFAC people, list names in
SURNAME, Givenorder are stored asGiven SURNAME, so a natural-order query works. - For OFAC companies,
Inc./LLCon the list name is stripped at ingest. A query ofAcme Incstill hasincas a token unless you omit it; it is a weak token, not a blocker. English legal-form expansions on the query (Incorporated,Limited) are rewritten to the same short tokens asInc./Ltd.duringNormalize(). - Send IDs and dates when you have them. Preparation does not invent them.
Debugging
debug=true (or debug=0.80) on /v2/search returns score pieces (name, identifiers, dates, addresses), not a dump of these pipeline stages. If a name does not match, check type (person vs business), whether you sent the same fields as the list record, and the prepared tokens implied above.