How DROP hashing works: standardizing records so they match
Every DROP identifier is hashed. If your records aren't formatted exactly to CalPrivacy's rules, matches fail without any error. The rules and example code.
- Standardize first: lowercase, strip special characters,
YYYYMMDDdates, five-digit ZIPs, ten-digit phones. - Hash with the same algorithm the list uses.
- For multi-field lists, hash each field, join the hashes, then hash the result.
Why hashing makes matching fragile
DROP never gives brokers raw identifiers. Every value on a consumer deletion list is hashed. To find a match, you hash your own records and compare hashes. A hash changes completely if a single character is different, so Jane.Doe@x.com and jane.doe@x.com produce unrelated values. Nothing tells you a near-miss happened. The request just looks like "record not found."
The regulations specify exactly how to standardize records before hashing.
The standardization rules (§ 7613)
| Rule | Example |
|---|---|
| Use all lowercase letters, including names | Jane → jane |
| Remove extraneous or special characters, except in email addresses | O'Connor-López → oconnorlopez |
| Convert non-English characters to the closest English character | Björn → bjorn |
| Date of birth as 8 digits: year, month, day | January 12, 1990 → 19900112 |
| ZIP code as the first 5 characters | 95811-6213 → 95811 |
| Phone as the last 10 digits, no dashes or country code | +1 (987) 765-4321 → 9877654321 |
| Any other standardization you know will improve matching | e.g. trimming stray whitespace |
The regulation's own example: Björn O'Connor-López becomes bjornoconnorlopez.
You only have to standardize for matching. You don't have to store your data in this format (§ 7613(a)(1)(C)).
Hashing single identifiers
After standardizing, hash each value with the algorithm specified in the deletion list. DROP uses SHA-256. Then compare against the list for that identifier type.
Hashing multi-field identifiers
Some lists combine several fields, like first name, last name, date of birth and ZIP code. The regulation says to:
- standardize and hash each field separately;
- join the hashes into one string, with no spaces or other characters;
- hash that combined string.
The result is what you compare to the list (§ 7613(a)(2)(A)). Combining the fields before hashing them individually gives a different value and won't match.
Example code
A minimal Python version of the rules above:
import hashlib, re, unicodedata
def sha256(value):
return hashlib.sha256(value.encode("utf-8")).hexdigest()
def std_name(value):
# Björn O'Connor-López -> bjornoconnorlopez
ascii_value = unicodedata.normalize("NFKD", value).encode("ascii", "ignore").decode()
return re.sub(r"[^a-z0-9]", "", ascii_value.lower())
def std_email(value):
return value.strip().lower()
def std_dob(date): # a datetime.date
return date.strftime("%Y%m%d")
def std_zip(value):
return re.sub(r"\D", "", value)[:5]
def std_phone(value):
return re.sub(r"\D", "", value)[-10:]
def email_hash(email):
return sha256(std_email(email))
def name_dob_zip_hash(first, last, dob, zip_code):
parts = [std_name(first), std_name(last), std_dob(dob), std_zip(zip_code)]
return sha256("".join(sha256(p) for p in parts))
Where matching usually goes wrong
- Whitespace and punctuation. Leading or trailing spaces, periods in names, and apostrophes all change the hash.
- Accents.
Josémust becomejose, notjosorjosé. - Phone formats. Numbers stored with extensions or country codes need trimming to the last ten digits.
- Dates. Dates stored as text in mixed formats (
1/12/90,1990-01-12) need parsing before formatting. - Combined lists. Hashing the joined raw values instead of joining the hashes.
When there's more than one match
If one hashed identifier matches several consumers in your records, you can't tell which one made the request. The rule is to opt every matched consumer out of sale and sharing, and report the request as "record opted out of sale" (§ 7613(a)(2)(B), § 7614).
Check every record against DROP with one API call
Purgepath keeps your DROP list current daily and tells you exactly what to delete. Unlimited scrubs by API or CSV upload, $500/month.
Sources
This article is general information as of October 5, 2026, not legal advice. Rules and fees can change; check the sources above and talk to counsel about your situation.