DeSpy Privacy Academy Lesson 3

August 4, 2026

AI & Digital Identity

How your identity is built, used, and shared in the age of AI.

Lesson 3 of 6 · Estimated read: 13 minutes · Level: everyone

You have a version of yourself you’ve never met.

Industry sometimes calls it a digital twin: a modeled copy of you, assembled from purchases, clicks, locations, photos, and the behavior of people who resemble you. It makes predictions about your health, your finances, and your intentions. Companies act on it as though it were you. You’ve never seen it, and in most cases you can’t correct it.

The term borrows from engineering, where a digital twin is a simulation of a physical machine used to predict failures before they happen. Applied to a person, the ambition is the same: model the system well enough to anticipate what it does next.

This lesson is about how that profile gets built, why AI changed the economics of building it, and what you can actually do, including several rights most people don’t know they have.

In this lesson

  1. How AI builds identity profiles
  2. Facial recognition and biometric data
  3. Behavioral tracking and scoring
  4. Data aggregation and identity linking
  5. Deepfakes and synthetic identity risks
  6. Protecting your identity in an AI world
  7. Tools to reclaim your digital identity

1. How AI builds identity profiles

The basics

A profile isn’t a filing cabinet of facts about you. It’s mostly inferences: predictions generated by comparing your patterns to millions of other people’s.

The distinction matters. Your data is the raw input. The valuable product is what gets guessed from it: your likely income band, whether you’re about to move, whether you’re managing a health condition, how price-sensitive you are, whether you’re likely to respond to a certain political message.

Go deeper

What most people don’t know

Deleting your data doesn’t remove you from a model that already learned from it.

This is the most important technical fact in the lesson. When you file a deletion request, a company removes records from databases. But if your data was used to train a model, your influence is baked into the model’s weights, diffused across billions of parameters, not stored in a row anyone can find and drop.

Removing it is an unsolved research problem called machine unlearning. Current approaches are approximate, expensive, or require retraining from scratch. In practice, companies delete the record and keep the model.

The strategic consequence is large and cuts against how most privacy advice is framed: prevention is worth far more than remediation. A deletion request is genuinely useful for stopping future flows and future sales. It is not a rewind button. Data you never gave up is in a different category from data you gave up and later asked to have deleted.

Do this today

Look at your own profile. Several companies will show you:

2. Facial recognition and biometric data

The basics

Facial recognition doesn’t compare photographs. It converts a face into a faceprint, a mathematical template, and compares templates. Two very different operations get called the same thing:

The second is the one with civil liberties implications.

Go deeper

What most people don’t know

You can decline facial recognition at U.S. airport security, and most people don’t realize it’s optional.

At TSA checkpoints using facial comparison, you can say “I decline facial recognition” and request a standard manual ID check instead. You’re entitled to do this without penalty or losing your place in line. Signage disclosing the option is often small or absent.

Second thing worth knowing: biometrics can’t be revoked. A breached password gets changed in thirty seconds. A breached faceprint or fingerprint template is permanent. You cannot issue yourself a new face. This is why using biometrics for convenience (unlocking a phone, where the template stays on-device in secure hardware) is a very different risk than surrendering biometrics to a third-party database.

And a nuance that surprises people: you don’t have to be in a database to be findable. Photos of you that others uploaded and tagged, images of relatives who resemble you, and recognition techniques that work from the area around the eyes alone all contribute. Masks and sunglasses are far less protective than assumed.

Do this today

3. Behavioral tracking and scoring

The basics

Behavioral tracking has moved past what you click to how you move. Typing rhythm, mouse path, scroll speed, how long you hover before deciding, the angle you hold your phone, how you swipe.

These patterns are individually distinctive enough to function as identification, and they’re collected without any permission prompt.

Go deeper

What most people don’t know

You are scored by companies you’ve never heard of, and you can request several of those scores.

Beyond credit, a whole layer of consumer scoring operates in the background:

The pattern to notice: these function as consumer reports in effect. When they do, federal law gives you access and dispute rights, but only if you know the score exists to ask about it.

Do this today

4. Data aggregation and identity linking

The basics

Aggregation is the step that makes everything else work: identity resolution. Separately, your email, phone number, ad ID, IP address, loyalty card, and browser cookie are fragments. Linked, they’re one continuous record.

Companies specializing in this maintain graphs connecting all your identifiers, plus your household, plus your offline purchases.

Go deeper

What most people don’t know

A hashed email address is not anonymous.

This is the misconception that survives even among technical people. Hashing your email with SHA-256 produces a fixed string, described in industry materials as privacy-preserving. But hashing is deterministic: the same email always yields the same hash. So any company that already has your email can compute the hash and match you instantly. It isn’t encryption with a key nobody holds. It’s a stable pseudonym that everyone who knows your email can reproduce.

The practical implication is that your email address is the master key to your profile, and using one address everywhere is what makes the graph cohere. Email aliasing is therefore one of the highest-leverage privacy practices available, and one of the least used.

Related: “data clean rooms,” now common in advertising, restrict what queries can be run against combined datasets. They don’t anonymize the underlying data.

Do this today

5. Deepfakes and synthetic identity risks

The basics

Two separate threats wear the same name.

Deepfakes of you: your likeness or voice replicated to deceive someone, usually your family, your employer, or your bank.

Synthetic identities: a fake person assembled partly from your real data, used to open accounts and take out credit.

Go deeper

What most people don’t know

Children are the preferred raw material for synthetic identity fraud, and freezing their credit is free and almost nobody does it.

Synthetic identity fraud combines a real Social Security number with a fabricated name and date of birth. The ideal SSN belongs to someone with no credit history and no reason to check. That means children. Fraud can accumulate for fifteen years before a teenager applies for their first loan and discovers the mess.

Under federal law, you can freeze a minor child’s credit at all three bureaus for free. It takes about thirty minutes total, requires documentation of guardianship, and eliminates the entire attack. It is one of the highest-return privacy actions available to a parent, and it’s largely unknown.

Do this today

6. Protecting your identity in an AI world

The basics

Three principles, in order of effectiveness: reduce what you produce, separate your identities so they can’t be linked, and verify before you act on anything urgent.

Go deeper

What most people don’t know

Your relatives can expose your identity without your involvement — permanently.

Consumer DNA databases made this concrete. If a cousin uploads their genome to a genealogy service, you become findable through them, because relatedness is inferable. Law enforcement has used exactly this technique to identify people who never took a test. You cannot opt out of your relatives’ choices.

This became urgent when a major consumer genetics company entered bankruptcy proceedings in 2025, and genetic data, the most permanent identifier there is, became an asset in a corporate sale. Multiple state attorneys general publicly urged customers to delete their data.

The general lesson, and the reason this section closes the loop with Section 1: for irrevocable data, the only real control is upstream. You can change a password, rotate an email, reset an ad ID. You cannot change your genome, your face, or the fact that a model already trained on you. Spend your effort where it compounds.

Do this today

7. Tools to reclaim your digital identity

Ranked by protection gained per hour spent

  1. Freeze your credit (and your kids’): free, permanent, blocks the most damaging fraud
  2. Email aliases going forward: breaks the key that links your profile
  3. Turn off ad personalization across Google, Meta, Amazon, Microsoft
  4. Face-search opt-outs and photo face-tagging off
  5. Request your data and your inferences: you can’t manage what you can’t see
  6. Broker deletion: via California’s DROP if eligible, or a removal service
  7. A family verification word: costs nothing, defeats voice cloning

Tools worth knowing

The honest limits

Broker removal is not one-time. Records repopulate from upstream sources, so it’s maintenance rather than a fix. That’s what the subscription services are actually selling.

Search removal hides a page from Google. The page remains.

And the limit that governs all of this: no tool removes you from a model already trained on your data. Every honest version of this lesson ends the same way: the leverage is in what you don’t hand over next.

Recap

Key terms

Check your understanding

1. You submit a deletion request and the company confirms your data is deleted. Are you out of their AI model? No. Deletion removes database records. Data already used in training is diffused across the model’s weights, and removing that influence is an unsolved problem. Your request stops future use. It doesn’t reverse past training.

2. An ad company says it only uses hashed email addresses, never your real one. Are you anonymous? No. Hashing is deterministic, so anyone who already has your email can compute the same hash and match you. It’s a stable pseudonym, which is why unique aliases per service are so effective.

3. You get a call from your daughter’s number. It’s her voice, she’s distressed, she needs money urgently. What do you do? Hang up and call her back on the number you already have, or ask for the family verification word. Voice cloning needs only seconds of audio, and urgency is the tool that stops people from checking.

Next lesson: Connected Cities — public surveillance, smart infrastructure, and your privacy.