Back to Exposure Report
Identity Services / BiometricsAugust 2026Global

ClarityCheck

Every other breach in this series involves data that can eventually be replaced. A face cannot.

Facial photographsEmail addressesPhone numbers
1

What happened?

ClarityCheck, an identity and reverse-lookup service, exposed a dataset reported to include more than nine million image files consisting of facial photographs, alongside email addresses and phone numbers. No threat actor has been publicly attributed as of this writing, and the company has not published a full accounting.

2

What data was actually inside?

Facial photographs, email addresses, and phone numbers, reported at a scale of more than nine million image files.

Photographs of faces are biometric data whether or not the holder ever described them that way. A facial image does not need to have been converted into a template to be biometrically useful — modern recognition systems work directly from photographs, and a large corpus of images paired with verified email addresses and phone numbers is a labeled dataset linking faces to identities.

3

Who gets hurt and how?

The permanence is the point. Card numbers get reissued, passwords rotate, and even a Social Security number can, with enormous difficulty, be changed. A face is fixed, and it has quietly become a credential — for phone unlock, banking app login, airport boarding, and identity verification at account opening.

A large corpus of labeled faces is material for building and testing systems designed to defeat exactly those checks. Liveness detection is an arms race, and training data is ammunition.

There is a second, sharper problem specific to a lookup service. People submitted photographs to check on someone else — to verify a dating match, to identify an unknown caller. A substantial share of the subjects in those images never interacted with the company at all. They cannot be notified, because no relationship exists through which to reach them, and they never consented to anything.

4

What did they think they were doing right?

While not publicly confirmed for this company, identity verification services typically operate with security programmes calibrated to the sensitivity of identity documents, because that is the data category the industry recognises as high risk and the one customers ask about.

The assumption that tends to break is that uploaded images are transient. A photograph submitted for a one-time lookup feels like a query input rather than a stored record — the user's mental model is that they searched, they got an answer, and the search is over. Systems retain the input anyway, for caching, for abuse prevention, for model improvement, or simply because the storage bucket has no lifecycle policy.

5

What did they not know about their own data?

Image stores are the least inventoried data type in common use. Structured databases have schemas, and a scan of a table tells you what columns exist. An object storage bucket holding nine million files tells you almost nothing without opening them, and most classification tooling was built for text.

So organisations routinely hold enormous volumes of images whose contents nobody has characterised — uploaded documents, support attachments, verification selfies, scanned forms. These accumulate faster than any other category because they are generated by users rather than by the business, and they are the least likely to appear on a data map with an accurate description.

The question worth asking in any organisation is not whether you collect biometric data deliberately. It is whether your users have uploaded photographs of faces into a bucket somebody set up for attachments.

If you use cloud storage, do you know what sensitive data lives in your buckets and blobs? Or would you find out the same way they did?

6

What does attribution look like the morning after?

Biometric regulation is uneven in a way that produces very different outcomes for identical data. The Illinois Biometric Information Privacy Act provides a private right of action with statutory damages per violation and has produced substantial settlements; Texas and Washington have their own regimes. Under GDPR, biometric data processed for identification is special category data under Article 9, carrying the highest tier of obligation. Most other jurisdictions have nothing specific at all.

The unresolvable part is notification of non-users. A person photographed by someone else and submitted to a lookup service is a data subject with rights and no contact record. There is no mechanism to inform them, and in most cases no way for the company to know who they are.

7

What would have changed the outcome?

Knowing what is inside the image stores — and applying a retention policy to uploads, which are the fastest growing and least classified data category most organisations hold.

Ask whether your organisation holds biometric data and most teams will say no. Ask whether users can upload photographs and the answer is usually yes, into a bucket with no lifecycle rule, containing files nobody has characterised. Data you cannot describe is data you cannot delete, and irreplaceable data that is never deleted eventually becomes an irreversible disclosure. Compare our analysis of the Organization for Transformative Works breach, where the harm also fell largely outside what notification statutes measure.

ClarityCheck found out the hard way.

Your team could spend the next 6 months rebuilding systems, notifying customers, and answering legal questions. Or you could spend 24 hours finding out what's actually at risk.