Back to Blog
ATSData QualityDuplicate DetectionRecruitment Technology

Duplicate Candidate Detection: Why Your ATS Needs It

Learn how ATS duplicate detection can compare email, phone and profile evidence, where false matches occur, and why recruiters should review merges before changing records.

Janis Kolomenskis

8 min readUpdated
Share

A duplicate rule should help a recruiter review records, not silently merge people. Names alone create false positives; strong signals such as a normalised email, phone number or stable profile URL are more useful, but each can be missing, shared or outdated. Keep the evidence and a human review step.

Do not assume what most ATS platforms do. Inspect the matching rules in your own system, test accented names and shared contact details, and measure false positives before changing merge policy.

The Hidden Cost of Duplicate Records

Do not import a generic duplicate-rate benchmark into your business case. Measure your own sample: draw 200 records across sources and years, review candidate pairs, then report confirmed duplicates, false positives and records that need more evidence.

The downstream effects compound quickly:

  • Double outreach: Two recruiters contact the same candidate about the same role. The candidate notices. Your agency looks disorganised.
  • Split history: Half the notes and interactions sit on one profile, half on the other. Nobody has the full picture.
  • Broken reporting: Pipeline conversion rates are inflated. Source attribution is wrong. You're making decisions on dirty data.
  • Rights-handling risk: A request is applied to one profile but misses a duplicate. Review all linked records and any applicable retention exceptions before closing the request.
Illustrative failure case: two recruiters progress separate records for the same person, while notes and deletion status remain split. The control is one reviewable identity record, not an unverified merge.

Why Name Matching Fails

These hypothetical records illustrate why names alone are insufficient. A shared name or spelling variant does not establish identity without further evidence:

ScenarioName in SystemSame Person?Name Match Catches It?
MarriageSarah Miller to Sarah ChenYesNo
Typo / transliterationMueller vs MullerYesMaybe
Common nameThomas Schmidt (x47 in DB)NoFlags all 47
NicknameRobert vs Bob WilliamsYesNo

The result: name matching produces both false positives (flagging different people as duplicates) and false negatives (missing actual duplicates with changed names). It's the worst of both worlds.

Strong Signal Matching: The Modern Approach

A practical approach is to compare several identifiers and expose the evidence to a reviewer. No single identifier is perfectly unique or permanent.

Three signals stand out:

1. Email Address

An exact normalised email can be a strong indicator, but shared inboxes, recycled addresses and provider-specific alias rules create exceptions. Lowercase and trim whitespace; do not apply Gmail-specific dot rules to other domains.

2. Phone Number

Phone numbers arrive in different formats and can be shared or reassigned. Parse them with the recorded country context into an international form; stripping a country code can turn different numbers into a false match.

3. LinkedIn Profile URL or Member ID

A matching canonical profile URL is useful evidence, but URLs can be edited, copied into the wrong record or become unavailable. Preserve the source and checked date, and combine it with another signal for automatic decisions.

Worked example: exact email plus the same verified profile URL can enter a high-confidence review queue; same name plus employer belongs in a lower-confidence queue. The reviewer still sees conflicting fields before merging.

How Modern Duplicate Detection Works in Practice

At import: When a CV is uploaded or a LinkedIn profile is imported, the system instantly checks email, phone, and LinkedIn against every existing record. If there's a match, the recruiter sees a clear alert: possible duplicate found with a link to the existing profile.

At manual creation: As a recruiter types a name, real-time suggestions surface similar records before they hit save. This prevents duplicates from being created in the first place.

In batch: For databases that have grown organically over years, a periodic batch scan identifies duplicate clusters across the entire database. The results go into a review queue for human confirmation.

The merge process matters too. A proper merge should preserve the complete activity history from both profiles — notes, emails, interview feedback, pipeline stages — and let the recruiter choose which data to keep when fields conflict. For a deeper look at database maintenance, check out our recruitment database cleanup guide.

GDPR and Duplicates: The Compliance Angle

GDPR Article 5(1)(d) requires personal data to be accurate and, where necessary, kept up to date. Duplicates are not automatically unlawful, but unresolved conflicting records can undermine accuracy and rights handling.

A practical risk arises when a rights request is applied to one record but a linked record is missed. Article 17 also contains conditions and exceptions, so route the request through the organisation's approved process rather than promising automatic deletion.

Building a Deduplication Culture

Technology solves most of the problem. But without process, new duplicates accumulate as fast as old ones are cleaned up. The agencies with the cleanest databases share a few habits:

  • Search-before-create rule: Every new profile starts with a search. No exceptions.
  • Import policies: Define who can bulk-import LinkedIn profiles and under what conditions. Mass imports without dedup checks are banned.
  • Quarterly audits: Run the batch deduplication process every quarter. Document the results.
  • Clear ownership: Someone is explicitly responsible for data quality — not as a side task, but as a core responsibility.

For teams choosing between an ATS and CRM approach, our ATS vs CRM comparison covers how each handles data quality differently. And if you're working with enriched candidate data, the database cleanup and enrichment guide dives deeper into keeping imported data clean.

Useful operating rule: search before creating a profile, record the import source, and assign every proposed merge to a named reviewer. Track results monthly to see whether the rule works.

FAQ: Duplicate Candidate Detection

How many duplicates does a typical recruitment database have?

There is no reliable universal percentage. Sample your own records and report confirmed duplicate pairs separately from stale or incomplete profiles.

Can duplicate detection work across different name spellings?

Name-based fuzzy matching can catch some variations, but it's unreliable for name changes like marriage or legal changes. That's exactly why strong signal matching — email, phone, LinkedIn — is more effective. These identifiers persist even when names change.

What happens to GDPR consent when merging duplicates?

Do not reduce lawful basis and communication status to one “consent” flag. Preserve the provenance, purpose, lawful basis, notices, objections and suppression status from both records; resolve conflicts under your approved policy and document the merge.

Should duplicates be merged automatically or manually reviewed?

Automatic linking may be suitable for exact, tested signals, but destructive merging should be reversible and policy-controlled. Use manual review whenever fields, identity evidence or rights status conflict.

How often should I run deduplication on my database?

Real-time checks on every new record, plus a quarterly batch scan across the full database. If your database is growing rapidly with more than 100 new profiles per week, monthly batch scans are worth the effort.

Detect duplicates by signal, not by name

When evaluating Yena, test duplicate handling with exact and conflicting signals from your own sample. Confirm what gets flagged, what remains reviewable and whether a merge can be audited or reversed.

See current plans

Janis Kolomenskis

March 24, 2026

Share
Yena

Turn a role brief into a qualified shortlist.

Describe who you need. Yena finds and ranks candidates, explains why they fit, surfaces available contact details for review, and keeps outreach in the same recruiting workspace.