Here's a number that should make you uncomfortable: according to Dun & Bradstreet's long-running data quality research, professional contact data decays at roughly 30% per year. People change jobs, change roles, change names, change email addresses, get promoted, leave the industry entirely.
If your ATS database is three years old and you haven't run a systematic cleanup, up to 65% of your candidate records may contain at least one stale data point. Email addresses that bounce. Phone numbers that go to someone else. Job titles that are two roles out of date.
Most recruiters know their database is messy. Few have a clear plan for fixing it — and even fewer understand the GDPR dimension, which in Europe turns "database cleanup" from an operational nicety into a legal obligation.
This guide covers the cleanup process end to end: what to fix, in what order, how to handle data deletion under GDPR, and how to prevent the mess from rebuilding itself.
Why Your Database Gets Messy (It's Not Just Your Team)
Before getting into the fix, it's worth understanding why recruitment databases degrade faster than most other business databases. Several factors compound together:
Multi-source candidate ingestion. A single candidate might enter your database from a job board application, a LinkedIn sourcing export, a referral email, and a CV attached to a speculative inquiry — all at different times with different data completeness. Without automatic deduplication, you have four records for the same person.
Consultant turnover. When a recruiter leaves, their naming conventions, tagging preferences, and note formats leave with them. What's left behind are records that the next person can't reliably interpret or search.
Import-heavy onboarding. Many agencies, when switching ATS platforms, bulk-import their old database without cleaning it first. Every existing duplicate, stale record, and inconsistent field format comes along for the ride and immediately degrades the new system's usefulness.
No structured data retention policy. Without a clear rule about how long candidate data is stored and for what purpose, records accumulate indefinitely. It's not unusual to find active ATS databases containing CVs from candidates who applied in 2010 — data you legally cannot justify retaining and practically cannot use.
The GDPR Dimension: This Isn't Optional
Let's address this directly before the operational cleanup steps, because in Europe, data purging isn't just good database hygiene — it's a legal requirement.
GDPR's storage limitation principle (Article 5(1)(e)) requires that personal data be kept in a form which permits identification of data subjects for no longer than is necessary for the purposes for which it is being processed. For recruitment databases, this means you need a defensible answer to: why are you still holding this candidate's personal data, and for how long?
Lawful Basis and Retention Periods
Most recruitment agencies process candidate data under one of two lawful bases: legitimate interest (you have a business reason to maintain a talent pool) or explicit consent (the candidate agreed to be in your database).
Either way, retention can't be indefinite. A practical, defensible retention schedule looks something like this:
- Active applicants (applied to a specific role, within the process): Data can be retained for the duration of the hiring process plus typically 6–12 months for "silver medallist" re-engagement.
- Unsuccessful applicants (no ongoing relationship): Most data protection authorities suggest 6–12 months as a reasonable retention period after a rejection, unless the candidate has explicitly consented to longer-term inclusion in a talent pool.
- Talent pool candidates (added proactively, consent obtained): Typically 2 years from the last positive engagement or explicit consent refresh — whichever is more recent.
- Placed candidates: Retention for the duration of any contractual obligation (placement guarantee period, typically 3–6 months) plus your organisation's standard business records retention period.
The UK ICO's guidance on GDPR in recruitment is worth reading in full if you haven't. The German BFDI has published similar guidance for DACH-based agencies. Both are clear that indefinitely retaining candidate data without a documented lawful basis is a compliance risk.
Right to Erasure — What It Actually Means Operationally
Under Article 17, candidates have the right to request deletion of their personal data in specific circumstances — including when the data is no longer necessary for the original purpose, or when they withdraw consent.
Operationally, this means your cleanup process must include a mechanism to identify and honour erasure requests. Not just delete the candidate record, but also:
- Remove their data from any backup systems within a reasonable timeframe
- Suppress their data so it doesn't get re-ingested (e.g., if they appear again on a job board)
- Log the deletion with a timestamp for audit purposes
- Notify any third parties you've shared their data with (referencing companies, background check providers)
This is easier to manage systematically than on an ad hoc basis. Build the erasure workflow before you need it, not in response to an angry LinkedIn message from a candidate.
The Cleanup Process: What to Do and in What Order
A database cleanup done wrong creates new problems — merged records that shouldn't have been merged, deleted data that had legitimate purposes, compliance issues introduced by the cleanup itself. Sequence matters.
Step 1: Audit Before You Act
Run a database health report before touching anything. You need to know:
- Total record count
- Percentage of records with a valid email address (test deliverability, not just format)
- Percentage with a valid phone number
- Percentage with no activity logged in the past 24 months
- Estimated duplicate rate (most ATS platforms have a built-in duplicate detection tool)
- Percentage added before your current retention policy began
This audit gives you your cleanup priority list and a baseline to measure against once you're done. It also tends to be a useful internal advocacy tool — showing leadership a concrete number like "31% of our candidate records have no activity in 3 years" is more persuasive than "our database is a bit of a mess."
Step 2: Purge Legally Non-Retainable Records First
Before any deduplication or enrichment work, remove the records you legally shouldn't be keeping. This usually means:
- Candidates who explicitly requested deletion but weren't properly removed
- Records older than your retention policy with no documented consent for longer retention
- Records where the original lawful basis no longer applies (e.g., the candidate was added under a client contract that ended years ago)
Do this first because any work you do enriching or deduplicating records that should be deleted is wasted effort — and potentially increases your compliance exposure by demonstrating you're actively working with data you shouldn't hold.
Step 3: Deduplicate
Duplicate records are the single most common problem in agency recruitment databases. HR Dive's 2024 ATS data quality survey found that the average recruitment agency database contains a 12–18% duplicate rate — meaning roughly 1 in 7 records is a duplicate of an existing one.
Deduplication has two stages: automated matching and manual review.
Automated matching flags probable duplicates based on combinations of email address, phone number, LinkedIn URL, and name similarity. This catches the obvious cases — same person, slightly different name format, two records with the same email.
Manual review handles the ambiguous cases. Two people with the same name at the same company. A candidate who changed their email and phone number so nothing matches automatically. Identical names but clearly different profiles. You cannot automate this layer; trying to will result in merged records that shouldn't have been merged.
When merging, establish a clear rule for which record's data takes precedence. Usually: most recent data wins for contact information; most complete record wins for history. Document your merge logic so you can explain decisions later.
Step 4: Standardise Fields
This is less glamorous than purging and deduplication but often has a bigger impact on search quality. If your job title field contains "Head of Engineering," "Head, Engineering," "VP Engineering," "Engineering Head," "Engineering VP," and "Eng VP" — all referring to the same type of role — none of your searches will return complete results.
Create a controlled taxonomy for your highest-variance fields: job titles, seniority levels, skills, and locations. Map existing variations to the canonical form. This is time-consuming the first time; automated normalisation tools can handle much of it once you've defined your taxonomy.
German and Polish location names deserve specific attention if you recruit across DACH and CEE markets. "München" vs "Munich," "Warszawa" vs "Warsaw," "Köln" vs "Cologne" — these should all resolve to the same searchable entity, or you'll consistently miss candidates based on which spelling a recruiter happened to use.
Step 5: Re-engagement Campaign Before Final Archival
Before archiving or deleting records that are old but not yet legally required to be deleted, consider a re-engagement campaign. A brief email to candidates who haven't had any activity in 18–24 months: "We still have your details in our system. Are you still open to hearing about opportunities, or would you prefer we removed your record?"
This serves two purposes. It refreshes consent for candidates who want to stay in your database, turning a grey-area legitimate-interest retention into a clear consent-based one. And it produces a clean suppression list — candidates who don't respond (or who explicitly opt out) can be archived with confidence.
Expect roughly 15–25% of candidates to actively confirm they want to remain. The rest are de facto confirmation that their record isn't worth maintaining.
Step 6: Archive, Don't Always Delete
There's a useful distinction between deletion and archival. Some records that you can't actively market to still have value — a candidate you placed three years ago, for example, where you need to retain the record for contractual and tax purposes even though you're no longer using their personal data for recruitment.
Archiving means removing the record from active search and recruitment use, logging why it's archived and when it should be permanently deleted, and ensuring it's not returned in normal search queries. This keeps your active database clean without creating compliance gaps in your historical records.
Preventing the Mess From Rebuilding Itself
A database cleanup is expensive in time and effort. The most important return on that investment is making sure you don't need to repeat it at the same scale in two years.
Three systemic changes make the biggest difference:
Mandatory deduplication check before record creation. Any recruiter adding a new candidate should be required to confirm no existing record exists. Your ATS should enforce this — not just suggest it. If the system lets people skip the check, they will.
Automated retention flags. Set your ATS to automatically flag records approaching their retention limit 90 days in advance. The recruiter responsible for the record gets a notification: review this candidate's status and either refresh consent or schedule for archival. Not an annual fire drill — a continuous rolling process.
Data entry standards, enforced technically. Documented style guides don't work if the system accepts non-standard input anyway. Use dropdown fields with controlled vocabularies for job titles, seniority levels, and skills. Accept free text only where genuinely necessary. Every field that allows arbitrary text is a future standardisation headache.
What a Clean Database Is Actually Worth
It's worth quantifying the upside, because database cleanup is the kind of project that gets deprioritised constantly in favour of more immediately visible work.
A recruiter who can reliably search their database and surface relevant candidates in under 5 minutes has a meaningful competitive advantage over one who spends 45 minutes searching the same database and still misses the best match because it was filed under the wrong job title.
More concretely: ERE Media's 2024 analysis of database reactivation campaigns at 15 mid-size European agencies found that agencies with clean, well-segmented talent databases filled an average of 23% of new roles from existing records — compared to 8% at agencies with unmanaged databases. That 15-percentage-point difference represents a significant reduction in job board spend and sourcing time.
Your database isn't just a compliance obligation. It's the most under-leveraged asset most agencies own. Yena's data enrichment tools automate the ongoing maintenance — keeping contact information fresh, flagging stale records, and surfacing re-engagement opportunities before records go cold.
The cleanup is a one-time effort. The compounding returns from a database you can actually use last for years.
About the author: Janis Kolomenskis is the founder of Yena, an AI-native ATS built for European recruitment teams. He wrote this guide because he's seen too many agencies ignore their database until a GDPR audit or a failed reactivation campaign forced the issue. Better to do it on your own timeline.