
The demo is always impressive. A vendor rep types a free-form description of a senior actuary with FIA credentials in Munich, hits enter, and twenty perfectly matched profiles appear in under a second. You leave the call excited. Three months into the contract, your team is manually tweaking queries to avoid a flood of irrelevant results, and time-to-shortlist hasn't moved. What happened?
Here's what happened: demos run on curated, pre-cleaned databases where the signal-to-noise ratio doesn't reflect your actual candidate pool. Your reality involves 50,000 partially completed profiles, inconsistent job titles, and mandates in niche specialisations where the candidate universe is genuinely small. No vendor shows you that version.
According to a 2025 AI in Hiring report from Insight Global, 62% of HR teams that adopted AI tools in the past two years felt the results didn't match what was shown during the sales process. That gap isn't always dishonesty — it's often a mismatch between demo conditions and production conditions. This guide gives you the framework to close that gap before you sign.
Why Vendor Demos Don't Tell You What You Need to Know
AI sourcing vendor demos are structurally misleading because they run on curated, pre-cleaned databases using pre-selected queries with no false-positive accounting. Your real candidate pool contains 50,000 partially completed profiles and niche mandates that expose accuracy gaps the demo was designed not to show.
Three things make AI sourcing demos structurally misleading for buyers.
Curated databases, not your database
Vendors demo against their hosted talent database — typically a well-structured, continuously enriched pool of profiles scraped from public sources. Your existing candidate database, by contrast, contains years of inconsistently entered data, stale profiles, duplicate entries, and job titles your industry invented six months ago. The AI's accuracy against a clean external database says nothing about how well it'll handle your legacy data.
This matters most if you're evaluating AI matching against your internal pool. Run the demo on your data. If the vendor won't allow it, that tells you something.
Pre-selected demo queries
Vendors pick queries they know will perform well. A "VP of Engineering in London" search will always surface a full, relevant shortlist — there are thousands of matching profiles and the role is well-understood by any language model trained on English professional data. Ask the rep to run "Head of Actuarial Risk for a Liechtenstein-based reinsurer, German-speaking, FSA or FIA qualified" and watch what happens. That's closer to what your team actually works on.
No false-positive accounting
Demos show you the top 10 results, and those 10 look great. What they don't show you: how many results were returned in total, and how quickly quality degrades past position 10. A system that returns 500 results, 10 of which are relevant, will waste hours of your team's time even though the demo looked clean. Precision at rank 10 is not the same as precision across the full result set.
The 4 Metrics That Actually Matter
The four AI sourcing metrics that translate directly into whether your team saves or wastes time are: Precision at Top 10 (how many of the first ten results are genuinely relevant), Recall (what fraction of truly qualified candidates are surfaced), Time-to-Shortlist (the whole-process duration), and False-Positive Rate (clearly wrong profiles reaching the top of results).
Forget accuracy as a standalone number. It's meaningless without context. Here are the four metrics that translate directly into whether your team saves time or wastes it.
Precision @ Top 10
Of the first 10 candidates the system surfaces for a given query, how many would your consultants actually consider relevant? This is the number that determines daily workflow. If your team has to scroll past irrelevant results to find useful ones, the tool is creating friction rather than removing it.
A reasonable benchmark: 7 out of 10 should be worth at least a second look. Anything below 5 means your sourcers will stop trusting the tool within weeks.
Recall
This is the metric vendors rarely volunteer. Recall asks: of all the truly relevant candidates in the database, what percentage does the system actually surface? A system with high precision but low recall will give you a clean shortlist — but miss half the candidates who could be excellent placements.
Recall is hard to measure without knowing the "true" set of relevant candidates. The bake-off method below gives you a practical proxy.
Time-to-Shortlist
How long does it take from starting a search to having a shortlist of 8-12 qualified candidates ready for client review? This is a whole-process metric that includes query construction time, result review, qualification, and outreach prep. The LinkedIn Future of Recruiting 2025 report found that AI sourcing tools reduced sourcing workload by 20% on average — but that average masks wide variance, with some teams seeing 40%+ reductions and others seeing no measurable improvement.
Measure this in your own bake-off, not the vendor's case study.
False-Positive Rate
How many results in the top 20 are genuinely off-target? Not borderline cases — clearly wrong profiles that made it to the top of the ranked list. Even one or two clearly irrelevant results in the top 10 signals that the ranking algorithm doesn't understand your query as well as the demo suggested.
False positives are expensive in executive search. Your team's time is the constraint. A single sourcer reviewing 15 irrelevant profiles per search, across 10 active mandates per week, burns 6-8 hours that should go to candidate engagement.
How to Design Your Own Vendor Bake-Off
A vendor bake-off for AI sourcing accuracy uses five closed roles you've already filled, runs each query across competing platforms, and checks whether the actual placed candidate appears in the top 20 results. This test takes roughly one day of senior consultant time and is the only reliable way to predict real-world performance before signing a contract.
This is the most valuable thing you can do before committing to any AI sourcing platform. It takes about a day of senior consultant time, and it's worth every minute.
Step 1: Select 5 closed roles
Pick 5 mandates you've already filled in the last 12-18 months. The criteria: each should have been a genuine search (not an easy lateral referral), and you should know who the actual placed candidate was. Ideally, pick roles at different seniority levels and in different functional areas.
Why closed roles? Because you have a ground-truth answer. You know who the right candidate was. You can check whether the AI would have surfaced them.
Step 2: Reconstruct the search query
For each of the 5 roles, write the query you'd actually use in the tool you're evaluating. Don't engineer it for the system — write it the way your consultants would naturally describe the ideal candidate. This tests real-world usability, not edge-case capability.
Step 3: Run each query, record the top 20 results
For each query, note the position at which the actual placed candidate appears (if they appear at all). Also note how many of the top 20 you'd consider genuinely relevant vs. clearly off-target.
The scoring framework:
- Placed candidate in top 5: Excellent recall for this role type
- Placed candidate positions 6-10: Acceptable — tool would still surface them in normal review
- Placed candidate positions 11-20: Concerning — tool found them, but would consultants review that far?
- Placed candidate not in top 20: Poor recall — the right answer wasn't surfaced
Step 4: Aggregate and compare
Run the same 5 roles across each vendor you're evaluating. The results are now directly comparable. You'll quickly see whether a platform that shone in the demo actually performs on your specific role types and seniority levels.
"The bake-off isn't about finding the perfect tool. It's about eliminating the tools that will waste your team's time and erode their trust in AI before it has a chance to work."
Good vs. Bad AI Sourcing Accuracy: The Comparison Table
Strong AI sourcing platforms return 7 to 9 genuinely relevant results in the top 10, surface the placed candidate in 4 of 5 bake-off test roles, handle synonym expansion correctly, and disclose their ranking methodology. Weak platforms produce round-number accuracy claims with no methodology, can't explain why a candidate ranked where it did, and have no performance data outside a single geography or industry.
| Signal | Strong AI Sourcing Platform | Weak / Overpromised Platform |
|---|---|---|
| Precision @ Top 10 | 7-9 of 10 results genuinely relevant | 3-5 of 10 results relevant; rest are demographic matches only |
| Recall on bake-off | Placed candidate appears in top 10 for 4 of 5 test roles | Placed candidate not surfaced in top 20 for 2+ test roles |
| Handling ambiguity | Expands synonyms correctly ("actuary" includes "risk modeller", "reserving analyst") | Returns only exact-title matches; misses equivalent roles |
| Methodology disclosure | Vendor explains ranking logic, data sources, refresh frequency | "Proprietary AI" with no further explanation |
| Accuracy claims | Specific, benchmarked claims with methodology footnotes | "95% accuracy" with no definition of what's being measured |
| Non-English performance | Similar result quality for German, Polish, French searches | Strong English, noticeably weaker for non-English profiles |
| Trial access | Willing to let you trial with your own data | Demo only, no access to your own database during trial |
Red Flags in Vendor Pitches
Red flags in AI sourcing vendor pitches include round-number accuracy claims (90%, 95%) with no methodology footnote, inability to explain individual candidate rankings, case studies confined to a single industry or geography, and no acknowledgement of false positives. Each signals either that performance hasn't been measured honestly or that it has and the vendor doesn't want you to know the result.
These aren't automatic disqualifiers — but each one should prompt a direct follow-up question before you go further.
Round-number accuracy claims
"90% accuracy." "95% match rate." Any claim ending in a round number with no methodology footnote is a marketing figure, not a measurement. Ask: what is being measured (precision, recall, or something else)? Measured against which benchmark? On whose data? What was the sample size?
A vendor who can answer those questions specifically has probably actually measured it. One who deflects to "our customers love the results" hasn't.
No ability to explain ranking
If your team asks "why did this candidate rank above that one?" and the vendor can't give a coherent answer, your consultants won't be able to improve their queries. The tool becomes a black box that either works or doesn't — with no path to getting better over time.
Case studies from a single industry or geography
A vendor whose entire proof portfolio is tech companies in London or SaaS startups in Berlin may not have meaningful performance data for financial services in Vienna or manufacturing in Kraków. Ask for case studies relevant to your mandate types.
No discussion of false positives
Every AI sourcing tool produces some irrelevant results. A vendor who only talks about the hits and never acknowledges the misses is either not measuring false positives or doesn't want you thinking about them. The question "what does a typical false positive look like in your system, and why does it happen?" reveals a lot.
"Adoption rates for AI hiring tools have reached 57% among UK businesses, according to a UK government survey — but widespread adoption doesn't mean widespread satisfaction. The gap between deployment and genuine utility is where most recruiting teams are still stuck."
The UK Government's 2025 AI Labour Market Survey found that 57% of UK businesses now use AI in at least one HR function — but satisfaction is unevenly distributed across tool categories, with sourcing tools receiving notably more mixed feedback than document processing tools.
Understanding what AI candidate matching actually does technically helps you ask smarter questions during vendor evaluation. The vendors who talk openly about vector embeddings and similarity thresholds are usually the ones who've actually built the system rather than licensed someone else's black box.
What Good Looks Like in Practice
Good AI sourcing accuracy in practice means transparency about data provenance and refresh frequency, explainability so consultants understand why a candidate ranked where they did, willingness to engage with failure cases during evaluation, and — for European teams — consistent multilingual quality across German, Polish, and French searches, not just English.
The platforms that tend to perform well in honest bake-offs share a few characteristics. They're transparent about data provenance — they can tell you where their profile data comes from, how often it's refreshed, and what percentage has verified contact information. They offer explainability: consultants can see why a candidate ranked where they did. And they're willing to engage with failure cases during evaluation, not just showcase wins.
For European teams specifically, the additional dimension is multilingual quality. A tool that works brilliantly for English-language searches but struggles with German compound nouns or Polish diacritics will create a two-tier experience for your team — great for UK mandates, frustrating for DACH or CEE work. Checking this during the bake-off with at least 2 of your 5 test roles in non-English markets is worth the extra time.
The ONS research on AI and employment found that 23% of UK businesses plan to expand AI tool usage in HR in 2025-26. The ones who get lasting value from it are those who evaluated rigorously before buying — not those who were most impressed by the demo.
If you're building out your sourcing tech stack more broadly, the full comparison of AI sourcing tools for European recruiters covers the major platforms side-by-side on the criteria that matter for non-US markets. And if you want to understand where agentic AI fits in the sourcing workflow — the autonomous version that takes a brief and runs an entire outreach campaign — that's worth reading alongside this framework.
Frequently Asked Questions
The most common AI sourcing evaluation questions cover how long a rigorous evaluation should take, whether bake-offs work without a large internal database, the difference between matching accuracy and sourcing accuracy, how much to trust vendor benchmarks, and which role types see the strongest AI sourcing results.
How long should a proper AI sourcing evaluation take?
The bake-off itself takes 3-5 working days if you're running it in parallel across 2-3 vendors. Before that, you need a week to select your test roles and reconstruct queries. Budget 3 weeks total from first vendor conversation to decision. Rushing this step is how firms end up locked into a 12-month contract with a tool that sounded better than it performed.
Can I evaluate AI sourcing accuracy if I don't have a large internal database?
Yes. If you're evaluating sourcing against an external talent database (rather than your own), use the closed-role bake-off against the vendor's database. You won't be able to check internal recall, but you can still measure precision at the top 10 and false-positive rate on your specific role types. The test roles don't need to have been filled from an AI platform — they just need a known ground-truth answer.
What's the difference between AI matching accuracy and AI sourcing accuracy?
Sourcing accuracy is about finding candidates you didn't already know about — precision and recall on searches against external databases. Matching accuracy is about ranking known candidates against a new role (typically within your existing ATS). They use similar underlying technology but test different capabilities. If you're evaluating both, run separate tests — a tool can excel at one and underperform on the other.
Should I trust vendor-provided accuracy benchmarks?
Use them as a starting point, not a decision driver. Vendor benchmarks typically measure performance on the vendor's test set, which is almost always more favourable than your real data. Treat them as a lower bound on variance: if a vendor claims 85% precision and you see 60% in your bake-off, that's a significant gap worth investigating. If you see 82%, the benchmark is probably honest.
Is AI sourcing accuracy better for some role types than others?
Yes. Roles with well-established titles, clearly defined skill sets, and high profile density (senior software engineers, finance professionals, HR directors) tend to return better AI sourcing results than roles with emerging titles, highly niche skill combinations, or thin profile coverage in specific geographies. This is worth factoring into both your evaluation and your eventual tool selection — a platform with excellent accuracy for technical roles may disappoint on specialist operations or compliance mandates.
If you want to see how a platform performs against your actual mandates — not a pre-selected demo scenario — book a working session with the Yena team. We'll run the bake-off with you using your own test roles, against Yena's 1B+ profile database, and show you the methodology behind the rankings. No curated data. No round-number claims.