Why every AI sourcing demo looks the same
You have probably sat through three or four of them this year. A polished deck. A live screenshare. A pre-built brief that runs in 90 seconds. A shortlist of suppliers that look real, in a country you care about, with names and emails. "Want to see another category?" "Sure." Another 90 seconds, another impressive list.
The problem is that every AI sourcing vendor can build that demo. The underlying search + LLM pipeline is not proprietary technology; it is commodity components arranged with different levels of care. What separates real capability from a good demo is what happens on the 50th campaign, the week your integration partner changes an API, the Tuesday a supplier's VAT returns invalid and you need to understand why.
This guide is 12 questions structured so vendors cannot dodge them. Each question has a "good answer," a "yellow flag," and a "walk-away" answer. Use it as a scorecard: 3 points for good, 1 for yellow, 0 for walk-away. Max 36. A vendor scoring under 24 is a no. A vendor scoring 36 is either perfect or lying, reference-check hard.
Structure: four themes, three questions each. Data sources, AI + transparency, integration + workflow, commercials + customer proof.
- Q01 · DataWhere does your supplier data come from?
- Q02 · DataHow fresh is each record?
- Q03 · DataHow many regions are covered?
- Q04 · AIWhich model powers your scoring?
- Q05 · AICan we see the reasoning?
- Q06 · AIHow do you handle hallucinations?
- Q07 · WorkflowHow does it integrate with our ERP?
- Q08 · WorkflowCan we export to our CRM?
- Q09 · WorkflowHow do we send outreach?
- Q10 · ProofWhat does a pilot really cost?
- Q11 · ProofShow me 3 references in our industry.
- Q12 · ProofHow long is a realistic ramp-up?
Theme 1, Data sources (3 questions)
Q1. Which web sources do you actually crawl?
Good answer: A specific list. Google + regional directories (wlw.de, panorama-firm.pl, europages) + trade registries (KRS, Handelsregister, Companies House) + language-specific sources per country. Vendor names sources and explains how each is incorporated.
Yellow flag: "We have partnerships with data providers" without naming them. This usually means one paid commercial dataset (Dun & Bradstreet, ZoomInfo, or an aggregator) rebranded as "AI search."
Walk-away: "Our AI finds suppliers across the web" with no specificity. Translation: we run Google Custom Search Engine and prompt an LLM. That is commodity infrastructure and will not out-perform your own searches at scale.
Q2. Do you scrape directories yourself, or license their data?
Good answer: Combination. Scrape the public web (Google results + crawlable sites), license structured data where it exists (registry APIs, VIES for VAT), supplement with verified directory scraping where ToS permits.
Yellow flag: "We only use licensed data." Licensed data is clean but narrow, it misses the long tail of manufacturers who do not pay for directory inclusion.
Walk-away: "We scrape everything including premium directories behind login walls." That is a ToS violation waiting to become a lawsuit, and if the vendor talks openly about it, they are either naive or careless with compliance.
Q3. How do you handle non-English sources?
Good answer: Multilingual search queries generated per country, native-language content scraping, translation done client-side with the original text preserved in the record. The vendor can show you a German-language supplier homepage feeding into the shortlist with the description in both German and English.
Yellow flag: "We translate queries with DeepL and search Google." It works, but it is thin, you miss suppliers whose homepages only rank in their local language SERPs.
Walk-away: "Our primary focus is English-language sources." You will miss 60-70% of the real European supplier base. No AI polish fixes that.
A missing certificate never drops a good maker. It lowers a score, and the score is a number you can argue with.
Theme 2, AI + transparency (3 questions)
Q4. Which LLM do you use and is it swappable?
Good answer: Specific model name (GPT-4, Gemini 2.0 Flash, Claude 3.5) with a multi-provider fallback. Vendor explains why they use the specific model for specific tasks (e.g., Gemini for search strategy, GPT-4 for relevance screening). Willing to swap if you have a policy preference.
Yellow flag: "We use a proprietary model." If they trained a foundation model from scratch, they will happily tell you. If "proprietary" means a fine-tune of an open-weights model, fine, but they should say so.
Walk-away: "Our AI is proprietary, we cannot discuss the details." This is usually a cost-optimization story, they are using GPT-3.5 or a small open model to keep margins high and do not want to admit it.
Q5. Can I see the reasoning for each supplier's relevance score?
Good answer: Yes. Vendor shows a per-supplier explanation like "scored 0.87 relevance because: homepage mentions injection molding, company size 45 employees matches your 20-500 range, ISO 9001 listed." You see the signals.
Yellow flag: "You can see the score but not the full reasoning." This usually means the pipeline works but logging was an afterthought. Workable; not ideal.
Walk-away: "The scores come from our AI model, we cannot decompose them." Translation: there is no rational audit trail. When internal audit asks how you qualified this supplier, you will have nothing to show.
Q6. What is your precision and recall on a category like mine?
Good answer: Specific numbers with a published methodology. "On metal fabrication in Poland, precision ~88%, recall ~72%." Vendor shows how they measured (sampling, blind-coding by procurement specialists, etc.). Admits where they perform worse and why.
Yellow flag: Generic benchmarks "95% accurate" with no definition. Accuracy is not precision or recall; using accuracy for classification tasks is a pedagogical red flag.
Walk-away: "We do not measure precision and recall." Any vendor serious about an AI pipeline measures these. If they do not, they are not operating a quality process, they are running a demo pipeline and calling it production.
Theme 3, Integration + workflow (3 questions)
Q7. Which ERPs have productized (not "we can build it") connectors?
Good answer: Specific list with status per ERP. "NetSuite, productized via Merge.dev. SAP S/4HANA, pilot. Microsoft Dynamics: CSV only. Oracle Fusion, not supported." Honest about what is shipping, what is pilot, what is custom.
Yellow flag: "We integrate with all major ERPs." Every vendor says this. It usually means they have done a custom integration for one customer with each ERP and will do the same for you at an unnamed cost.
Walk-away: "ERP integration is part of our enterprise tier, details in the statement of work." Translation: they have no productized connectors; every integration is bespoke and will take 3-6 months at consulting rates.
Q8. What does your CSV export schema look like?
Good answer: Here is the schema. Specific field names. Shows you a real CSV from a real campaign. Fields include everything you need for an ERP import (company name, VAT, address, primary contact, category, certifications, trust score).
Yellow flag: "We support CSV export" without showing the schema. Usually means the export exists but is thin.
Walk-away: "Export is on our roadmap." If export is roadmap, the tool is a prototype. Walk.
Q9. Which plan includes API access?
Good answer: Specific plan names and what the API exposes. "API access on our Growth plan (€299/month), read access to campaigns and supplier records. Write access on Enterprise." Documentation link provided.
Yellow flag: "API access available on Enterprise plans." Usually fine, but ask for the API docs. If they cannot share docs before the contract, the API is either not stable or not real.
Walk-away: "API access is custom per customer." This means there is no API; every customer gets a snowflake endpoint that breaks when the vendor changes internals.
Theme 4, Commercials + customer proof (3 questions)
Q10. Is pricing transparent on your website?
Good answer: Yes. Specific plan names, specific prices, what each includes. Enterprise custom pricing is fine as long as Starter and Growth tiers are public.
Yellow flag: "Contact sales for pricing." Everywhere. For every plan. Usually means they price-discriminate heavily, and you should expect the first quote to be 2-3x the right number.
Walk-away: "Our pricing is based on your spend volume." This is almost always a % fee on sourced spend, which turns into a 5-8% tax on every category you put through the tool. For a €20M procurement spend, that is €1M/year, which is Ariba money.
Q11. Credit-based or seat-based pricing?
Good answer: Clear model. Credit-based (pay per campaign, buy in packs) is friendlier for variable-usage teams. Seat-based (flat per-user per-month) is friendlier for heavy daily users. Either is fine if the vendor explains why they chose it and the economics match your usage.
Yellow flag: Hybrid models with 4 different meters. Usually a sign the pricing was designed by a CFO optimizing revenue capture, not a product team optimizing customer value.
Walk-away: "It is custom per customer." You will be the price-discrimination test case.
Q12. Show me three live customers in my size range.
Good answer: "Here are three customers, similar industry, similar size. They have agreed to 30-minute reference calls. Here are their names." The vendor schedules the calls and does not sit in on them.
Yellow flag: "We have customer logos we can share." Logos are marketing, not references. Always push for actual calls.
Walk-away: "We cannot share customer details due to NDA." Every customer signs an NDA; that does not prevent reference calls. If the vendor cannot produce three customers willing to talk, they probably do not have three customers the right size for you.
| # | Question | Theme | What a good answer looks like |
|---|---|---|---|
| Q1 | Which web sources do you actually crawl? | Data | Names specific sources: Google + directories + registries |
| Q2 | Scrape or license directory data? | Data | Hybrid: public web + licensed registries (VIES, IAF) |
| Q3 | How do you handle non-English sources? | Data | Multilingual queries per country, native-language scraping |
| Q4 | Which LLM, and is it swappable? | AI | Specific model named (GPT-4, Gemini) with multi-provider fallback |
| Q5 | Can I see reasoning for each score? | AI | Per-supplier score decomposition visible in UI |
| Q6 | Precision and recall on a category like mine? | AI | Published numbers with methodology (e.g. 88% / 72%) |
| Q7 | Which ERPs have productized connectors? | Integration | Per-ERP status (productized / pilot / CSV / unsupported) |
| Q8 | What does your CSV export schema look like? | Integration | Shows a real CSV; schema documented |
| Q9 | Which plan includes API access? | Integration | Specific plan + public docs link before contract |
| Q10 | Is pricing transparent on your website? | Commercial | Public Starter + Growth tiers; Enterprise custom OK |
| Q11 | Credit-based or seat-based pricing? | Commercial | Clear model; rationale matches your usage pattern |
| Q12 | Show me three live customer references. | Commercial | Schedules 30-min calls; vendor does not sit in |
The honesty test (ask every vendor this)
Beyond the 12 questions, one meta-question reveals more about the vendor than anything else: "What is the number-one thing your product is NOT good at?"
A mature vendor will have an answer. "We are weak on categories under 10 global suppliers, our AI assumes a larger universe than those categories have." "Our Salesforce integration is pilot, not productized, so if you need real-time bidirectional sync on day one, we are not there yet." "We do not support Oracle Fusion at all, if that is your ERP, do not buy us."
Those answers are selling. They demonstrate that the vendor has self-awareness, measured their product, and respects your decision-making. A vendor who cannot name a weakness is either lying about having no weaknesses, or does not understand their product well enough to have found them. Both are disqualifying.
The other meta-question worth asking: "What would cause you to recommend we not buy your product?" The best vendors have this answer ready. "If your procurement is 80% services, if you have 40+ suppliers per year across complex categories but your budget is under €5k/year, if you need deep S/4HANA sync on week one, in all three cases, we are not the right fit."
A vendor willing to talk you out of a bad-fit purchase is more likely to still care about you in year two. A vendor who pretends every customer is a good fit will lose interest the moment the contract signs.
The reference-check protocol that actually works
A logo slide is not a reference. A 30-minute call with an actual customer is. Three references, 30 minutes each, specific questions.
Ask each customer:
- How long have you been using the tool? Less than 6 months = too new to matter. 6-18 months = peak honeymoon. 18+ months = real signal.
- What did you try before? Reveals what problem the tool actually solved versus what the vendor claims it solves.
- What is the #1 thing you wish the tool did better? If the answer is "nothing" or a feature request that sounds like marketing fluff, the reference was coached. If the answer is specific and technical ("the email follow-up cadence is too aggressive for our German suppliers"), the reference is real.
- Has the vendor ever failed you? How did they handle it? Vendors handle failures either well (clear root cause, timely fix) or badly (deflection, blame). Customers remember.
- Would you buy it again today? Ask this at the end. The answer and the hesitation both matter.
Three calls, 90 minutes total, will tell you more than five hours of vendor demos. Do not skip this step; the vendors who are confident in their product will facilitate it, and the vendors who dodge it are telling you what you need to know.
Score each vendor 1-5 on all 12 questions
3 = good answer, 1 = yellow flag, 0 = walk-away. Max 36. Fill the scorecard during the demo, not after, memory degrades fast.
Weight by your top 3 priorities
Pick the 3 questions most critical for your use case (e.g. Q6 precision, Q7 ERP, Q12 references). Double-weight them in the total.
Calculate weighted total per vendor
Total + weighted bonus. Anyone under 24 (unweighted equivalent) is a no. Anyone over 34 is suspicious, reference-check hard.
Eliminate bottom 2, demo final 2
If you have 4-5 vendors, drop the bottom 2. Book 60-min deep demos with final 2 using your own real brief, not theirs.
Reference calls + honesty test
3 references × 30 min per finalist. Ask meta-question: "what is your product NOT good at?" Mature vendors answer honestly.
Procurea answers all 12, and fails two of them
Applying the scorecard to ourselves. This section used to award us 34 out of 36. That was written when we intended to build things we have since decided not to build, and it stayed on the site long after it stopped being true. Here it is against what exists today.
- Q1 data sources. The live web through Serper, company pages read directly, regional directories, and VIES for VAT numbers. Named, not "proprietary". 3/3.
- Q2 scrape vs license. We read the public web and respect robots.txt; VIES is the European Commission's own service. 3/3.
- Q3 non-English. Queries and reading in local languages, which is the whole point. The published runs turn up German, Italian, Turkish and Latvian makers a monolingual search does not reach. 3/3.
- Q4 LLM swappable. Vertex primary with an AI Studio fallback, documented. 3/3.
- Q5 relevance reasoning. Every supplier carries a score and the reason behind it, and the rejections carry theirs. 3/3.
- Q6 precision and recall. We have not run a formal precision and recall study, and we will not quote one we have not done. What we publish instead is whole campaigns with their rejections, which is weaker as a statistic and stronger as evidence. 2/3.
- Q7 productized ERP connectors. None. Not pilot, not roadmap, none. You export a file and load it yourself. By this article's own standard that is a walk-away for anyone who needs an integration. 0/3.
- Q8 export schema. Documented spreadsheet export, shippable today. 3/3.
- Q9 API access. There is no public API and no public docs for one. 0/3.
- Q10 transparent pricing. $69, $159 and $299 a month, on the pricing page, no call required. 3/3.
- Q11 pricing model. A monthly subscription with a campaign allowance. Not per seat, not per user. 3/3.
- Q12 three live references. We do not have three. We have published every campaign we have run instead, which is a different kind of evidence and not a substitute for a customer who will take your call. 1/3.
Total: 27 out of 36, with two outright failures. By the threshold this article sets that is a pass with material gaps, and the gaps are exactly where you would expect them in a small product: no integrations, no API, not enough customers yet.
Run this scorecard on every vendor you evaluate. Include us. That is the point, and the point is worth less if we mark our own homework generously.
472 suppliers across 8 briefs, 1053 rejections still on the record with the reason each one was given.