procurea.Book a call
AI Sourcing Automation

Buyer's Guide: 12 Questions to Ask Before Picking an AI Sourcing Tool

Every AI sourcing demo looks the same. These are the 12 questions that separate real capability from cosmetic polish, and the specific answers that should make you walk away.

4656
Pages read
472
Shortlisted
1053
Rejected with a reason
300
With an email

ALL 8 PUBLISHED CAMPAIGNS, SUMMED. THE FAILED ONES INCLUDED.

Why every AI sourcing demo looks the same

You have probably sat through three or four of them this year. A polished deck. A live screenshare. A pre-built brief that runs in 90 seconds. A shortlist of suppliers that look real, in a country you care about, with names and emails. "Want to see another category?" "Sure." Another 90 seconds, another impressive list.

The problem is that every AI sourcing vendor can build that demo. The underlying search + LLM pipeline is not proprietary technology; it is commodity components arranged with different levels of care. What separates real capability from a good demo is what happens on the 50th campaign, the week your integration partner changes an API, the Tuesday a supplier's VAT returns invalid and you need to understand why.

This guide is 12 questions structured so vendors cannot dodge them. Each question has a "good answer," a "yellow flag," and a "walk-away" answer. Use it as a scorecard: 3 points for good, 1 for yellow, 0 for walk-away. Max 36. A vendor scoring under 24 is a no. A vendor scoring 36 is either perfect or lying, reference-check hard.

Structure: four themes, three questions each. Data sources, AI + transparency, integration + workflow, commercials + customer proof.

12
Questions to ask
4
Evaluation themes
3-5
Typical vendor shortlist
60-90 d
Full evaluation cycle
12 questions across four themes, data, AI, workflow, proof.
One query, five of twenty six
🇨🇳CHINESE二甲双胍原料药 GMP 生产商
🇩🇪GERMANMetformin Wirkstoff Hersteller GMP
🇯🇵JAPANESEメトホルミン 原薬 GMP 製造
🇸🇪SWEDISHmetformin API tillverkare GMP
🇮🇹ITALIANmetformina API produttore GMP
FIG. 01 · THE SAME BRIEF, DISPATCHED IN ITS MARKETS' OWN LANGUAGES

Theme 1, Data sources (3 questions)

Q1. Which web sources do you actually crawl?

Good answer: A specific list. Google + regional directories (wlw.de, panorama-firm.pl, europages) + trade registries (KRS, Handelsregister, Companies House) + language-specific sources per country. Vendor names sources and explains how each is incorporated.

Yellow flag: "We have partnerships with data providers" without naming them. This usually means one paid commercial dataset (Dun & Bradstreet, ZoomInfo, or an aggregator) rebranded as "AI search."

Walk-away: "Our AI finds suppliers across the web" with no specificity. Translation: we run Google Custom Search Engine and prompt an LLM. That is commodity infrastructure and will not out-perform your own searches at scale.

Q2. Do you scrape directories yourself, or license their data?

Good answer: Combination. Scrape the public web (Google results + crawlable sites), license structured data where it exists (registry APIs, VIES for VAT), supplement with verified directory scraping where ToS permits.

Yellow flag: "We only use licensed data." Licensed data is clean but narrow, it misses the long tail of manufacturers who do not pay for directory inclusion.

Walk-away: "We scrape everything including premium directories behind login walls." That is a ToS violation waiting to become a lawsuit, and if the vendor talks openly about it, they are either naive or careless with compliance.

Q3. How do you handle non-English sources?

Good answer: Multilingual search queries generated per country, native-language content scraping, translation done client-side with the original text preserved in the record. The vendor can show you a German-language supplier homepage feeding into the shortlist with the description in both German and English.

Yellow flag: "We translate queries with DeepL and search Google." It works, but it is thin, you miss suppliers whose homepages only rank in their local language SERPs.

Walk-away: "Our primary focus is English-language sources." You will miss 60-70% of the real European supplier base. No AI polish fixes that.

A missing certificate never drops a good maker. It lowers a score, and the score is a number you can argue with.

Theme 2, AI + transparency (3 questions)

Q4. Which LLM do you use and is it swappable?

Good answer: Specific model name (GPT-4, Gemini 2.0 Flash, Claude 3.5) with a multi-provider fallback. Vendor explains why they use the specific model for specific tasks (e.g., Gemini for search strategy, GPT-4 for relevance screening). Willing to swap if you have a policy preference.

Yellow flag: "We use a proprietary model." If they trained a foundation model from scratch, they will happily tell you. If "proprietary" means a fine-tune of an open-weights model, fine, but they should say so.

Walk-away: "Our AI is proprietary, we cannot discuss the details." This is usually a cost-optimization story, they are using GPT-3.5 or a small open model to keep margins high and do not want to admit it.

Q5. Can I see the reasoning for each supplier's relevance score?

Good answer: Yes. Vendor shows a per-supplier explanation like "scored 0.87 relevance because: homepage mentions injection molding, company size 45 employees matches your 20-500 range, ISO 9001 listed." You see the signals.

Yellow flag: "You can see the score but not the full reasoning." This usually means the pipeline works but logging was an afterthought. Workable; not ideal.

Walk-away: "The scores come from our AI model, we cannot decompose them." Translation: there is no rational audit trail. When internal audit asks how you qualified this supplier, you will have nothing to show.

Q6. What is your precision and recall on a category like mine?

Good answer: Specific numbers with a published methodology. "On metal fabrication in Poland, precision ~88%, recall ~72%." Vendor shows how they measured (sampling, blind-coding by procurement specialists, etc.). Admits where they perform worse and why.

Yellow flag: Generic benchmarks "95% accurate" with no definition. Accuracy is not precision or recall; using accuracy for classification tasks is a pedagogical red flag.

Walk-away: "We do not measure precision and recall." Any vendor serious about an AI pipeline measures these. If they do not, they are not operating a quality process, they are running a demo pipeline and calling it production.

Theme 3, Integration + workflow (3 questions)

Q7. Which ERPs have productized (not "we can build it") connectors?

Good answer: Specific list with status per ERP. "NetSuite, productized via Merge.dev. SAP S/4HANA, pilot. Microsoft Dynamics: CSV only. Oracle Fusion, not supported." Honest about what is shipping, what is pilot, what is custom.

Yellow flag: "We integrate with all major ERPs." Every vendor says this. It usually means they have done a custom integration for one customer with each ERP and will do the same for you at an unnamed cost.

Walk-away: "ERP integration is part of our enterprise tier, details in the statement of work." Translation: they have no productized connectors; every integration is bespoke and will take 3-6 months at consulting rates.

Q8. What does your CSV export schema look like?

Good answer: Here is the schema. Specific field names. Shows you a real CSV from a real campaign. Fields include everything you need for an ERP import (company name, VAT, address, primary contact, category, certifications, trust score).

Yellow flag: "We support CSV export" without showing the schema. Usually means the export exists but is thin.

Walk-away: "Export is on our roadmap." If export is roadmap, the tool is a prototype. Walk.

Q9. Which plan includes API access?

Good answer: Specific plan names and what the API exposes. "API access on our Growth plan (€299/month), read access to campaigns and supplier records. Write access on Enterprise." Documentation link provided.

Yellow flag: "API access available on Enterprise plans." Usually fine, but ask for the API docs. If they cannot share docs before the contract, the API is either not stable or not real.

Walk-away: "API access is custom per customer." This means there is no API; every customer gets a snowflake endpoint that breaks when the vendor changes internals.

Theme 4, Commercials + customer proof (3 questions)

Q10. Is pricing transparent on your website?

Good answer: Yes. Specific plan names, specific prices, what each includes. Enterprise custom pricing is fine as long as Starter and Growth tiers are public.

Yellow flag: "Contact sales for pricing." Everywhere. For every plan. Usually means they price-discriminate heavily, and you should expect the first quote to be 2-3x the right number.

Walk-away: "Our pricing is based on your spend volume." This is almost always a % fee on sourced spend, which turns into a 5-8% tax on every category you put through the tool. For a €20M procurement spend, that is €1M/year, which is Ariba money.

Q11. Credit-based or seat-based pricing?

Good answer: Clear model. Credit-based (pay per campaign, buy in packs) is friendlier for variable-usage teams. Seat-based (flat per-user per-month) is friendlier for heavy daily users. Either is fine if the vendor explains why they chose it and the economics match your usage.

Yellow flag: Hybrid models with 4 different meters. Usually a sign the pricing was designed by a CFO optimizing revenue capture, not a product team optimizing customer value.

Walk-away: "It is custom per customer." You will be the price-discrimination test case.

Q12. Show me three live customers in my size range.

Good answer: "Here are three customers, similar industry, similar size. They have agreed to 30-minute reference calls. Here are their names." The vendor schedules the calls and does not sit in on them.

Yellow flag: "We have customer logos we can share." Logos are marketing, not references. Always push for actual calls.

Walk-away: "We cannot share customer details due to NDA." Every customer signs an NDA; that does not prevent reference calls. If the vendor cannot produce three customers willing to talk, they probably do not have three customers the right size for you.

#QuestionThemeWhat a good answer looks like
Q1Which web sources do you actually crawl?DataNames specific sources: Google + directories + registries
Q2Scrape or license directory data?DataHybrid: public web + licensed registries (VIES, IAF)
Q3How do you handle non-English sources?DataMultilingual queries per country, native-language scraping
Q4Which LLM, and is it swappable?AISpecific model named (GPT-4, Gemini) with multi-provider fallback
Q5Can I see reasoning for each score?AIPer-supplier score decomposition visible in UI
Q6Precision and recall on a category like mine?AIPublished numbers with methodology (e.g. 88% / 72%)
Q7Which ERPs have productized connectors?IntegrationPer-ERP status (productized / pilot / CSV / unsupported)
Q8What does your CSV export schema look like?IntegrationShows a real CSV; schema documented
Q9Which plan includes API access?IntegrationSpecific plan + public docs link before contract
Q10Is pricing transparent on your website?CommercialPublic Starter + Growth tiers; Enterprise custom OK
Q11Credit-based or seat-based pricing?CommercialClear model; rationale matches your usage pattern
Q12Show me three live customer references.CommercialSchedules 30-min calls; vendor does not sit in
12 questions scorecard, what separates real capability from demo polish

The honesty test (ask every vendor this)

Beyond the 12 questions, one meta-question reveals more about the vendor than anything else: "What is the number-one thing your product is NOT good at?"

A mature vendor will have an answer. "We are weak on categories under 10 global suppliers, our AI assumes a larger universe than those categories have." "Our Salesforce integration is pilot, not productized, so if you need real-time bidirectional sync on day one, we are not there yet." "We do not support Oracle Fusion at all, if that is your ERP, do not buy us."

Those answers are selling. They demonstrate that the vendor has self-awareness, measured their product, and respects your decision-making. A vendor who cannot name a weakness is either lying about having no weaknesses, or does not understand their product well enough to have found them. Both are disqualifying.

The other meta-question worth asking: "What would cause you to recommend we not buy your product?" The best vendors have this answer ready. "If your procurement is 80% services, if you have 40+ suppliers per year across complex categories but your budget is under €5k/year, if you need deep S/4HANA sync on week one, in all three cases, we are not the right fit."

A vendor willing to talk you out of a bad-fit purchase is more likely to still care about you in year two. A vendor who pretends every customer is a good fit will lose interest the moment the contract signs.

The reference-check protocol that actually works

A logo slide is not a reference. A 30-minute call with an actual customer is. Three references, 30 minutes each, specific questions.

Ask each customer:

  • How long have you been using the tool? Less than 6 months = too new to matter. 6-18 months = peak honeymoon. 18+ months = real signal.
  • What did you try before? Reveals what problem the tool actually solved versus what the vendor claims it solves.
  • What is the #1 thing you wish the tool did better? If the answer is "nothing" or a feature request that sounds like marketing fluff, the reference was coached. If the answer is specific and technical ("the email follow-up cadence is too aggressive for our German suppliers"), the reference is real.
  • Has the vendor ever failed you? How did they handle it? Vendors handle failures either well (clear root cause, timely fix) or badly (deflection, blame). Customers remember.
  • Would you buy it again today? Ask this at the end. The answer and the hesitation both matter.

Three calls, 90 minutes total, will tell you more than five hours of vendor demos. Do not skip this step; the vendors who are confident in their product will facilitate it, and the vendors who dodge it are telling you what you need to know.

  1. Score each vendor 1-5 on all 12 questions

    3 = good answer, 1 = yellow flag, 0 = walk-away. Max 36. Fill the scorecard during the demo, not after, memory degrades fast.

  2. Weight by your top 3 priorities

    Pick the 3 questions most critical for your use case (e.g. Q6 precision, Q7 ERP, Q12 references). Double-weight them in the total.

  3. Calculate weighted total per vendor

    Total + weighted bonus. Anyone under 24 (unweighted equivalent) is a no. Anyone over 34 is suspicious, reference-check hard.

  4. Eliminate bottom 2, demo final 2

    If you have 4-5 vendors, drop the bottom 2. Book 60-min deep demos with final 2 using your own real brief, not theirs.

  5. Reference calls + honesty test

    3 references × 30 min per finalist. Ask meta-question: "what is your product NOT good at?" Mature vendors answer honestly.

Procurea answers all 12, and fails two of them

Applying the scorecard to ourselves. This section used to award us 34 out of 36. That was written when we intended to build things we have since decided not to build, and it stayed on the site long after it stopped being true. Here it is against what exists today.

  • Q1 data sources. The live web through Serper, company pages read directly, regional directories, and VIES for VAT numbers. Named, not "proprietary". 3/3.
  • Q2 scrape vs license. We read the public web and respect robots.txt; VIES is the European Commission's own service. 3/3.
  • Q3 non-English. Queries and reading in local languages, which is the whole point. The published runs turn up German, Italian, Turkish and Latvian makers a monolingual search does not reach. 3/3.
  • Q4 LLM swappable. Vertex primary with an AI Studio fallback, documented. 3/3.
  • Q5 relevance reasoning. Every supplier carries a score and the reason behind it, and the rejections carry theirs. 3/3.
  • Q6 precision and recall. We have not run a formal precision and recall study, and we will not quote one we have not done. What we publish instead is whole campaigns with their rejections, which is weaker as a statistic and stronger as evidence. 2/3.
  • Q7 productized ERP connectors. None. Not pilot, not roadmap, none. You export a file and load it yourself. By this article's own standard that is a walk-away for anyone who needs an integration. 0/3.
  • Q8 export schema. Documented spreadsheet export, shippable today. 3/3.
  • Q9 API access. There is no public API and no public docs for one. 0/3.
  • Q10 transparent pricing. $69, $159 and $299 a month, on the pricing page, no call required. 3/3.
  • Q11 pricing model. A monthly subscription with a campaign allowance. Not per seat, not per user. 3/3.
  • Q12 three live references. We do not have three. We have published every campaign we have run instead, which is a different kind of evidence and not a substitute for a customer who will take your call. 1/3.

Total: 27 out of 36, with two outright failures. By the threshold this article sets that is a pass with material gaps, and the gaps are exactly where you would expect them in a small product: no integrations, no API, not enough customers yet.

Run this scorecard on every vendor you evaluate. Include us. That is the point, and the point is worth less if we mark our own homework generously.

Read the campaigns behind these numbers
All 8 campaigns, published unedited.

472 suppliers across 8 briefs, 1053 rejections still on the record with the reason each one was given.

Questions buyers ask about this

What score on the 12-question scorecard means a vendor is a pass?
Minimum 24 out of 36. Below 24 means more than a third of answers were walk-aways or yellow flags, the product has material gaps. A perfect 36 is actually suspicious, it means either a very mature vendor (possible) or one who said yes to everything (more common). Calibrate by asking the honesty meta-question: "what is the #1 thing your product is not good at?" A vendor scoring 36 with a confident weakness answer is real; one with no weakness answer is dodging.
Should I disqualify a vendor that uses GPT-4 or Gemini under the hood?
No. Most AI sourcing vendors use commodity foundation models, that is fine and expected. What matters is (a) they are transparent about which model, (b) they can swap if you have a policy preference, (c) they add value on top (multilingual query strategy, verification layers, ERP integrations, audit trails). The model is the commodity; the pipeline around it is the product. A vendor hiding which model they use is the red flag, not the model choice itself.
How many reference calls should I do before buying?
Three, at minimum. Each 30 minutes. Different industries if possible, definitely different company sizes within your range. Ask the vendor to schedule them and not attend. If the vendor cannot produce three, that is signal. If the vendor wants to sit on the calls, that is also signal. The 90 minutes total will tell you more than five hours of vendor demos.
Is a pilot status disqualifying for ERP integration?
Not automatically. It depends on your tolerance for being an early customer and your timeline. Pilot means the integration works but is still stabilizing field coverage and edge cases. If you can start with a CSV export path and move to the pilot integration in month 3-6, pilot is fine. If you need bulletproof real-time sync on day one for a compliance-critical workflow, pilot is not for you, pay for a vendor with productized connectors even at a higher cost.
What is the single best question to ask an AI sourcing vendor?
"What is the number-one thing your product is NOT good at?" A mature vendor has an answer, specific, honest, not a marketing-adjacent weakness like "we are growing so fast we cannot keep up with demand." A vendor who cannot name a real weakness either does not know their product or is willing to hide things from you. Both are disqualifying. This one question reveals more than any individual capability question.