Akonita Resources

AI Vendor Selection: How to Evaluate Without Getting Burned

Cover image for AI Vendor Selection: How to Evaluate Without Getting Burned

AI Vendor Selection: How to Evaluate Without Getting Burned

TL;DR: Every AI vendor sounds impressive in a demo. That is the demo's job. The difference between a vendor that delivers and a contract you regret is almost never visible in the pitch — it shows up six months later, when the system meets your messy data, your real usage volumes, and your first 2 a.m. incident. Traditional software procurement was not built for this: you cannot evaluate a probabilistic system with a feature checklist. This article lays out a practical evaluation framework — why AI vendor selection is a different discipline from software procurement, the seven criteria that actually predict success, the red flags hiding inside polished pitches, how to calculate the real total cost of ownership beyond the subscription line, when vendor lock-in is an acceptable trade and when it is fatal, and the reference checks that surface what case studies are designed to hide.

AI vendor selection — evaluating AI vendors without getting burned

Introduction

There is a meeting that happens in almost every company evaluating AI vendors. Three vendors have pitched. All three demos were impressive. The chatbot answered every question. The document extraction worked flawlessly. The analytics agent produced a beautiful summary of a dataset it had clearly seen before. The evaluation team leaves the room with three scorecards that all say roughly the same thing, and the decision comes down to whoever had the most confident salesperson or the lowest headline price.

Eighteen months later, half of those companies are quietly looking for an exit.

The failure is rarely that the vendor lied. It is that the evaluation was designed for the wrong product category. Traditional software procurement rests on assumptions that hold for deterministic systems: a feature either works or it does not, a demo is a reasonable proxy for production behaviour, and an uptime SLA covers the main reliability risk. AI systems violate all three. An AI feature works at some level of accuracy, on some distribution of inputs, at some cost per call — and all three of those numbers shift depending on your data, not the vendor's demo environment. A demo tells you what the system does on curated inputs. Production tells you what it does on yours.

This article is the evaluation framework we use when clients ask us to help them choose between AI vendors — or to decide whether a vendor is the right answer at all. It is not a list of questions to copy into an RFP. It is a way of thinking about what actually predicts whether a vendor relationship will work, built from watching these decisions go right and wrong across dozens of evaluations.

Why AI vendor evaluation is different from software procurement

Before the framework, it is worth being precise about why the old playbook fails. Four properties of AI systems break traditional evaluation methods.

Performance is probabilistic, not binary. A traditional software feature passes or fails. The export button either produces a valid file or it does not. An AI feature produces outputs that are right some percentage of the time, on some kinds of inputs, with some distribution of failure modes. A vendor claiming "95% accuracy" is making a statement about a specific test set — usually one they chose. The number that matters is accuracy on your data, your edge cases, your users' phrasing. That number does not exist until you measure it.

The product is a moving target. Traditional software changes at the pace of releases. AI products change at the pace of model upgrades, which the vendor may not even fully control. The system you evaluate in March may run on a different underlying model by September — with different behaviour, different failure modes, and different costs. You are not evaluating a product. You are evaluating a trajectory, and the vendor's discipline around model transitions matters as much as the current feature set.

The demo-production gap is structural. Every vendor demo runs on clean data, rehearsed prompts, and examples chosen because they work. This is not dishonesty — it is what demos are. But for AI systems the gap between curated and real inputs is not a rounding error. It is often the difference between a product that works and one that does not. The evaluation has to close that gap deliberately, because the pitch process never will.

Value depends on your integration, not just their product. The same AI vendor produces spectacular results at one company and mediocrity at another, because the outcome depends on the customer's data quality, integration depth, and internal adoption as much as on the software. This means reference checks matter more than feature lists — and it means part of what you are evaluating is your own readiness, not just the vendor.

The seven evaluation criteria that actually matter

Most AI vendor scorecards measure what is easy to measure: feature counts, pricing tiers, the polish of the sales deck. The criteria below are harder to evaluate and far more predictive. They are ordered roughly by how often each one is the actual cause of a failed vendor relationship.

1. Performance on your data, not theirs

This is the single most important criterion, and the one most evaluations skip. The only performance number that matters is how the system behaves on your data, your documents, your customers' phrasing, your edge cases. Everything else — benchmark scores, demo results, case study metrics — is marketing material with a methodology section.

The practical test: require every shortlisted vendor to run a structured evaluation on a dataset you provide. Pick 50 to 100 representative examples from your real operations, including the ugly ones — the incomplete records, the ambiguous requests, the inputs that break your current process. Define the scoring rubric before the vendor sees the data. If a vendor will not do this, the evaluation is over. They have told you the demo is the product.

2. Production track record at your scale

There is a difference between a vendor with customers and a vendor with customers like you, in production, at your volume, for more than a year. A startup that has run three pilots is not the same risk profile as one running fifty production deployments — even if the pilot demos were better. Ask specifically: how many customers are live in production, in my industry, at my transaction volume, for twelve months or longer? Then ask to talk to two of them. Logos on a slide are not evidence. A customer who has survived a renewal is.

3. Data security and compliance posture

For AI vendors, the security question has an extra layer: your data does not just sit in their database — it flows through their models, their prompts, their logs, and often their third-party model providers. The questions that matter: Where does our data go when the system processes it? Is it used to train any model, theirs or a provider's? What is the retention policy, and can we enforce deletion? Which subprocessors touch the data, and in which jurisdictions? What happens to our data when the contract ends? A serious vendor answers these in writing, with contractual backing — not in a sales call. If your data is regulated, this criterion is pass-fail and it comes first, not third.

4. Pricing transparency and unit economics at scale

AI pricing models hide cost in ways traditional SaaS does not. A per-seat price is predictable. A per-token, per-call, or per-resolution price scales with your success — which means the vendor's incentive is to grow your bill, and your cost at 10x usage is a guess unless you model it. The test: take your projected usage at launch, at 12 months, and at 24 months, and ask the vendor to price all three scenarios in writing. If the year-two number is not survivable, the year-one discount does not matter.

5. Exit terms and portability

You will eventually leave every vendor, or they will leave you — acquisition, shutdown, pivot, or a price change you cannot accept. What you take with you determines how painful that day is. For AI systems the portable assets are: your data (in standard formats), your prompts and evaluation sets, any fine-tuned model weights, and your conversation histories. Ask at evaluation time what happens to each of these on termination. A vendor confident in their product makes leaving easy. A vendor whose retention strategy is your switching cost will be vague here, and that vagueness is information.

6. Integration depth and API quality

The AI product you buy is only as good as the surface it exposes to your systems. Weak APIs, sparse documentation, missing webhooks, no SSO, no audit logs — these do not show up in demos, and they dominate your actual cost of ownership. Have an engineer, not a procurement lead, read the API documentation before the shortlist is final. If the docs are thin, the integration will be expensive. If there is no sandbox environment, your team will be testing in production.

7. The team and the support model

When the system misbehaves at 2 a.m. before a board demo, the question is not what the SLA says — it is who answers. Understand the support model in concrete terms: named solutions engineer or a ticket queue? Response times by severity, in the contract or on a web page? Do customers get roadmap input, and does it ever ship? Ask references specifically about support, because it is the one criterion you cannot test in a pilot — the vendor is never more attentive than during the sale.

AI vendor evaluation criteria — the scorecard that predicts success

Red flags in vendor pitches

Some warning signs are obvious. Most are not — they are engineered to look like strengths. These are the patterns that correlate most strongly with vendor relationships that end badly.

The demo never touches your data. Every example is the vendor's dataset, the vendor's documents, the vendor's scripted flow. When you ask to see it on your material, the answer involves scheduling, preparation, or a follow-up call that never quite happens. A system that only works on curated inputs is not a product. It is a performance.

Benchmark scores stand in for evidence. "We scored 94% on the industry benchmark" is a claim about a test the vendor selected, on data they did not show you, with a methodology you cannot inspect. Benchmarks are useful for researchers. For buyers, the only benchmark that matters is the one you run yourself, on your own evaluation set.

Everything is "AI-powered" and nothing is explained. Ask how the system actually works — what models, what retrieval, what guardrails, what happens when the model is wrong — and the answer stays at the level of adjectives. Vendors with real engineering can explain their architecture to an executive in plain language. Vendors hiding a thin wrapper around someone else's API cannot, because the explanation would reveal how little you are paying for.

No referenceable customer at your scale. The case studies are impressive but anonymous, or from companies a tenth your size, or from industries with nothing in common with yours. If nobody like you has bet on this vendor and survived a year of production, you would be the experiment.

Resistance to a scoped pilot. A vendor confident in their product welcomes a time-boxed pilot on your data with agreed success criteria — it is their fastest path to a signed contract. A vendor who pushes hard to skip straight to the annual agreement is telling you the product does not survive scrutiny.

Pricing that is vague in every dimension. "It depends" is a fair answer to one or two pricing questions. When it is the answer to all of them — usage, overages, support tiers, price increases at renewal — you are not being quoted a price. You are being sized for one.

The roadmap is sold as the present. "That capability is coming next quarter" is fine, as long as you are not paying for it this quarter. Evaluate and contract against what exists today. Roadmap features are options, not assets — and vendor roadmaps slip more often than they ship.

Total cost of ownership: the subscription is the tip

The subscription fee is the most visible cost of an AI vendor and usually the smallest third of the real number. Year-one total cost of ownership typically splits into seven categories, and most evaluations model only the first.

Cost categoryWhat it coversTypical year-one share
Subscription / usage feesThe contract price, including usage overages30–40%
Integration engineeringConnecting the vendor to your systems, data pipelines, auth, and workflows20–30%
Data preparationCleaning, structuring, and migrating the data the AI needs to work10–15%
Evaluation and tuningBuilding your test sets, tuning prompts, measuring accuracy on your data5–10%
Internal operationsThe people who monitor outputs, handle failures, and manage the vendor10–15%
Training and adoptionGetting your team to actually use the system, and fixing the process around it5–10%
Exit and migration reserveThe eventual cost of leaving — data export, re-integration, parallel runningunbudgeted, always real

Three implications follow from this. First, a vendor with a higher subscription and better APIs can be cheaper in total than a cheap vendor with a thin integration surface — the engineering line dwarfs the price difference. Second, usage-based pricing means your success is the vendor's growth lever, so model the 24-month cost at realistic adoption, not the launch month. Third, the categories the vendor cannot see — your data prep, your internal ops, your adoption work — are the ones most likely to kill the project's ROI. Budget for them before you sign, not after the first invoice surprise.

Total cost of AI vendor ownership — the subscription is only the tip of the iceberg

Vendor lock-in: when it is acceptable and when it is fatal

Lock-in is not binary. It is a spectrum, and some positions on it are reasonable trades. The mistake is not accepting lock-in — it is accepting it accidentally, without pricing what you are giving up.

Acceptable lock-in is commodity capability behind standard interfaces. If the vendor provides something you could rebuild or re-source in a quarter — a hosted vector database behind an abstraction layer, a transcription API, a commodity classification service — then lock-in costs you, at worst, a migration project. Annoying, budgetable, survivable. You trade some future flexibility for present speed, and the trade is often worth it.

Fatal lock-in is when the vendor ends up owning the assets that make the system valuable to you. Your proprietary data schemas locked in a proprietary format. Fine-tuned model weights you cannot export. Prompts, evaluation sets, and agent workflows encoded in the vendor's orchestration layer with no way out. Conversation histories and behavioural data — the record of how your customers actually talk — trapped behind an export function that produces a CSV nobody can use. In this position, the vendor does not need to win your renewal on merit. They just need to make leaving expensive enough that you stay by default.

The test is a single question, asked in the evaluation and answered in the contract: if this vendor doubled its price or was acquired tomorrow, how long would migration take, and what would we permanently lose? If the honest answer is "months, and our fine-tunes, our evals, and our history," the subscription price is the least important term in the deal.

Two negotiating points matter here, and both are cheapest at signing — the moment your negotiating power is highest, before you are dependent. First, contractual data portability: full export of your data, prompts, configurations, and any fine-tuned artefacts in documented, standard formats, at no additional charge, within a defined period after termination. Second, a transition assistance clause: the vendor commits to reasonable support for migration for a fixed window after the contract ends. Vendors who plan to keep you by being good will sign both. Vendors who plan to keep you by being sticky will resist — and now you know which one you are dealing with.

Vendor lock-in vs portability — knowing which door you are walking through

Reference checks that reveal the truth

Reference checks are the highest-signal step in any AI vendor evaluation, and the step most teams perform worst. The vendor hands you two curated contacts, you ask them if they are happy, they say yes, and the box is ticked. This is theatre. Done properly, reference checks surface everything the demo was designed to hide.

Get the right references, not the offered ones. Ask for customers matching your size, industry, and use case — and push past the first list. Then do the harder work: find customers who left. The vendor will never introduce you to a churned customer, but your network, LinkedIn, and industry communities will. Thirty minutes with a company that evaluated the same vendor and walked away — or signed and later exited — is worth more than five calls with references from the vendor's highlight reel.

Ask questions with answers that cannot be rehearsed. "Are you happy?" invites a polite yes. These do not:

  • What broke in the first 90 days of production, and how did the vendor respond?
  • What did the second year actually cost compared to what you were quoted?
  • What did the vendor promise during the sale that has not shipped?
  • How does the system perform on your worst data, not your average data?
  • When you had your worst incident, who answered, and how fast?
  • What did you negotiate into your renewal that you wish you had negotiated at signing?
  • If you were evaluating today, knowing what you know now, what would you ask?

The pattern to listen for is specificity. Happy reference customers talk in specifics: the incident, the response time, the number. Rehearsed references talk in adjectives. And the single most predictive answer in any reference call is the response to "what did year two cost?" — because it compresses pricing honesty, usage growth, and the vendor's renewal behaviour into one number.

Check the vendor's stability as a company. An AI vendor is not just a product — it is a bet on that company's next three years. Funding position, customer concentration, leadership turnover, and acquisition rumours all matter, because the failure mode is not just the product getting worse. It is the vendor getting acquired by a company that sunsets the product, or running out of runway mid-contract. This is not a reason to only buy from giants — it is a reason to make portability and exit terms non-negotiable when buying from anyone who is not one.

Build vs buy: the evaluation lens most teams skip

There is a framing error baked into most vendor evaluations: they compare vendors against each other, when the real question is vendor against alternative. Every AI vendor selection is secretly a three-way decision — buy, build, or partner — and evaluating only the buy column produces the right answer to the wrong question.

Buy when the capability is commodity and speed matters. If the function is not core to your differentiation — meeting transcription, document OCR, standard support deflection — a vendor is usually right. The capability is well-understood, multiple vendors compete on price and reliability, and your engineering time is worth more than the subscription. The evaluation criteria above still apply, but the stakes of a wrong call are lower because the exit cost is lower.

Build when the capability is your differentiation and your data is the moat. If the AI system encodes something that makes your business better than competitors — your underwriting judgement, your logistics optimisation, your clinical triage logic — then a vendor relationship puts your most valuable asset inside someone else's product roadmap. Worse, your usage data trains the vendor's understanding of the problem, which they then sell to your competitors. When the capability is the business, own it.

Partner when you need custom but lack the team. The most common real-world position: the capability matters enough that an off-the-shelf product fits badly, but hiring a full AI engineering team to build it is a twelve-month detour. This is what engineering partners exist for — a team that builds the system on your data, inside your infrastructure, and hands you the keys. You get the ownership economics of build with the speed of buy, at the cost of picking the partner with the same rigour you would apply to a vendor.

The test that separates the three: if a competitor bought the same vendor subscription tomorrow, would they have what you have? If yes, the vendor is selling you table stakes — fine, as long as you are not paying differentiation prices for it. If no — if what makes the system valuable is your data, your workflow, and your tuning — then ask why you are renting it instead of owning it.

FAQs: AI Vendor Selection

How long should an AI vendor evaluation take?

Four to eight weeks for a serious evaluation: two to three weeks for desk research, demos, and reference checks, then two to four weeks for a scoped pilot on your data with pre-agreed success criteria. Anything under two weeks means you are buying the demo. Anything over six months usually means the organisation is avoiding a decision — and the market will have moved by the time you make one. The pilot phase is non-compressible; the meeting count around it almost always is.

Should the pilot be paid?

Yes — scoped, time-boxed, and paid. Free pilots get neither side's best effort: the vendor assigns junior staff, your team treats it as background noise, and the results mean nothing. A paid pilot with defined success criteria, a fixed dataset, and a decision meeting scheduled before it starts produces a real signal. The pilot fee is the cheapest insurance in the entire procurement process — it buys you the answer to the only question that matters before you commit to the annual contract.

How many vendors should we evaluate?

Three to five on the longlist, two to three through full evaluation, one or two into pilot. More than five on the longlist means the requirements are not clear enough yet — go back and define them. And resist the enterprise instinct to include a vendor you have no intention of choosing just to satisfy a process; every additional serious evaluation costs your team weeks of attention, which is the scarcest resource in the whole exercise.

What if a vendor will not let us test on our own data?

Walk away. This is the single most decisive red flag in AI procurement. There are legitimate edge cases — a vendor may reasonably ask for an NDA first, or scope the pilot to a data subset for privacy reasons. But a vendor whose answer to "run it on our data" is "trust the demo" is telling you that the gap between demo and production is one they cannot afford to show you. Believe them, and move on.

The vendor is just wrapping OpenAI or Anthropic — how do we evaluate that?

Evaluate the layer they add, because the model is not theirs. Many AI vendors are thin wrappers: a prompt template and a UI on top of a foundation model API you could call yourself. Others add real value on top — retrieval pipelines tuned to your domain, evaluation frameworks, guardrails, workflow integration, compliance infrastructure, and support that has seen your failure modes before. The test: ask your engineering team how long it would take to replicate the vendor's product by calling the same underlying API directly. If the honest answer is "a few weeks," you are paying enterprise prices for a prompt. If the answer is "six months plus everything we would get wrong," the layer is real, and the vendor earns their margin.

Conclusion

AI vendor selection fails in a predictable way: the evaluation measures what is easy to show, and the contract commits to what was never tested. The demo was impressive because demos are built to impress. The benchmark was strong because the vendor chose it. The price looked reasonable because nobody modelled year two. The exit terms looked standard because nobody read them until the exit.

The fix is not cynicism — good AI vendors exist, and the right one can compress years of capability building into months. The fix is an evaluation designed for what AI systems actually are: probabilistic, data-dependent, and only as good as their behaviour on your inputs. Test on your data or do not buy. Talk to customers who left, not just the ones on the slide. Price the full cost of ownership, including the work your own team will do. Negotiate portability and exit terms at signing, when your negotiating position is strongest. And before comparing vendors against each other, make sure buying is even the right answer — sometimes the capability is valuable enough to own.

The vendor question is never "which demo impressed us most." It is "which of these options survives contact with our data, our scale, our lawyers, and our year-two budget." Evaluate for that, and the burn rate on bad vendor decisions drops to nearly zero.

If you are evaluating AI vendors right now — or stuck between a build decision and a buy decision — we can help. We run vendor evaluations the way we run engineering projects: structured pilots on your real data, reference checks that go past the highlight reel, total-cost modelling that includes the lines the vendor cannot see, and an honest recommendation even when the recommendation is that you do not need a vendor at all.

Get help evaluating AI vendors — we will pressure-test the shortlist, run the pilot on your data, and give you a decision you can defend in the boardroom.

Related reading

A

Aria

Akonita AI · Online

Hi, I'm Aria — Akonita's AI assistant. I can answer questions about our services or help you figure out the best next step. What brings you here today?

Powered by Akonita AI