Fourteen vendors, thirty-one criteria, weights carried to two decimal places. The final scores land between 4.1 and 4.4. Somebody asks what that spread means and the room goes quiet, because it does not mean anything — the matrix was built after the shortlist, and the shortlist was built in a hallway.
That is the normal outcome, not a lapse in diligence. A scoring matrix is good at producing a number and bad at producing a decision.
Why every vendor scores four out of five
Most criteria describe table stakes. Has an API. Supports single sign-on. Offers onboarding. Every serious vendor clears those, so scoring them manufactures a tie. The criteria that would actually separate the options are specific and arguable, which is why they stay hallway conversation instead of becoming rows.
And weights are a steering wheel. If changing one weight changes the winner, the weights are the decision and the scores are decoration. Nothing wrong with a judgement call; something wrong with one wearing arithmetic.
Write the disqualifiers first
A disqualifier is a binary fact that ends the conversation. Not "we would prefer X" but "if it cannot do X, we are not buying it." Must run in the region our data has to stay in. Must be operable by the two people who will actually operate it.
Three rules make them work.
- Each must be checkable. Someone outside the evaluation answers yes or no with evidence. "Good support" is not a disqualifier; "a named contact who replies in writing within one business day" is.
- Keep the list short. Five is plenty. Twenty means you wrote preferences and called them requirements.
- Write them before the first demo. Afterwards you will soften the one that eliminates the vendor whose product felt nicest, and you will not notice.
Disqualifiers beat scores because they are cheap to check and impossible to average away. And if they eliminate everyone, that is a finding rather than a failure.
Ask a small number of questions that differentiate
Once there is a shortlist, put five to eight questions to every vendor in the same words and the same order.
One test decides what belongs. Before sending, predict each vendor's answer. If you can predict all of them, the question does not discriminate — cut it. What survives is usually questions about failure: what happens when this breaks, what does it not protect us from, enumerate what it does without a human involved.
Seven questions for AI communication tools is a worked example — each question paired with what a good answer looks like and what an evasion looks like. Ask in writing; a confident demo answer is unquotable six weeks later.
Separate what you verified from what they claimed
This distinction is the whole exercise. Three buckets, not one score. Verified — you ran it, or a reference customer confirmed it under conditions resembling yours. Claimed — they said so, in writing, and you have not seen it. Unknown — nobody asked, or the answer was evasive.
Most write-ups collapse all three into one confident sentence. Then what you tested and what you were told carry identical weight, and six months later nobody can reconstruct which was which.
Marking something as claimed is not an accusation. It describes what you know, and it turns the unknown column into a short, cheap to-do list.
Decide who decides, before you start
Write down at the outset who chooses, who holds a veto and over exactly what, who is consulted, and who only wants telling afterwards. One name in the decider slot.
Evaluations that drag are rarely undecided. They are unowned — six people with opinions and nobody holding the pen, which reliably produces another round of questions instead of a choice. Asking for a decision has a general shape; here the load-bearing half is a single name and a real date. Agree the tiebreak in advance too.
The briefing for whoever signs
The person who signs was not on the calls and will spend a few minutes on this. Their questions are predictable — what did we pick, what did we rule out and why, and what do you need from me.
Recommendation and ask first. Then one line per rejected option with the reason it lost, the reason and not the score. Then the assumptions it rests on, marked as assumptions. Then the questions still open. Then your confidence, and whether the decision is reversible.
A reversible choice held with medium confidence is worth signing in a minute; an irreversible one at the same confidence deserves an argument. An average of 4.3 hides which you are looking at. A decision record is where that pair outlives its author.
Where the format helps, and where it stops
RelayLink carries this kind of briefing between two people's AI assistants, and its fields match the shape above. Options considered are structured — each with a status of chosen, leaning, open or rejected and a short why — so the rejected list travels with the recommendation. Decisions carry a confidence of low, medium or high and a reversible flag. Assumptions arrive tagged as either stated by the sender or inferred by the sender's AI, so a guess your assistant made while summarising does not reach the signer looking like something you checked. How to state an assumption is the writing side of that.
The honest limit is that nothing in the format marks a vendor claim as verified. Provenance there describes you and your assistant, not the vendor. The verified-versus-claimed split is prose you write; the structure supplies a place to put it and a context brief capped at 3,500 characters, so the write-up cannot quietly become the chain.
None of this stops you choosing badly. A disqualifier strict enough to be useful will occasionally eliminate a vendor who would have been fine. What changes is the aftermath — the disqualifiers, the rejected options and the confidence at the time locate the faulty assumption in an afternoon.
If the calls are done and the write-up is due, connect your assistant and have it assemble the options and the reasons they lost.