The secret of good evaluation scoring matrices

Dollars And Sense

We have all heard the phrase “You can’t have your cake and eat it, too!” If you stop to think about it, you CAN have your cake and eat it, too. What you CAN’T do is eat your cake and have it, too. 

Read on as Paul Rogers stops to think about one of the shibboleths of tender evaluation; the scoring guide that ‘helps’ evaluators decide whether to score a proposal six out of ten or seven out of ten.

Problem one: Vague descriptors

Procurement is a profession without standards. As an example, if you Google ‘tender evaluation scoring matrix’ and select ‘images’ you will see thousands of results, with the only commonality being they all include some numbers and some words. 

Here’s how Dr. Google defines a tender evaluation scoring matrix:

“A tender evaluation scoring matrix is a tool used to objectively compare bids by assigning scores to different criteria, such as price, technical capability and experience, which are weighted by their importance.”

Objectively, eh?

Let’s analyse one part of one of the matrices that I chose at random:

I have highlighted in bold the only words that differ between these three scores. The question I would like you to ask yourself is this: Do these words help the evaluation team score the responses objectively? 

We all know the answer to that question! 

The distinction between ‘high’, ‘very high’ and ‘superior’ is subjective. And don’t get me started on the word ‘superior’. Isn’t that a relative descriptor, scoring this bidder’s response against another bidder’s response? A cardinal sin!

The research evidence is that “…when rubrics contain unclear or subjective language, (evaluators) are likely to base their evaluations on an overall impression of the (document) rather than the specified criteria.”

This means that vague language not only introduces subjectivity (because there is no clarity as to what is the difference between ‘high’ and ‘very high’), but it causes evaluators to abandon criterion-based judgements entirely and revert to subjective impression-based scoring. Yikes!

The use of quantification helps reduce variation in scoring. So, instead of ‘some’, ‘partially’ or ‘mostly’, consider specifics: “Meets at least 80 percent of the requirement.” 

A good scoring matrix will reduce the spread of scoring. If you are interested in this topic, one metric is a percentage of scores that differ by no more than one point from other evaluators’ scores.

Problem two: Score compression

Commercial businesses that offer unsatisfactory performance go out of business pretty quickly. Anyone who has served on an evaluation panel will recognise the distribution of scoring in the chart below.

Most scores cluster in the range of 6 to 8, with fewer scores outside this range. Occasionally, there may be a lower score, while scores of 10 are infrequent. Let’s call this ‘score compression’. 

The problem presented by score compression is that when we total all of the scores for each criterion, we can end up with totals of (for example) 72, 73 and 74. There is no “clear blue water” between the respondents, and then the evaluation panel needs to search for other methods to distinguish between the respondents. 

The question we have to ask is this: Is the scoring matrix fit for purpose?

The purpose is to help us distinguish between the respondents, and one cause of score compression is the location of the ‘tipping point’. Let’s have a look at an example.

The first thing to note is that the threshold between satisfactory and unsatisfactory responses (the ‘tipping point’) is 5. I understand that this creates symmetry within the scoring matrix, but the problem is that locating the tipping point at 5 compresses the satisfactory scores in the range of 5 to 10. 

It is my experience that most of the responses will score more than 5, so the ‘useful range’ is from 6 to 10.

If you were wondering what is the difference between a ‘minor deficiency’, a ‘significant deficiency’ and a ‘critical deficiency’, let me refocus you on the tipping point. 

An unsatisfactory response doesn’t need degrees of ‘unsatisfactory-ness’. The evaluators need to have evidence of why the response is unsatisfactory, but I’m not clear why we need to gradate ‘unsatisfactory-ness’. The response is either satisfactory or unsatisfactory.

Why not set the tipping point at 1? A score of 0 is unsatisfactory, while 1 and above is satisfactory. If we set the threshold between unsatisfactory and satisfactory at a score of 1, this would give us a ‘useful range’ of 10, which is twice that of the conventional wisdom.

Problem three: The Rolls-Royce problem

Why pay for a Rolls-Royce when a Mini would do?

We’ve all heard this gem. It explores the question of comparing solutions against the need, rather than comparing solutions against each other. 

If offered a Mini or a Rolls-Royce for free, most people would choose the Rolls-Royce, but if we had to pay for the vehicle, our script may change. 

“Why would I pay so much more for the Rolls-Royce when the Mini meets my needs?” 

Hardly groundbreaking stuff but work with me on this.

Do we really want to give a higher score to a response that exceeds our requirements than a response that meets our requirements? Yes, this is a philosophical point, but imagine that the one point that “Rolls-Royce” solution won for exceeding our requirement was the deciding factor in awarding the contract. 

My suspicion is that the challenge of crafting intermediate descriptors has caused the person who designed the tool to latch onto ‘exceeding the requirement’ like a drowning person grabbing a passing buoyancy aid. It isn’t easy to come up with meaningful descriptors!

But what if the extra functionality is a ‘nice-to-have’? What if the Rolls-Royce is at the same cost as the Mini?

Yeah, right. It’s more likely that the expression of requirements is incomplete or inaccurate and one bidder has correctly addressed the client’s actual needs. 

The problem is that the respondent who scored 9 may object. 

“What do you mean I only scored a 9, even though I met all of your requirements without any deficiencies? If you had asked for those value adds that you are telling me made all the difference, I could have offered them, too! But you didn’t ask for them!”

Yes, I’m being pedantic, but responses that meet all of our needs should score a 10, and responses that meet all of our needs and exceed them in some respects should also score a 10.

Paul Rogers is Practice Manager at Landell. The opinions expressed in this article are the author’s alone and are not necessarily the opinions of Landell.