How we score use cases
Every use case in the compendium is scored the same way, in the open. Here is the whole framework: the six dimensions, the evidence rubric, and the formula behind the start-here score. No black box.
The six dimensions we score
Three dimensions feed the start-here score directly. The other three describe the shape of the work so you can judge fit.
Value potential
1 to 5, where 1 is marginal and 5 is transformational.
The only dimension that raises the score.
Implementation complexity
1 to 5, where 1 is straightforward and 5 is very hard.
A cost: a higher score lowers the start-here score.
Data readiness burden
1 to 5, where 1 works with what most firms already have and 5 needs substantial data groundwork.
A cost: a higher score lowers the start-here score.
Autonomy level
L0 assistive, L1 human approves each step, L2 human approves the outcome, L3 human on exceptions, L4 fully autonomous.
How much the agent does unsupervised. Descriptive, not scored.
Time to value
Under 6 weeks, 6 to 12 weeks, or a quarter or more.
How long to a first working result. Descriptive, not scored.
Evidence tier
Tried and tested, emerging, or frontier.
How well-proven the use case is. Weights the score (see below).
The evidence-tier rubric
The evidence tier is the honesty control. It sets how much confidence a use case has earned, and it directly weights the score.
The highest bar. A use case earns this only when it is backed by two or more independent, non-vendor sources, or by a QuantSpark case study we have actually delivered. This is the only tier that carries full confidence weight in the score.
Credible and gaining traction, but the evidence base is still early. Vendor claims and single sources sit here, not in tried and tested.
Experimental. We flag it as unproven so you can weigh the risk honestly, rather than take a demo at face value.
The start-here score, in plain English
The score is a single number from 0 to 100. Higher means a better first candidate: proven, high value, and light on complexity and data groundwork.
- Step 1 · Inputs
We take three scores, each on the 1 to 5 scale: value, complexity and data burden. If a score is missing, it counts as a neutral 3 rather than helping or hurting.
- Step 2 · Merit
Merit rewards value twice and penalises the two burdens equally: merit = (2 × value) − complexity − data burden. That runs from −8 (low value, high cost) to +8 (high value, low cost).
- Step 3 · Map to 0 to 100
We map merit linearly onto a 0 to 100 scale, so the worst possible merit lands at 0 and the best at 100.
- Step 4 · Weight by evidence
We scale that by a confidence weight from the evidence tier: tried and tested counts in full (×1.0), emerging counts at ×0.75, and frontier at ×0.5. An unknown tier is treated as ×0.75. So at equal merit, proven work always ranks above experimental work.
Three worked examples
- A tried-and-tested use case with value 5, complexity 1 and data burden 1 scores 100.
- A neutral tried-and-tested case at 3 / 3 / 3 scores 50.
- A case with value 1, complexity 5 and data burden 5 scores 0.
How we label the number
The bands are a reading aid on top of the score, not part of the formula.
| Score | Label |
|---|---|
| 70 to 100 | Start here |
| 50 to 69 | Strong candidate |
| 30 to 49 | Worth exploring |
| 0 to 29 | Longer horizon |
The score is computed at read time and never stored. Re-scoring the whole library is a change to this formula, not a data migration, so the ranking stays consistent and auditable.