Back to the compendium

How we score use cases

Every use case in the compendium is scored the same way, in the open. Here is the whole framework: the six dimensions, the evidence rubric, and the formula behind the start-here score. No black box.

The six dimensions we score

Three dimensions feed the start-here score directly. The other three describe the shape of the work so you can judge fit.

Value potential

1 to 5, where 1 is marginal and 5 is transformational.

The only dimension that raises the score.

Implementation complexity

1 to 5, where 1 is straightforward and 5 is very hard.

A cost: a higher score lowers the start-here score.

Data readiness burden

1 to 5, where 1 works with what most firms already have and 5 needs substantial data groundwork.

A cost: a higher score lowers the start-here score.

Autonomy level

L0 assistive, L1 human approves each step, L2 human approves the outcome, L3 human on exceptions, L4 fully autonomous.

How much the agent does unsupervised. Descriptive, not scored.

Time to value

Under 6 weeks, 6 to 12 weeks, or a quarter or more.

How long to a first working result. Descriptive, not scored.

Evidence tier

Tried and tested, emerging, or frontier.

How well-proven the use case is. Weights the score (see below).

The evidence-tier rubric

The evidence tier is the honesty control. It sets how much confidence a use case has earned, and it directly weights the score.

Tried and tested

The highest bar. A use case earns this only when it is backed by two or more independent, non-vendor sources, or by a QuantSpark case study we have actually delivered. This is the only tier that carries full confidence weight in the score.

Emerging

Credible and gaining traction, but the evidence base is still early. Vendor claims and single sources sit here, not in tried and tested.

Frontier

Experimental. We flag it as unproven so you can weigh the risk honestly, rather than take a demo at face value.

The start-here score, in plain English

The score is a single number from 0 to 100. Higher means a better first candidate: proven, high value, and light on complexity and data groundwork.

  1. Step 1 · Inputs

    We take three scores, each on the 1 to 5 scale: value, complexity and data burden. If a score is missing, it counts as a neutral 3 rather than helping or hurting.

  2. Step 2 · Merit

    Merit rewards value twice and penalises the two burdens equally: merit = (2 × value) − complexity − data burden. That runs from −8 (low value, high cost) to +8 (high value, low cost).

  3. Step 3 · Map to 0 to 100

    We map merit linearly onto a 0 to 100 scale, so the worst possible merit lands at 0 and the best at 100.

  4. Step 4 · Weight by evidence

    We scale that by a confidence weight from the evidence tier: tried and tested counts in full (×1.0), emerging counts at ×0.75, and frontier at ×0.5. An unknown tier is treated as ×0.75. So at equal merit, proven work always ranks above experimental work.

Three worked examples

  • A tried-and-tested use case with value 5, complexity 1 and data burden 1 scores 100.
  • A neutral tried-and-tested case at 3 / 3 / 3 scores 50.
  • A case with value 1, complexity 5 and data burden 5 scores 0.

How we label the number

The bands are a reading aid on top of the score, not part of the formula.

ScoreLabel
70 to 100Start here
50 to 69Strong candidate
30 to 49Worth exploring
0 to 29Longer horizon

The score is computed at read time and never stored. Re-scoring the whole library is a change to this formula, not a data migration, so the ranking stays consistent and auditable.