Workshops / Use-case shortlist + scoring

AI use-case shortlist and scoring: from a pile of ideas to a ranked plan

Most companies that ask about AI already have a list of twenty or thirty ideas collected from department heads, vendor demos and a board slide, and no defensible way to pick the first two. This is for the person who has to make that pick, fund it, and answer for it six months later.

Why the pile exists and why it does not shrink on its own

Ideas are cheap to collect and expensive to compare. Sales wants lead scoring, finance wants invoice extraction, support wants a chatbot, and the CEO saw a demo of an agent that books meetings. Each idea arrives with its own advocate and its own implied metric, so any discussion about priority turns into a discussion about whose department matters more.

Left alone, the pile resolves in one of two ways: the loudest idea gets built first, or the most technically interesting one does. Both are lotteries. The first AI project a company ships sets the budget, the tolerance and the credibility for every one after it, and a lottery is a poor way to choose it.

The five scoring factors

A scoring model replaces advocacy with a shared rubric. We score every idea on five factors, each from 1 to 5, in the room with the people who own the process. The numbers matter less than the argument that produces them.

FactorThe question it answersWhat a 5 looks like
ValueWhat changes in hours, euros or errors if this works?A measurable line item: 40 hours a week, or €8,000 a month in write-offs
FeasibilityCan this be built with current tools and current models?A known pattern: extraction, classification, drafting, RAG over documents
Data availabilityDoes the input exist, and can a workflow reach it?Structured, behind an API, with history to test against
RiskWhat happens when the model is wrong?Errors are visible, reversible and caught before a customer or a ledger sees them
Time-to-impactHow soon does anyone notice a difference?A first version live and measured within 30 days

Two rules keep the model honest. First, value and risk are scored by the process owner, not by the technical team; feasibility and data are scored the other way around. Second, any idea that scores 1 on data availability or 1 on risk (an error that is expensive and hard to catch) is parked regardless of its total. A high-value idea with no data is a data project, not an AI project, and the data readiness check is where it goes next.

How a workshop runs the scoring

  1. Collect and de-duplicate. Every idea goes on the wall in the same format: who triggers it, what goes in, what comes out, who checks it. "A chatbot" is not an idea in this format; "answer warranty questions from the product manual, with escalation to support" is. Thirty ideas usually collapse to fifteen.
  2. Attach a number to each. Volume per week, minutes per item, error rate, cost per error. If nobody in the room knows, that is a finding: the idea gets a placeholder and a name responsible for finding out.
  3. Score in pairs. Process owner and technical lead score each idea together, factor by factor, and disagreements are argued until both can defend the number. This is the slow part and the useful part.
  4. Apply weights and thresholds. The default is equal weights. A company under regulatory pressure doubles risk; a company with a runway problem doubles time-to-impact. The parking rules apply before any total is read out.
  5. Pick three to five, and write down why the rest lost. Rejected ideas are not deleted. Each gets a one-line reason ("no data until the CRM migration finishes in Q1") so the list can be re-scored later without re-arguing.

A half-day session is enough for most companies under a few hundred people. In Hilluter's workshops the scoring takes the middle third of the day; the rest goes to the data check and the architecture sketch for the top candidates.

A worked example: a wholesaler with 22 ideas

A building-materials wholesaler with about 120 staff came in with 22 ideas. After de-duplication there were 14. Five of them, with the scores as they left the room:

  • Order-confirmation extraction (supplier PDFs into the ERP): value 4, feasibility 5, data 4, risk 4, time 5. Total 22.
  • Support email triage (intent, urgency, routing to the right desk): value 4, feasibility 5, data 5, risk 4, time 4. Total 22.
  • Demand forecasting per branch: value 5, feasibility 2, data 2, risk 3, time 1. Total 13.
  • Automatic quote generation: value 5, feasibility 3, data 3, risk 1, time 3. Parked on risk.
  • Public website chatbot: value 2, feasibility 4, data 3, risk 3, time 3. Total 15.

Forecasting had the most enthusiasm and the lowest score. The data lived in four branch spreadsheets with different column names, and nobody could say what a 10% better forecast was worth in stock terms. The quote generator was parked because a wrong price in a quote is a binding offer, and the room could not name a cheap way to catch it. The chatbot scored low on value because the same 40 questions made up most of the volume and a rewritten FAQ page would answer half of them for free. The two winners were unglamorous, measurable and live within a month.

What a bad shortlist looks like

The signs of a weak shortlist:

  • Every item is a platform ("an AI layer", "a knowledge assistant") rather than a process with a trigger and an output.
  • No item has a number attached. Value is described with adjectives.
  • The top item is the one the sponsor mentioned in the kickoff, and nothing below it was seriously argued for.
  • Nothing on the list can be live inside 30 days, so the first proof of value is a quarter away.
  • Only the technical team found the ideas interesting; the process owners were informed rather than consulted.
  • No idea was rejected. A shortlist of twelve is a long list with a new name.

A shortlist that fails these tests is not necessarily wrong, but it is undefended. The first time a project slips, there is no record of why it was chosen over the alternatives, and the argument starts over.

What the shortlist feeds into

The scored list is the input to everything that follows. The top candidates go through a data readiness check, get a one-page architecture sketch, and are laid out on a 30/60/90-day roadmap with owners and kill criteria. The scoring sheet stays in the repository next to the roadmap; when a model gets cheaper or a data source gets cleaned up, the rejected ideas are re-scored rather than re-pitched.

If you have a list and no ranking, send us the list and we will come back with the questions we would ask before scoring it.

Frequently asked questions

How many ideas should end up on the shortlist?

Three to five. Fewer than three leaves no fallback when the first one hits a data problem; more than five means nothing was really decided.

Who should be in the room for the scoring?

The people who own the processes being scored, someone who can speak for the data and systems, and whoever controls the budget. Vendors and enthusiasts can present ideas but should not hold the pen.

Can the scoring be done without a workshop?

Yes, with a spreadsheet and the same rubric. The workshop's value is that the process owner and the technical lead score each idea in the same conversation, which is where most wrong assumptions surface.

This article expands Use-case shortlist + scoring from the Workshops service on the main page.

WANT THIS APPLIED TO YOUR PROCESS?

Tell us what the workflow does, where it hurts and which tools are involved. We reply with next steps and a proposed approach.