Skip to content

How to choose the first operational AI use case worth funding

Five gates and four practical prioritizers for choosing an internal AI pilot that a small operating team can evaluate and maintain.

Fulton Ring8 min readAI strategy · Operations · Evaluation

Most companies can produce a long list of things AI might do. That list is a poor basis for spending money.

A useful first project is narrower than that. It covers one recurring piece of work for a known group of users, over evidence the system can already reach, inside a boundary where a mistake is survivable, and it ends in a decision the company is willing to make. Selection means finding that project before implementation starts.

The method below has two stages. Five gates remove candidates that cannot yet support a responsible pilot. Four prioritizers help rank the candidates that remain. There is no universal score or pass mark. The final choice belongs to the people responsible for the work.

Describe the work in one sentence

Use this form:

When [situation] occurs, [role] needs to [complete a task or make a decision] using [named sources]. Today, [observable problem] gets in the way.

For example:

When a machine stops, the maintenance supervisor needs to identify the relevant procedure and recent service history. Today, that search interrupts senior technicians and delays diagnosis.

The sentence should name work that already happens. “Give everyone an AI assistant” describes a deployment. It leaves the work, expected evidence, and success condition undefined.

This sequence follows the context-of-use logic in the NIST AI Risk Management Framework: understand the intended purpose, affected people, operating setting, and risk before selecting measurements or controls.

Apply five must-pass gates

A candidate can be valuable and still fail a gate. Record the reason and revisit it when the underlying condition changes.

1. Work

Can an operator point to the task as it happens today?

Pass this gate when you can name the trigger, the person doing the work, the output they need, and the current path from request to completion. Observe at least one normal operating cycle. Record elapsed time, handoffs, interruptions, rework, and unresolved cases where those measures fit.

The baseline does not need a speculative dollar conversion. A consistent count of expert interruptions can be more useful than an invented productivity estimate.

2. Sources

Does the company have enough usable evidence for the task?

List the manuals, records, tables, messages, or systems an experienced person actually consults. For each source, identify who owns it, how current it is, who may see it, and which source takes precedence when two disagree. Test whether the material can be retrieved and cited at the passage or record level.

If essential knowledge exists only in one person’s memory, include knowledge capture in the pilot. A retrieval system cannot recover a fact the company has never recorded.

The NIST Generative AI Profile calls for attention to data provenance, suitability, permissions, and realistic testing. NIST MEP’s 2025 guidance for manufacturers likewise places data readiness and cybersecurity among the conditions for effective adoption.

3. Judgment

Can a qualified person say whether the result is useful?

Name the reviewer before the build starts. Ask that person to assemble representative examples, expected answers or decisions, acceptable sources, and known edge cases. Define what the system should do when evidence conflicts or runs out.

This gate prevents a demonstration from becoming its own evaluation. A fluent answer is easy to admire. A reviewer with the underlying record can determine whether it helps with the work.

4. Safety

Can mistakes be contained and corrected?

Document the consequence of a wrong answer, an omitted exception, stale evidence, or access by the wrong role. Define human review, escalation, logging, and rollback. Keep consequential actions behind an approval step until the evidence supports a wider boundary.

The first pilot should allow the team to inspect failures without exposing customers, employees, equipment, or finances to an unacceptable outcome.

5. Evaluation

Will enough real work occur to support a decision?

Choose measures tied to the task. Possible measures include time to a reviewed answer, completed eligible cases, rework, expert interruptions, correct handling of unanswerable questions, and operating hours required from the owner. Set the definitions before launch.

Also ask when enough eligible cases are likely to occur. Thirty calendar days tells you very little if the target task appears twice during that period.

The UK’s 2026 guidance on evaluating AI interventions emphasizes a clear theory of change, credible comparison, and evidence designed around the intervention. It does not provide a universal pilot template. Its useful contribution here is the discipline of stating how the proposed capability is expected to change the observed outcome.

Prioritize the candidates that pass

Compare the survivors in a working session with the process owner and reviewers. Use four dimensions. Plain language is usually more honest than a weighted total.

DimensionQuestionEvidence to bring
FrequencyHow often does eligible work occur?A count from a normal operating period
Operational consequenceWhat changes when the work is faster or better?A downstream delay, quality issue, service effect, or risk the owner recognizes
Current burdenWhere does the present process consume scarce attention?Search time, handoffs, interruption, queueing, or rework
Operating effortWhat will it take to connect, review, secure, and maintain the capability?Named systems, access dependencies, reviewer time, and expected upkeep

Favor a candidate with meaningful consequence and enough repetition to learn quickly. Treat operating effort as real work. Small firms often have less internal capacity to absorb implementation and governance overhead; the OECD’s 2025 report on AI adoption by SMEs documents that adoption barriers and use in core business activities differ materially from the experience of large firms.

Learn from task-level evidence

Research supports careful boundaries, though it cannot predict a company’s return.

In a preregistered experiment involving 758 Boston Consulting Group consultants, GPT-4 improved performance on tasks judged to be within the model’s capability frontier and reduced performance on a task outside it. The researchers describe this uneven boundary as a “jagged technological frontier”. The experiment ran in June 2023 on researcher-designed tasks, so its effect sizes do not transfer directly to a plant, distributor, or field-service team.

A separate field study followed 5,172 customer-support agents at one large software company. The peer-reviewed results found an average productivity increase, with larger gains among less-experienced workers. That finding argues for measuring the specific user group and workflow. It is not a forecast for another company.

The NIST MEP account of CJB Industries is useful for another reason. The reported work combined AI with broader changes to data collection and production processes. A pilot review should preserve that attribution boundary rather than credit every improvement to the model.

Write a one-page pilot contract

Complete this with the operator who owns the work. Keep it to one page. If a field cannot be answered, the candidate may need more discovery before funding.

FieldWhat to record
SituationThe event that creates eligible work
UserThe role expected to use the capability
TaskThe answer, report, classification, or decision support required
Current pathHow the work is completed today
BaselineMeasures collected with their definitions and period
Source boundaryIncluded systems and documents, authority rules, freshness, and permissions
ReviewerPerson accountable for the evaluation set and failure review
Allowed behaviorWhat the system may answer, draft, recommend, or initiate
EscalationConditions that require a person or the existing process
Evaluation setReal task families, edge cases, and unavailable or unauthorized questions
MeasuresWorkflow outcome, answer quality, use, and operating effort
Evidence windowWhen enough eligible work is expected to occur
Stop conditionsSafety breach, operating burden, or missing evidence that ends the pilot
DecisionExpand one boundary, repair a named defect, stop, or collect more evidence
Cost and capacityExternal spend plus named internal time for access, review, and maintenance

Attach the source inventory, the baseline record, the evaluation set, and the decision log. Without them the contract is a statement of intent rather than something anyone can act on.

Make the boundary smaller before launch

Limit the first release to one user group and one task family. Include only approved sources. Preserve the existing route for cases that need escalation. Select a decision date based on expected work volume.

A well-chosen pilot gives management a defensible answer about the next investment. It can also reveal that the source base needs repair, the task is too rare, or the operating burden is too high. Those are useful results because they arrive before a broad rollout.

Sources

This method is Fulton Ring’s synthesis of public research and implementation guidance. No client results or interview findings are represented in this article, and the framework has not been validated as a universal scoring instrument.

Primary references: NIST AI RMF 1.0, NIST Generative AI Profile, NIST MEP’s 2025 manufacturing guidance, UK guidance on AI impact evaluation, OECD’s 2025 SME report, the CJB Industries case published by NIST MEP, Dell’Acqua et al., and Brynjolfsson, Li, and Raymond.