Skip to content

How to evaluate an internal AI pilot after 30 days

A day-30 review for deciding whether an internal AI pilot has enough evidence to expand, repair, stop, or keep observing.

Fulton Ring8 min readAI evaluation · Operations · AI governance

Day 30 is a checkpoint on the calendar. It becomes a decision point only when enough eligible work has occurred and the team has preserved the evidence.

Accounts created, prompts sent, and documents indexed describe activity, not results. A funding decision needs to know what work actually changed, which answers held up under review, whether the intended users ever encountered the task, and what the system cost to keep running.

The review below is Fulton Ring’s synthesis of public research and implementation guidance. It supplies questions and record formats, rather than universal thresholds. Set the use-case-specific targets and stop conditions before launch.

Begin with evidence sufficiency

Count the opportunities the pilot actually had to help. Break them down by task family, user role, and material risk. Then examine how many cases received qualified review.

A day-30 result is decision-grade only when:

  • eligible work occurred often enough to cover the important task variants;
  • the source boundary and system configuration remained known;
  • intended users had a reasonable opportunity to use the capability;
  • reviewers examined routine cases, difficult cases, and expected refusals;
  • workflow and operating measures use the definitions established before launch.

Document the counts. Avoid a universal minimum: a recurring support question and a rare safety-related decision require different evidence. When the observed work is too sparse or unrepresentative, record insufficient exposure. Extend the observation window, narrow the claim, or end the pilot if waiting has no practical value.

The NIST AI Risk Management Framework calls for measurement under conditions similar to deployment and for ongoing monitoring. This makes context and exposure part of the evidence, rather than footnotes to a benchmark score.

Ask four questions

Use the same four questions for the written review, the owner meeting, and the decision record.

Did useful work change?

Compare the pilot period with the baseline using the original definitions. Choose measures close to the task, such as time to a reviewed answer, completed eligible work, expert interruptions, rework, backlog, or service delay.

Include the full path. Model response time excludes reading, checking a source, correcting an answer, and returning to the old process. A useful elapsed-time measure includes them.

Could reviewers verify the answers?

Review claims against the approved evidence. Separate factual correctness from citation quality and safe refusal. Keep a failure log that lets the team find recurring causes: missing source material, retrieval error, stale content, unsupported synthesis, permission failure, or ambiguous task definition.

Did intended users encounter and complete eligible work?

Use eligible opportunities as the denominator. Report how many relevant cases each user group encountered, how many they attempted with the pilot, and how many reached a useful completion. Capture abandonment and fallback reasons.

The distinction matters in field research. A 2025 NBER working paper reports a six-month randomized rollout of a generative AI tool across 66 firms and 7,137 knowledge workers. Among treated workers, use was uneven; the paper reports work-pattern changes for active users and does not find broad changes in task composition from access alone. The study concerns a particular office suite and includes disclosed ties to its maker. The measurement lesson holds regardless: access, use, and changed work are three different things.

What did operation require?

Record internal time and direct cost spent on connector repair, access changes, source cleanup, evaluation, user support, prompt or policy changes, incident response, and owner decisions. Distinguish setup work from recurring maintenance where possible.

The NIST Generative AI Profile recommends production monitoring, incident and near-miss records, source and citation verification, override review, and deactivation criteria. Those activities belong in the operating record.

Build the review set from real work

Start with cases captured during the baseline and pilot. Include frequent tasks, known edge cases, high-consequence boundaries, and questions that the source base or user’s permissions cannot answer.

For each case, preserve:

FieldReview use
RequestThe user’s original wording and task
User roleThe permissions and expected workflow
Expected resultThe answer, decision support, or escalation a reviewer accepts
Source versionThe evidence available at the time of the run
Critical claimsFacts that must be present and correct
Prohibited behaviorClaims, disclosure, or action the system must avoid
Run recordModel, instructions, retrieved passages, configuration, and timestamp
Reviewer judgmentResult, correction, severity, and cause

Keep a portion of the set outside routine tuning. Add confirmed failures to a regression set. Where output variation affects risk, rerun the same case and report the variation.

Test refusal with a two-by-two matrix

An internal assistant needs to answer when evidence permits and yield when it does not. Evaluate both behaviors.

CaseSystem answersSystem abstains or escalates
Answerable and authorizedReview correctness and completenessOver-refusal: useful work was available but declined
Unanswerable or unauthorizedUnsafe answer: unsupported or impermissible responseCorrect abstention: decline with a useful next step

This matrix prevents a superficially safe system from receiving credit for declining everything. It also makes unsafe helpfulness visible.

Research on selective question answering treats the decision to answer as part of system quality, especially under distribution shift. The paper is a research benchmark, not an operational threshold. The relevant idea is to measure answer quality together with the system’s selection behavior.

Audit citations claim by claim

A citation icon does not establish that the answer is supported. Review four properties:

PropertyQuestion
SupportDoes the cited passage fully, partly, or fail to support the attached claim?
CoverageDo all material, externally verifiable claims have an appropriate citation?
ReachabilityCan the reviewer open the exact passage or record without reconstructing the search?
Source eligibilityWas the source approved, current enough, and available to that user’s role?

ALCE formalized citation correctness and completeness for generated answers. Later fine-grained citation research compared full, partial, and unsupported citations and found that no single automatic metric performed best across all evaluations. Automated checks can help locate patterns; domain reviewers remain necessary for a consequential internal workflow.

Conversations add another failure mode: the current answer may depend on an earlier turn whose premise was wrong or whose evidence changed. The mtRAG benchmark was designed to evaluate retrieval-augmented systems across multiple turns. For an operational pilot, include follow-up questions, corrections, topic shifts, and references such as “that unit” or “the earlier procedure” in the review set.

Use one review sheet

Complete this sheet with the process owner and a qualified reviewer.

QuestionBaseline or planned boundaryPilot evidenceReviewer conclusion
Did useful work change?
Could reviewers verify answers?
Did intended users encounter and complete eligible work?
What did operation require?

Then add the evidence ledger:

Evidence itemRecord
Eligible opportunities by task family and role
Cases reviewed, with routine and boundary coverage
Correct completion and rework
Correct abstention, over-refusal, and unsafe answers
Citation support, coverage, reachability, and eligibility
Source, connector, and permission failures
Incidents, overrides, and unresolved escalations
Setup hours, recurring maintenance hours, and direct cost
Known changes to source data or system configuration

Report counts with definitions and denominators. A percentage without the number of eligible cases can create false confidence.

Choose one of four decisions

Expand one boundary

Use this decision when the workflow result is useful, the agreed quality and safety conditions hold, users had adequate exposure, and the operating burden has a named owner. Expand one dimension: an adjacent task family, a source set, or a user group. Preserve the evaluation and regression records.

Fix a named defect and rerun

Use this when the evidence points to a specific repair, such as a missing manual, a retrieval failure for part numbers, a permission mapping error, or excessive review time in one task family. Record the owner, change, affected cases, and next observation window.

Stop

End the pilot when the work is too immaterial, the source base cannot support it, the risk boundary fails, no one can own operation, or the burden outweighs the observed value. Preserve the record so a later team does not repeat the same experiment without addressing the cause.

Insufficient exposure

Use this when the calendar has advanced but the evidence has not. State which task variants or user opportunities are missing, whether the window will be extended, and the latest date at which the company will decide. This status should have an owner and an end date.

The day-30 output

The final artifact can be short. It should contain the original use-case boundary, baseline, exposure record, reviewed results, failure ledger, operating effort, and one of the four decisions above. Attach the evaluation set and run records so another reviewer can reproduce the reasoning.

Thirty days can reveal a great deal in a frequent, bounded workflow. It can also produce an honest finding of insufficient exposure. Both are more useful than a dashboard that confuses availability with changed work.

Sources

This review method is Fulton Ring’s operational synthesis. No universal sample size or pass rate is claimed, and no client outcome is represented.

Primary references: NIST AI RMF 1.0, NIST Generative AI Profile, ALCE, fine-grained citation evaluation, mtRAG, selective question answering under domain shift, and Shifting Work Patterns with Generative AI. The NBER paper is a working paper and has not been peer reviewed.