How to Test Whether an AI Workflow Is Actually Better

To evaluate the impact of an AI solution, the current process needs to be understood first. Pilot results also need to be measured in an appropriate and clear way.

Rikke Egeberg Tankard Written byRikke Egeberg Tankard
Read time7 min read Published
Three pharmaceutical professionals compare current and AI-supported workflow records beside a laptop showing paired pilot results
Current-process evidence and AI-supported results are compared in the same review.HTO & Beyond
Overview

In August 2026, researchers at Rambam Health Care Campus and the Technion tested an AI decision-support system across two emergency department units. Expert reviewers judged 99 of 100 sampled outputs clinically appropriate. Use of the system still fell from about 68% in the first week to 30% in the fourth. Average length of stay was 4.9 hours in both units.

The result was mixed because the AI's answers were only one part of the test. No adverse events were found. Reviewers judged 99 of 100 sampled outputs appropriate. Yet the average length of stay was 4.9 hours with or without the system. A possible 9.4-minute reduction in consultation time was too uncertain to count as a proven improvement.

Use also fell as the test continued. The study linked this to workload. For every additional hour into a shift, the odds that a doctor used the system fell by about 28%. Doctors were more likely to use it for radiology questions than for other tasks. The detail matters. A good answer does not improve a process if people cannot fit the tool into their work. It may also show that the tool is useful for one narrow task, rather than for every task in the process.

AI firms now make clear claims about speed. IQVIA says its service helped start studies 33% faster. It also says data cleaning was 50% faster. Its release gives no details on the test plan, number of studies, or how it compared the results. Local tests would therefore be needed before applying those figures to another company or process.

Practical pilot

Six checks before deciding whether AI is better

  1. 01Problem and use caseName the work that needs to improve and why AI may help.
  2. 02Current baselineRecord today's time, errors, rework, and handoffs.
  3. 03Hypothesis and measuresSet the expected improvement and the quality limit.
  4. 04Representative casesInclude normal work and difficult cases that drive risk.
  5. 05Real workflow testMeasure quality, time, human effort, and serious failures.
  6. 06Honest decisionAdopt, narrow, retest, or stop.

Start with the current process

A baseline records how the current process performs before AI enters it. It should cover the same outcome that the pilot claims to improve.

If the claim concerns speed, record total time from the first input to the final approved output. Include waiting, rework, and escalation. If the claim concerns quality, define the error types and their severity. Count the work that reviewers correct and the errors that escape them.

In lab work, researchers test the base material before they add the active material. An AI pilot needs the same care. A faster result tells us little when the pilot gets easier cases, more staff, or a shorter review than the current process.

Looking closely at the current process also helps to choose the right AI use case. A slow handoff may be the problem. So may repeated data entry or a review step with many simple corrections. Each gives a company something specific to improve. The baseline shows how often the problem happens and what it costs today.

This makes an established process a stronger starting point than a completely new one. Without a current way of working, there is no reliable baseline for deciding whether AI made it better. A new process can still be tested for feasibility, but an improvement claim needs a fair comparison.

A clear problem and use case also support risk management from the start. The first pilot can focus on work where value can be measured. Mistakes should also be found before they affect a GxP decision. The company can then avoid beginning with an attractive idea that has no clear benefit or puts a critical control at risk.

Define the goal and how it will be measured before starting

Define the hypothesis and the expected change before starting. For example: “The AI-supported route will cut normal review time by 20%, while serious missed cases stay at or below the current level.” The process owner should set the target and quality limit to fit the use and its risk.

Use cases that represent the real work, including the difficult ones that drive risk or rework. Keep definitions, time windows, and quality checks the same on both sides. For a higher-risk use, the AI can first run in shadow mode while the existing process continues to make the live decision.

AWS now gives each record match its own confidence score. A fair test would use records with known answers and compare several score levels. It should count wrong matches, missed matches, and review work. Two pilots can report the same headline score while making very different mistakes.

Test the way people will really use it

The test must include the people who do the work, the steps they follow, and the pressure they face during a normal day.

FDB’s Script Agent turns a patient visit into a draft prescription for a clinician to review. FDB says clinicians can change or approve each draft. Its release gives no outside error rate. A local test should record how often clinicians make changes, which fields they change, how serious the errors are, and how long the review takes.

At Rambam, the pressure of a normal shift changed the result. The sampled outputs scored well, but use fell during the four-week study and as work built up later in each shift. The system could only add value if doctors kept using it under real pressure.

The same point applies to IQVIA's service across several trial steps. A company should measure the path from study planning to the final locked data. Total time, quality, manual work, and work passed to the next person belong in the same result.

Stay true to the method

When the pilot ends, return to the written hypothesis and the original risk analysis. Did the AI create the expected improvement? Did quality remain within the agreed limit? Did the work become easier, or did review and handoffs move the effort somewhere else? The test may also reveal new risks or show that the first controls need to change.

The result may support wider use, a narrower use, another test, or a stop. A mixed or negative result is useful when it shows where the idea, workflow, or choice of use case needs to change. Reporting that result honestly is more valuable than selecting one flattering number.

NIST, a US standards body, gives similar advice. It calls for clear benchmarks and tests that match the place where the AI will work. It also includes the way people and AI work together. A research guide called DECIDE-AI asks early clinical tests to study real use, human behavior, and fit with the current setting.

Takeaways

What the pilot should establish

  • A baseline turns an AI improvement idea into a claim that can be tested.
  • Real users, difficult cases, and normal work pressure can change the result.
  • The final decision should return to both the hypothesis and the original risk analysis.

A good pilot gives a clear answer about a defined application: whether AI improved the process, under what conditions, and at what risk. The data and reasoning need to be understandable to business leaders, data specialists, QA, process owners, and other affected stakeholders. That shared understanding supports sound decisions and helps keep the use compliant as it moves into routine work.

Sources

  1. Prospective evaluation of a large language model clinical decision support system in the emergency department, Nature Medicine, 19 August 2026.
  2. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI, Nature Medicine, 18 May 2022.
  3. AI Risk Management Framework Core, NIST.
  4. IQVIA’s Predictive Clinical Development provides sponsors with significant efficiencies, IQVIA, 3 September 2026.
  5. AWS Entity Resolution adds record-level confidence scores for ML matching, AWS, 9 September 2026.
  6. FDB moves ambient listening beyond notes with AI-powered prescribing agent, FDB, 24 August 2026.
Rikke Egeberg Tankard

About the author

Rikke Egeberg Tankard, PhD

Rikke is GxP & Compliance Lead at HTO & Beyond. She makes AI innovation practical, compliant, and trusted in GxP environments.