Overview
A promising AI idea needs an operating home. Someone has to own the problem, define what better work looks like, decide which evidence counts, and bring the people affected by the change into the process. Teams also have to agree on the decision the system will support, the data it can use, the limits people must understand, and the way the new process fits into daily work.
This is where implementation begins. A clear connection between the problem, the people, and the operating conditions gives the technical team something concrete to build toward. It also gives the organization a basis for judging whether the work deserves to continue, which matters because a model can perform well in a test and still create extra work once it enters a live process.
Start with the real work
The first useful question is simple: which decision or task should improve? A team reviewing safety cases needs a different result from a team preparing regulatory documents, even when the teams use similar language models. A precise problem statement creates the boundary for the project and gives every later choice a point of reference.
A workable problem statement should tell the team:
- which person or group experiences the problem in daily work;
- which decision, handover, or task should change;
- which evidence would show a useful improvement;
- which data and exceptions belong inside the first test; and
- who owns the process when the pilot ends.
I would keep the statement close to the work. “Use AI to improve regulatory operations” gives the team a theme. “Help submission teams find approved evidence for a response and show the source behind every suggested passage” gives them a process, a user, and a result they can test.
Applied AI becomes useful when a team can point to the exact decision, the person making it, and the evidence that should improve.
The narrower formulation also makes disagreement useful. Quality, regulatory, technology, and operations can discuss the same proposed change because the team has named the work. That discussion usually exposes assumptions early, while the team can still change the scope at low cost.
Build an evidence plan before the prototype
A prototype answers whether a team can make something work under selected conditions. An evidence plan answers what the organization must learn before it can trust the result in a real process. I would write that plan while the team defines the use case, because the plan determines which data, reviewers, comparison, and test cases the prototype needs.1
The evidence plan can stay short. It should name the current baseline, the change the team expects, the people who will judge the result, and the boundary for the next decision. A baseline can include quality, time, reviewer effort, rework, or the number and type of exceptions, as long as the measure connects directly to the selected task.
Implementation readiness
The checks that turn an AI idea into a controlled first pilot
| Area | Core question | Evidence | Status |
|---|---|---|---|
| Problem | Which task or decision should improve? | Current workflow and baseline | Defined |
| Users | Who works with the output? | Named roles and working sessions | Defined |
| Review | Which outputs require a person’s judgment? | Review criteria and escalation route | Partial |
| Ownership | Who decides what happens after the test? | Named process owner | Open |
| Learning | What must the pilot settle? | Decision gate and recorded findings | Partial |
The table gives a project team a common view of readiness and preserves each question separately. That distinction matters because one unresolved ownership question can outweigh several completed technical checks.
Design with the people who carry the process
The people closest to the work can show where information arrives late, where a handover breaks down, and which exceptions consume the most time. Their knowledge shapes the problem definition and the conditions the system has to handle. Early involvement also gives the team a more honest view of adoption because people can explain which parts of the proposed workflow feel credible and where responsibility becomes unclear.
A useful working session starts with a real case. The team can follow how information enters the process, who checks it, which exceptions change the route, and what happens when the available evidence remains unclear. This gives the technical team details they can use and gives the people in the process a direct way to shape the design.2
“The team needs to explain who checks the output and what happens when confidence drops before the process is ready.”
Illustrative workshop comment for this article template
The quotation above demonstrates the treatment for a person’s exact words. The attribution stays attached to the quote, while the surrounding paragraphs carry the article’s argument. A pull quote works differently: it repeats one of the article’s own strongest lines as a visual pause.
Make responsibility visible inside the process
Implementation changes how people work, who reviews an output, and who answers when the system behaves unexpectedly. The team should place those responsibilities inside the workflow at the moments when they guide a decision. A policy document can support the work, while the process itself has to show each person what to do.
I would define the human role in this order:
- Name the decision. State which choice the AI output informs and who makes that choice.
- Define the review. Explain what the reviewer checks and which evidence they need.
- Set the boundary. Describe the conditions that send a case to another route or a more experienced reviewer.
- Record the outcome. Keep the reviewer’s decision and the reason available for learning and oversight.
Clear ownership also improves the pilot itself. Reviewers can tell the team which errors matter, where the explanation helps, and which part of the output creates extra work. The team can then change the system or the process using evidence from the people who carry the responsibility.
Learn through a controlled pilot
A controlled pilot should answer a small number of decisions that the organization has already named. The pilot might test whether reviewers can find relevant evidence faster, whether the system keeps source links intact, or whether the proposed review route handles common exceptions. Each test should connect to the next decision and to a person who owns that decision.
The example keeps the claims close to what the pilot can show. It gives the team a defined set of cases, qualified reviewers, observable outputs, and a decision at the end. It also leaves room for honest findings, including extra work or conditions the first design failed to cover.
Implementation sequence
From question to next decision
- 01DefineName the work and the expected result.
- 02TestUse representative cases and clear review criteria.
- 03ReviewRecord quality, effort, exceptions, and feedback.
- 04DecideChange, extend, pause, or stop the work.
The decision at the end should match the evidence stage. Early findings can support another controlled test. Repeated pilot evidence can support a limited rollout with defined review. Production evidence collected over time can support a wider implementation discussion.
Run the implementation check
Before committing to the next stage, I would want a clear answer in five areas. Together, these answers show whether the idea has enough structure to move from possibility to practice and whether the organization can learn responsibly from the next test.
- Problem
The team has named the task, user, current baseline, and expected change.
- Evidence
The test will produce information that supports a defined next decision.
- Ownership
One person owns the result, the process, and the decision after the pilot.
- Human role
Reviewers know what to check, when to escalate, and which authority they carry.
- Learning
The team will record findings, unresolved questions, and changes for the next cycle.
The checklist gives a team a practical pause before it commits more time and attention. It also creates a common language for the people who sponsor, build, review, and use the system, which makes the next decision easier to explain.
Notes and references
- National Institute of Standards and Technology, AI Risk Management Framework. NIST organizes AI risk work around governance, context, measurement, and management.↩
- HTO & Beyond, Applied AI Use Cases for Pharma & Life Sciences. The library connects AI capabilities to specific work, value, requirements, and implementation conditions.↩
- HTO & Beyond, The HTO Applied AI Method. The method begins with a real problem and carries the work through testing, implementation, and adoption.




