1. Choose a real recurring task
Do not test an AI tool with a clever prompt that has no operational value. Choose a task that happens often enough to matter: drafting a first outline, classifying non-sensitive feedback, turning meeting notes into actions, or comparing a document against a checklist.
Write one sentence describing the desired outcome. Then define what “good enough” means before seeing the AI output. This prevents a polished demo from moving the goalposts.
Use non-sensitive or approved test data. Do not paste confidential, personal, regulated, or proprietary information into a tool without authorization and an appropriate data review.
2. Measure the current workflow
Complete the task once without the new tool. Record active work time, waiting time, number of corrections, and the quality standard you reached. That baseline is imperfect, but it is more useful than comparing the tool to a vague memory.
Run the same task type several times during the week. A single excellent result may be luck; a single bad result may be an unusual input. Note the conditions around each run.
3. Set boundaries before speed
Map who could be affected if the output is wrong, biased, disclosed, or used outside its intended context. Decide which inputs are permitted, which outputs require review, who is accountable, and when the workflow must stop.
NIST’s AI Risk Management Framework organizes responsible risk work around Govern, Map, Measure, and Manage. A small pilot does not replace formal governance, but these questions help reveal whether a seemingly simple tool touches larger risks.
4. Score the full workflow
| Dimension | Question | Evidence |
|---|---|---|
| Output quality | Did it meet the pre-set standard? | Errors, omissions, reviewer rating |
| Total effort | Did prompting and correction actually save time? | Minutes per completed task |
| Reliability | Did quality hold across normal inputs? | Pass rate across the week |
| Risk fit | Can the workflow stay inside approved boundaries? | Data, privacy, bias, and review notes |
| Cost fit | Is the recurring value greater than total cost? | License, setup, training, review time |
5. Make a decision, not a vague recommendation
Adopt when evidence shows durable value within acceptable boundaries. Constrain when the tool helps only with certain inputs or requires mandatory review. Improve when the workflow design—not necessarily the tool—needs another test. Stop when value is weak, risk is unacceptable, or the old process is simply better.
Record the decision date and the evidence used. AI products and terms change, so a past adoption is not a permanent approval.
Primary framework source
See the NIST AI Risk Management Framework and Generative AI Profile. Sources checked July 25, 2026.
