An operation with agents should be monitored through completed work and exceptions that need follow-through. Define what counts as a result, who takes on each outstanding item, and how the comparison will be made. Messages, calls, and responses are activities; the effect on the process needs its own confirmation and measure.
What should count as completion?
Choose a task recognized by the team and describe its input and output. For a quotation, the output may be a proposal ready for review; for an appointment, it may be confirmation in the system. These results are not equivalent and should not use the same criterion. Identify the record that allows completion to be checked and who handles discrepancies.
Consider partial completion and an unknown result. A request with three items may have two confirmed and one pending. Ending the conversation does not eliminate that outstanding item. The team should be able to recognize the remaining work and resume from there, without repeating completed parts or announcing as finished what has not yet been confirmed.
Which metrics help with decisions?
Start with a few measures related to the objective: time to completion, proportion of confirmed tasks, human interventions by reason, and rework. Write down how each measure will be calculated, which situations are included, and which period will be observed. If there is no baseline, measuring the current process is part of the first delivery.
Technical costs need to be related to completed work. A reduction in model consumption may accompany an increase in review effort. Compare the full path and the conditions of use, including changes in volume, request type, or team training. Time saved does not automatically become an expense reduction; interpretation depends on how that capacity is used.
How do you get an exception to someone who can resolve it?
Define a destination for requests outside the scope, missing information, rules that require approval, and integration failures. The handoff should carry the reason, necessary context, what has already been tried, and the expected next step. A visible queue without an owner or a follow-up condition may simply move the problem elsewhere.
Observe whether the person can take over without reconstructing the entire conversation. Confirm which decisions they can make and where they record the resolution. When no one is available, the operation needs to maintain an understandable waiting state and service expectations consistent with the actual service. Timeframes should be agreed with whoever is responsible for the team.
How do you turn monitoring into improvement?
Separate incident handling from pattern investigation. The first maintains continuity for a request; the second seeks a hypothesis that explains recurring problems. Describe the evidence, choose a change with a defined scope, identify the owner, and agree on how to observe the effect. Changing a prompt without examining the source, process, and integration may leave the cause intact.
Before expanding, check whether the change improved the team’s work and whether it created effort at another stage. Record the decision to maintain, adjust, or return to the previous version. In Starya’s work, engineering participates in this cycle alongside those running the operation; the products support work and monitoring according to the context and enabled functions.
Illustrative example · no customer data
The queue received the service interaction but lost the reason for contact
Fictional example: a team needs to repeat questions after receiving conversations from the agent. The investigation finds the deadline in the original request, but it does not appear in the queue’s context. The hypothesis is that the incomplete handoff increases repeated information collection.
The change includes the deadline and the reason for referral in the team’s incoming request. The test observes whether these fields arrive correctly and are usable. It then tracks comparable situations to decide whether there was an improvement. The hypothesis does not become a result simply because the screen was changed.
| Measure | What to check | Limit of interpretation |
|---|---|---|
| Handoff with usable context | The agreed fields arrive and help the person continue. | A completed field does not prove that the data is correct. |
| Repeated questions | Did the person need to collect something already provided again? | Confirming a piece of information may be a necessary step. |
| Time to resolution | The same start, the same end, and comparable requests. | Changes in demand or team affect the comparison. |
| Outstanding items by reason | Responsible destination and state of each exception. | Quantity alone does not indicate severity or effort. |
The team needs to agree on the definitions before comparing. The metrics are a proposal for this scenario, with no values or gains attributed to customers.
When the approach needs to change
If the operation does not yet have enough data, start with a monitored sample and common criteria across evaluators. If the team cannot handle the exceptions, reduce the scope or review capacity before increasing volume. Automating one stage requires monitoring the effect on subsequent stages.
A point to watch: Using message volume, fewer calls, or the end of a conversation as an automatic substitute for results.
The worksheet is ready to move forward when…
- Completion and outstanding items are recognizable to the team and verifiable in the target system.
- Metrics have a definition, scope, and baseline or measurement plan.
- Each exception and improvement has an owner and a next step.
Choose a recent incident and use the roadmap to discuss an improvement with operations and engineering.
Your worksheet
Complete it with your team.
Record what you know and what still needs confirmation. Answers stay in memory and are not submitted.
References for further reading
- NIST AI RMF Playbook ↗
A voluntary reference for organizing governance, understanding context, and measuring and managing AI risks. It is not a Starya certification and does not replace assessment of the specific case.