Skip to content

Practical guide · Assessment · Criteria

AI agents in production in Brazil: how to assess a platform

What to require from a platform or partner to run AI agents in production in Brazil: owner, verification, evidence, integration, deployment, and LGPD.

Who it is for

Operations, technology, security, and procurement leaders assessing platforms and partners for AI agents

What you take away

A matrix of criteria, questions, and evidence to compare proposals against the same requirements

To run AI agents in production in Brazil, assess the platform and the partner by what they can demonstrate in your process, not by how fluent a demo sounds. Require a named owner, a verification criterion agreed before the pilot, and evidence of the outcome in the target system, not only in the transcript. Check the bounded scope, the handoff to the team with context, the integration with source systems, the deployment option, the data map for the LGPD, and who follows the operation after go-live. This guide is written by Starya and does not rank vendors.

Is there a named owner for the process?

An AI agent carries out steps within a process, but it is not accountable for it. Ask any vendor how the project defines who, on the company side, accepts the scope, defines the verification criterion, receives the exceptions, and can pause the use. A proposal that treats “the team” or “the AI” as the owner leaves those decisions without one.

What to ask for. A table with the process, the owner by name and role, the stand-in, and the owners of permissions and of each target system. The guide on the owner of an AI agent in production details what this person decides and how to name them before the pilot.

Is the verification criterion agreed before the pilot?

Before measuring any result, the company and the vendor need to write down what counts as a completed task and where that is checked. Without that criterion, each area reads the result its own way, and the pilot assessment becomes a debate of opinions.

What to ask for. The written criterion for each flow in scope, with the source of the check, the person who approves it, and what happens when the result falls short. An indicator without a criterion agreed in advance does not allow proposals to be compared.

Is the evidence in the system, and not only in the transcript?

A “done” in the conversation is a report. Completion shows up in the target system: the record created, the lookup returned, the delivery confirmed. The transcript shows what was said; it does not prove that the action was carried out or that the result is correct.

What to ask for. An example record that shows, for the same request, the agent’s response, the action carried out, and the confirmation in the target system. Also ask how the follow-up separates completion, an attempt without confirmation, and a handoff.

Is the scope bounded, and does the handoff carry context?

An operation in production starts with requests that have a clear beginning and end, a defined data source, and written limits. Whatever is out of scope needs a known destination. In healthcare, for example, looking up the status of an authorization request is not deciding whether to authorize care, and that decision stays with the responsible team.

What to ask for. The list of included and excluded requests, the handoff conditions, and what reaches the person who takes over: identification, request, and what the agent has already tried or looked up. Without that context, the team restarts the conversation and the reason for the exception is lost.

Does the integration reach the source systems?

An agent without integration recognizes the request but cannot complete it. The information it delivers should come from the source system, through the interface defined in the project, and not be generated by the model. When the lookup fails, the correct response is to say it could not be confirmed and hand off with context.

What to ask for. The map of systems the agent reads and changes, the permissions for each integration, who approves them, and who can revoke them. Also ask how control reaches the point that carries out the action: observing the call to the model does not prove the authorization, execution, or result of the next call to a system.

Where does the operation run, and what changes for the LGPD?

Ask which deployment options exist, who operates each component, and where data and model calls flow. At Starya, the options are Shared SaaS and Dedicated. The purchasing channel is a separate decision from the runtime environment: buying through a marketplace does not define where the operation runs.

No deployment option resolves LGPD questions on its own. In any design, we recommend mapping which personal data the agent processes, for what purpose, through which systems, models, and vendors it passes, who accesses it, and how long it is retained, and bringing that map to the company’s legal and privacy review. The LGPD section of the Shared SaaS or Dedicated guide organizes this map.

Who follows the operation after go-live?

Going into production is the start of the operation, not the end of the project. Changes in source, permission, model, or volume can alter the agent’s behavior. Someone needs to follow completions, exceptions, and handoff reasons, and decide when to adjust, expand, or pause.

What to ask for. How the vendor takes part in the follow-up, which dashboards or records are available, how often results are reviewed together, and how new fronts are assessed. Timelines, coverage, and support conditions need to be agreed; they cannot be inferred from the name of the offer.

Portuguese and operations in Brazil: what to check

Ask to see the agent with the vocabulary, documents, and channels of your operation, in Portuguese, and not only in a demo script. Ask which operations in Brazil the vendor publishes as case studies, with what scope and with what limits.

The case studies Starya publishes include Unicall, in the Sistema Unimed Paraná, with voice service in the IVR, and dr.consulta, with patient relationship over WhatsApp. Each case describes its own scope, and its results are not a guarantee for other contexts.

Questions to bring to any vendor

Use the same questions for every proposal and record the answer and the evidence shown. An answer without evidence remains outstanding. The worksheet at the end of this page lets you fill in and download these questions with your team.

Questions to bring to any vendor of AI agents in production
CriterionQuestionVerifiable answer
OwnerWho, on the company side, is accountable for the process and can pause the use?Table with name, role, stand-in, and owners of permissions and systems.
Verification criterionWhat counts as a completed task and where is that checked?Written criterion for each flow, approved before the pilot.
EvidenceHow do I see that the action was carried out, beyond the transcript?Record linking response, action, and confirmation in the target system.
Scope and handoffWhat is out of scope, and what reaches the person who takes over?List of included and excluded requests and an example of a handoff with context.
IntegrationWhich systems does the agent read and change, and with which permissions?Integration map with who approves and who can revoke.
Deployment and LGPDWhere does the operation run and where does the data flow?Deployment design with routes, vendors, and owners for the legal and privacy review.
Follow-upWho follows results and exceptions after go-live?Review routine, available records, and agreed support conditions.

How Starya meets these criteria

Owner and criterion. Starya starts with a bounded process and asks the company to name who is accountable for it. Before starting, the company chooses the process, identifies the data and systems involved, and agrees on what will count as a completed outcome. The guide on the owner of an AI agent in production and the article on response and outcome detail this preparation.

Evidence and control. NebulaOS connects models, sources, and tools and helps monitor the points included in the deployment. To control an action, the design also integrates the point that carries it out. The guide for CISOs covers the audit trail and the evidence that allows the operation to be released or resumed.

Scope, integration, and handoff. At Unicall, in the Sistema Unimed Paraná, the agent Júlia was integrated into the IVR to identify the member and handle bounded flows, such as looking up the status of an authorization request and sending a duplicate payment slip. The status provided comes from the source system through the API, and the call is handed off with the identification and the information collected when the lookup fails or the request is out of scope. Average handling time went from 8 to 2 minutes in the automated workflows, and in the first month of operation 12% of calls were fully resolved by the AI. The indicators belong to those workflows, not to Unicall’s total call volume or to Starya’s other operations. The guide on an AI-powered IVR that identifies the member and looks up the authorization status without hallucinating applies these criteria to a health plan operator.

Follow-up after go-live. At dr.consulta, Gabi supports patient relationship over WhatsApp in the Care Pathways, and more complex situations go to the human team. The evolution was followed by the dr.consulta and Starya teams, and the analysis of the interactions revealed the demand that gave rise to Lara, focused on scheduling.

Deployment. Shared SaaS and Dedicated are the two options. Dedicated can use OCI, AWS, or on-premises infrastructure, subject to technical assessment and the agreed scope, and the purchasing channel is defined separately from the runtime environment. The deployment option does not make the operation LGPD compliant automatically; the data map goes to the company’s legal and privacy review.

About this page

This guide is written by Starya. It organizes criteria to assess any platform or partner for AI agents in production and does not compare, rank, or name other vendors. Where Starya appears, the facts are those published on its own case, product, and guide pages, with the limits described on each one.

Illustrative example · no customer data

Two proposals for the same status lookup

Fictitious example: a company receives two proposals for an agent that answers order status lookups. Both demos sound fluent. The team decides to compare the proposals against the same criteria, recording the evidence shown for each one, and not by the impression the demo left.

The matrix below shows how the record would look. The proposals are fictitious and do not represent real vendors or a Starya offer.

Assessment matrix for two proposals · fictitious example
CriterionProposal AProposal BOutstanding item
OwnerAsks for an owner and a stand-in to be namedDoes not mention who is accountable for the processProposal B needs to state how the owner will be defined
Verification criterionCriterion written before the pilotCriterion defined after the pilotProposal B needs to agree on the criterion before measuring
EvidenceShows response, action, and confirmation in the order systemShows only the conversation transcriptProposal B needs to demonstrate confirmation in the target system
Scope and handoffLists excluded requests and the handoff contextHands off without defined contextProposal B needs to describe what reaches the team
IntegrationRead access to the order system, with approved permissionsIntegration to be definedBoth need security approval
Deployment and dataCandidate deployment option with data routes describedEnvironment named without data routesBoth need legal and privacy review
Follow-upJoint review of results and exceptionsOn-demand supportBoth need agreed support conditions

Fill in one column for each proposal with the evidence shown, not with the promise. A cell without evidence is an outstanding item. In the example, neither proposal is approved while integration, data, and support have not been verified.

When the approach needs to change

If no proposal meets a mandatory criterion, narrow the scope or keep the decision pending until the criterion is demonstrated. If the company does not yet have an owner for the process, resolve that before comparing vendors. If the data requirement is not clear, start with the data map and the privacy review, and only then choose the deployment option.

A point to watch: Comparing proposals by how fluent a demo sounds, without checking evidence in the target system or who is accountable after go-live.

The worksheet is ready to move forward when…

  • Each proposal was assessed against the same criteria and the same questions.
  • Owner, verification criterion, and evidence in the target system are recorded for the scope.
  • Scope, handoff, integrations, and permissions have owners and written conditions.
  • Deployment option, data map, and follow-up after go-live were brought to the technical, legal, and privacy review.

Bring the matrix to the initial assessment and compare it with the owner guide and the guide for CISOs. Confirm the outstanding items before choosing the platform.

Your worksheet

Complete it with your team.

Record what you know and what still needs confirmation. Answers stay in memory and are not submitted.

Standalone HTML file with the full guide, your answers and an option to print / save as PDF. It does not inspect systems or verify controls.

References for further reading

Further reading

Continue exploring this context

Guide

Who is the owner of an AI agent in production?

The owner of an AI agent is a named person, not “the AI”. See what they decide, when they take over, and how to name them before the pilot.

GovernanceOperationsActions and evidence
Published

Article

Your AI answered “done”. But was the task completed?

“Done” in the chat does not prove completion. The difference between the agent's response, the execution of the action, and confirmation in the target system, with an owner and evidence.

Actions and evidence
Published

Guide · CISO

What a CISO asks before putting Applied AI into production

How to audit what an AI agent did: connect scope, data, RBAC/ABAC, the audit trail, and recovery in a matrix of controls, evidence, owners, and outstanding items.

SecurityGovernanceActions and evidenceContext and data
Published

Guide

Shared SaaS or Dedicated: how to choose a deployment

Compare operations, isolation, networking, data, models, changes, and support. Record requirements, personal data, conditions, and owners, and bring the map to the LGPD assessment.

DeploymentGovernance
Published

Continuity

Applied AI at Starya

Connect this worksheet to the products, deployment conditions and assessment of your operation.