The New Role of CX QA in 2026: Monitoring the Humans, the AI, and the Journey Between Them
As AI and human agents increasingly share responsibility for customer interactions, CX QA must evolve from evaluating individual agents to monitoring quality across the entire customer journey.
.png)
For decades, contact center quality assurance was primarily designed to evaluate people. A customer interacted with a human agent, and a QA evaluator reviewed that interaction to determine whether the agent followed the correct process, communicated effectively, complied with relevant policies, and ultimately helped the customer.
That model made sense when the human agent was the primary owner of the interaction. But in 2026, many customer journeys no longer belong to a single agent, or even to a single type of agent.
A customer may begin a conversation with an AI agent, move through an automated workflow, escalate to a human agent, and then trigger additional automated actions after the conversation ends. Along the way, information may pass between systems, channels, teams, and technologies.
From the customer’s perspective, however, these are not separate experiences. They are one journey.
This creates a fundamental challenge for CX QA teams. If quality assurance continues to evaluate only individual human interactions, it can provide a detailed view of one part of the customer experience while missing the systems and transitions that increasingly determine whether that experience succeeds or fails.
The role of CX QA is therefore expanding. It still needs to evaluate human performance, but it must also monitor AI performance and, critically, the journey between the two.
Contact Centers Did Not Just Add AI. They Changed How Customer Journeys Work.
The introduction of AI into customer service is sometimes discussed as though contact centers simply added another channel or productivity tool. In practice, the change is much more significant.
AI is increasingly participating directly in customer interactions. It can interpret intent, answer questions, retrieve information, recommend actions, route customers, trigger workflows, update records, and determine when an interaction should be escalated to a human.
As a result, a single customer journey can involve multiple owners.
Consider a customer contacting an insurance company about a claim. An AI agent may identify the customer, determine the reason for contact, retrieve policy information, and attempt to answer an initial question. If the issue requires judgment or an exception, the AI may escalate the case to a human agent. The human agent then reviews the information, makes a decision, and triggers another workflow that sends a confirmation or initiates a follow-up action.
The quality of that experience depends on much more than the performance of the human agent who eventually handles the case. It also depends on whether the AI understood the customer correctly, whether the information it provided was accurate, whether it escalated at the appropriate time, whether the relevant context transferred successfully, and whether the systems involved completed the correct next steps.
Traditional QA was not designed around this type of journey. It was designed around an interaction that could be isolated, reviewed, scored, and attributed to an individual agent.
In an AI-enabled contact center, that unit of analysis is often no longer enough.
Why Traditional CX QA Creates Blind Spots
Traditional contact center QA typically focuses on individual interactions and human agent performance. Evaluators review a sample of calls, chats, emails, or cases using a scorecard that measures criteria such as communication, process adherence, compliance, and resolution.
This approach remains valuable. Human performance still matters enormously, particularly as agents increasingly handle complex, sensitive, or high-value interactions that automation cannot resolve.
The problem is not that traditional QA measures the wrong things. The problem is that it measures only part of the experience.
When an interaction moves between AI, automation, and human support, a poor customer outcome may have multiple contributing factors. An agent might receive an incorrectly routed case. An AI agent might provide inaccurate information that a human later has to correct. A customer may have to repeat information because context was lost during an escalation. An automated workflow may fail after both the AI and human agent performed correctly.
If QA looks only at the human interaction, these failures can be difficult to identify.
Worse, the wrong cause may be assigned to the problem. An agent’s handle time might increase because they have to reconstruct information that was lost during an AI-to-human handoff. A human agent might receive lower resolution scores because the customer was routed to the wrong team. Repeated coaching might be assigned for an issue that is actually caused by a broken workflow or poor AI decision logic.
Modern CX QA needs enough context to distinguish between an agent performance problem and a system performance problem.
The Challenge of Fragmented Visibility
As contact centers introduce more AI and automation, the data needed to understand quality is often spread across multiple systems. AI performance may be monitored in one platform, human agent evaluations in another, customer satisfaction data somewhere else, and workflow or case outcomes in yet another system.
Each source can provide useful information, but looking at them independently makes it difficult to understand what actually happened across the complete customer journey.
A team might know that an AI agent successfully identified a customer’s intent and escalated the interaction. The QA team might also determine that the human agent who received the escalation performed well. Yet neither evaluation necessarily reveals whether the transition between the two created a poor experience for the customer.
For example, imagine an AI agent that collects a customer’s account information, identifies the reason for contact, and correctly determines that the issue requires a specialist. The interaction is then transferred to a human agent, but the information already gathered by the AI does not transfer with it. The customer has to provide their account details again and repeat the entire reason for contacting support.
Viewed separately, both parts of the interaction may appear successful. The AI correctly identified the issue and initiated an escalation, while the human agent communicated effectively and resolved the problem. However, the customer still experienced unnecessary repetition and effort because the handoff itself failed.
This is one of the biggest challenges for modern CX QA. When AI performance, human performance, and operational outcomes are evaluated separately, the gaps between them can become blind spots. Organizations need visibility not only into how each component performs, but also into how those components work together to create the overall customer experience.
The New Role of CX QA
In 2026, CX QA is becoming a system-level quality function.
This does not mean abandoning traditional agent evaluation. Instead, it means expanding the scope of QA to reflect how customer service is actually delivered.
A modern quality program needs to monitor three connected layers: the performance of human agents, the performance of AI agents and automated systems, and the quality of the journey between them.
Each layer answers a different question. Human QA asks whether agents are communicating effectively, making good decisions, following required processes, and resolving customer issues. AI QA asks whether automated systems are understanding customers correctly, providing accurate responses, making appropriate decisions, and escalating when necessary. Journey-level QA asks whether all of these components work together without creating unnecessary effort, inconsistency, or risk.
Looking at only one or two of these layers can produce an incomplete picture of quality.
1. Human Performance Remains Critical
The rise of AI does not make human QA less important. In many contact centers, it may make the interactions handled by humans more difficult.
As AI takes on routine questions and straightforward transactions, human agents are increasingly likely to receive the interactions that require judgment, empathy, negotiation, exception handling, or complex problem-solving. They may also inherit customers who are already frustrated because an automated interaction failed.
That changes what good human performance looks like.
Traditional measures such as communication quality, process adherence, compliance, and resolution remain important. However, QA teams may also need to place greater emphasis on an agent’s ability to understand complex context, make sound decisions, and recover an experience that has already gone wrong.
For example, an agent who receives an escalation from an AI agent may need to quickly understand what the customer has already attempted, recognize any incorrect information previously provided, acknowledge the customer’s frustration, and move the interaction toward resolution without creating additional effort.
Evaluating the human agent without visibility into what happened before the handoff makes it difficult to assess that performance fairly. An interaction that appears inefficient in isolation may actually represent excellent recovery from a poor automated experience.
Modern human QA therefore needs context. Evaluators should understand not only what the agent did, but also the journey the customer took before reaching them.
2. AI Performance Requires Its Own Quality Framework
AI is now part of the frontline customer experience, which means it requires systematic quality assurance.
Many organizations monitor AI using operational metrics such as containment, deflection, response time, or task completion. These metrics can be useful, but they do not necessarily indicate whether the customer received a high-quality experience.
An AI agent can successfully contain an interaction while providing an incomplete answer. It can respond quickly but misunderstand the customer’s intent. It can avoid escalation even when human intervention would have produced a better outcome.
For this reason, AI quality needs to be evaluated across multiple dimensions.
Intent recognition measures whether the AI correctly understood what the customer wanted. A mistake at this stage can affect everything that follows, from the information provided to the workflow or team selected.
Response accuracy evaluates whether the information given to the customer was correct, relevant, and grounded in approved sources. This is particularly important for organizations operating in regulated industries or dealing with complex products and policies.
Resolution effectiveness asks whether the AI actually solved the customer’s problem. This is different from simply ending the interaction without human involvement.
Decision quality becomes increasingly important as AI agents gain the ability to recommend actions, update records, trigger workflows, or make decisions within defined parameters.
Compliance adherence ensures that AI interactions follow required policies, disclosures, and procedures consistently.
Finally, escalation behavior evaluates whether the AI recognizes when it should hand the interaction to a human. A successful AI agent is not necessarily one that minimizes escalation at all costs. In many situations, the best outcome is an accurate and timely handoff.
The goal of AI QA should therefore be to determine whether the AI contributed to a correct, compliant, and effective customer outcome, not simply whether it reduced human involvement.
3. The Journey Between AI and Humans Is Now a Critical Part of Quality
The transition between AI and human support is one of the most important areas for modern CX QA because it is also one of the easiest to overlook.
Organizations often evaluate AI performance and human performance separately. The AI team may monitor automation metrics while the QA team scores human interactions. Yet many customer experience failures occur at the point where responsibility moves from one to the other.
A handoff can fail in several ways. The AI may escalate too early or too late. The customer may be routed to the wrong team. Information collected during the automated interaction may not transfer to the human agent. The agent may be unable to see what the customer has already tried. The customer may have to repeat information or restart the entire conversation.
None of these failures can be fully understood by evaluating the AI or human interaction in isolation.
Modern QA therefore needs to measure the quality of the transition itself. That includes whether the escalation happened for the right reason, whether the customer reached the correct destination, whether relevant context was preserved, whether the customer had to repeat information, and whether the next agent could continue the interaction without unnecessary friction.
This is where journey-level QA becomes particularly valuable. It allows organizations to identify failures that sit between traditional ownership boundaries.
A Practical Example: When the Agent Is Not the Problem
Imagine a customer contacting a telecommunications provider because their internet service is not working.
An AI agent incorrectly classifies the problem as a billing issue and routes the customer to a billing specialist. The billing agent recognizes the mistake and transfers the customer to technical support, where another agent eventually resolves the issue.
A traditional QA process might review the technical support interaction and conclude that the agent performed well. It might also review the billing agent and find that they followed the correct transfer process.
Both evaluations could be accurate, but neither would explain why the customer experienced unnecessary transfers, a longer resolution time, and greater effort.
The root cause occurred earlier in the journey when the AI misclassified the customer’s intent.
This distinction matters because the appropriate action depends on the cause. Coaching the human agents would not solve the problem. The organization would need to investigate the AI’s intent recognition or routing logic.
System-level QA changes the central question from, “Did the agent perform correctly?” to, “Where did the customer journey succeed or fail, and why?”
That is a much more useful question for improving the customer experience.
Why Broader QA Coverage Matters
Traditional QA has historically relied on sampling because manual evaluation is resource-intensive. A QA team may review only a small percentage of total customer interactions and use those results to identify trends in agent performance.
Sampling can still play a role in quality management, but it becomes more limiting as customer journeys become more complex.
A small sample of human interactions may not reveal a recurring problem with AI routing. Randomly selected calls may miss a compliance issue that appears only in a specific type of automated interaction. A handoff problem affecting a relatively small percentage of customers may still create significant risk if the contact center handles millions of interactions.
AI-powered QA makes it possible to monitor a much broader volume of customer interactions and identify patterns that would be difficult to find through manual sampling alone.
This does not mean every interaction needs to receive the same full evaluation or that human QA teams become unnecessary. Instead, automation can help monitor large volumes of interactions for defined signals, such as compliance risks, negative sentiment, escalation failures, unusual patterns, or specific customer experience problems.
Human evaluators can then focus their expertise where judgment is most valuable: investigating complex cases, validating findings, understanding root causes, and determining what action should follow.
The objective is not simply to produce more scores. It is to increase visibility into the issues that matter.
Modern CX QA Requires Connected Metrics
As the scope of QA expands, the measurement framework must expand with it.
Human quality metrics may include communication, compliance, resolution quality, process adherence, and coaching trends. AI quality metrics may include intent recognition, response accuracy, decision quality, appropriate escalation, and compliance. Journey metrics may include handoff success, context transfer, customer repetition, number of transfers, and time to resolution across the full experience.
These measures become significantly more useful when they can be connected to customer and business outcomes.
For example, an organization might identify that customer satisfaction declines sharply when an interaction moves from AI to a human agent. That observation is useful, but the next question is why.
Is the AI escalating too late? Is it routing customers incorrectly? Is context being lost during the transfer? Are human agents receiving incomplete information? Are customers repeating themselves?
Connected QA data can help move the organization from identifying a symptom to understanding the underlying cause.
This is the difference between reporting on quality and using quality data to improve operations.
Salesforce Can Provide the Context Needed for End-to-End QA
For organizations running customer service in Salesforce, much of the context required to understand the full customer journey may already exist within the Salesforce environment.
Depending on the organization’s architecture, Salesforce can connect customer records, case data, human agent activity, AI interactions, workflow activity, escalations, and business outcomes. When QA operates within this environment, quality data can be evaluated alongside the operational context in which the interaction occurred.
This makes it possible to connect an evaluation to more than an isolated call or chat. Teams can understand the customer involved, the case history, the AI activity that preceded a human interaction, the workflows that were triggered, and the eventual outcome.
It also makes transitions easier to analyze. Rather than treating AI interactions and human interactions as unrelated datasets, organizations can evaluate them as connected parts of the same customer journey.
For QA teams, this creates a more complete view of what happened. For operations teams, it makes quality insights more actionable because findings can be connected to the systems and processes that need to change.
CX QA Is Moving From Monitoring to Continuous Improvement
The broader role of CX QA is not simply to monitor more interactions. It is to help organizations understand where customer journeys break and what needs to improve.
A QA program may identify that a particular AI workflow frequently results in customers repeating information after escalation. The solution may not involve coaching human agents at all. Instead, the organization may need to improve how context is transferred during the handoff.
Similarly, if human agents repeatedly correct the same inaccurate AI response, QA data can reveal a problem with the AI’s knowledge source, instructions, or decision logic.
If customers consistently require multiple transfers before reaching the correct team, the issue may lie in routing rather than agent performance.
In each case, QA provides the evidence needed to identify a pattern, investigate the cause, and direct improvement efforts toward the right part of the system.
This represents an important evolution in the role of quality assurance. QA is no longer only a retrospective function that tells organizations how agents performed last month. It can become an operational feedback loop that helps improve humans, AI, workflows, and the customer journey continuously.
What Happens When CX QA Remains Focused Only on Humans?
Organizations that continue to evaluate only human agents risk developing an incomplete view of customer experience quality.
AI errors may go unmeasured. Handoff failures may remain invisible. System problems may be attributed to individual agents, resulting in coaching that does not address the actual cause. AI performance may be optimized around metrics such as containment without sufficient visibility into whether customers received accurate answers or successful outcomes.
Perhaps most importantly, leadership can develop false confidence.
Human agent QA scores may improve while customers continue to experience broken journeys caused by routing errors, lost context, poor escalation decisions, or unreliable automation. If those parts of the experience sit outside the QA program, the organization may not see the problem clearly enough to fix it.
The purpose of expanding CX QA is not to create more complexity. It is to ensure that quality measurement reflects how customer service is actually being delivered.
How to Start Building a System-Level CX QA Program
Moving from agent-level QA to system-level QA does not require rebuilding the entire quality program at once. A practical starting point is to map how customers currently move through the contact center and identify where quality is already being measured and where blind spots remain.
Start by identifying every point where AI participates in the customer journey. This includes not only customer-facing AI agents, but also systems that classify intent, recommend actions, route interactions, summarize conversations, or trigger workflows.
Next, examine the transitions between AI and human support. What information is transferred? Can the human agent see what the customer has already said or attempted? How often do customers repeat themselves? Are escalations reaching the right team?
QA teams should then compare these journeys with their current evaluation framework. If the program measures only the human portion of the interaction, which decisions and experiences are currently invisible?
Finally, organizations should consider how quality findings lead to action. A human performance issue may require coaching. An AI accuracy problem may require changes to knowledge, instructions, or logic. A handoff failure may require a workflow or integration change.
The value of system-level QA comes from being able to identify not only that something went wrong, but where it went wrong and what should change as a result.
The Bottom Line
CX QA was built for a world in which human agents owned most customer interactions. That is no longer the reality for many contact centers.
In 2026, customer experiences are increasingly created by a combination of humans, AI agents, automated workflows, and the transitions between them. Quality assurance must evolve to reflect that reality.
Human performance still matters. AI performance now matters too. But neither can be fully understood without visibility into how the entire journey works.
The future of CX QA is therefore not simply about scoring more interactions. It is about creating a connected view of quality across the systems responsible for delivering the customer experience.
When organizations can see how humans perform, how AI performs, and what happens when customers move between them, QA becomes more than a scorecard. It becomes a way to identify root causes, improve operations, reduce risk, and continuously strengthen the customer journey.
Because customers do not experience your AI, your human agents, and your workflows as separate systems.
They experience one customer journey.
CX QA needs to measure it that way.

