AI-Assisted Quality Assurance

Getting a real picture:
QA process tooling for full statistics

Designing an AI-Assisted customer service quality assurance review system that consolidated a severely tedious 5 tool workflow and boosted output exponentially.
ROLE
Product Designer
Team
1 Designer (Me!)
2 Administrators
4 Developers
2 QA Engineers
Timeline
9 months
Scope
V1: Chat reviews with AI-
assisted form completion
Summary
The company's customer service QA team was responsible for reviewing every CS rep once a month, but the process of doing that review required navigating five separate tools, manually hunting for average-length interactions across disconnected databases, downloading files to re-upload them elsewhere, and assembling feedback emails with custom summaries for every rep they evaluated.

The result was a process so time-consuming that meaningful QA coverage was impossible to achieve. A single review cycle for one rep took an average of 60–75 minutes, most of which was spent on tool-switching and manual file handling rather than actual quality assessment.

I designed a unified, AI-assisted review environment that consolidated the entire workflow into one place. The AI handles the mechanical assessment work. The human handles the judgment calls. V1 designs covering chat-based reviews were at handoff stage when I departed the company.

The result: an estimated 4–5x increase in review capacity that moved the QA team from monthly spot-checks toward genuine continuous performance monitoring.
Increased review capacity an estimated
4 - 5x
Set up the MVP and future versions for seamless
AI incorporation
Streamlined the workflow from across
5 systems → 1
Problem
The customer service quality assurance team had a process to tedious that it was impossible to truly get enough information to know how the customer service team was actually performing.

From a schedule in an Excel sheet updated for each week and month, to looking up agents in a PowerBI app to gather some averages and number relevant to their customer interactions, to the three separate places those interactions were housed, to the clunky and confusing PowerApp form to rate each interaction example pulled, to then downloading the examples and uploading them for a link, to pulling in the automatic emails sent by every submitted form into another outlook email where all the submission results and sharepoint interaction example links were attached and finally sent to the customer service rep.

It was such a convoluted process that each agent was reviewed each month using only 4 examples, one of each kind and an extra of the interaction type they completed the most of. Add to this the minimum requirements that were not always met as chats, calls, and case email chains needed to be a certain length to qualify and the internal case needed to be closed and completed to be reviewed, and you have a recipe for extreme inefficiency.

The result was that depending on the ease of access the review process often took an hour or often two hours per rep, most of it spend navigating tools and not evaluating performance.
Understanding the process
Before designing anything, I spent time observing the QA team's actual workflow end to end, which is the same approach I took on the Returns project. Understanding why the process worked the way it did, including the workarounds that had become necessary, was essential to designing something the team would actually love.
User flow of original review process.
Some things were immediately clear.

The Excel scheduling spreadsheet was load-bearing and a real hindrance. It was the only place anyone could see who was reviewing whom and whether reviews were on track. Any replacement had to account for that function, not just eliminate the tool.

The Power Apps form was universally disliked but the underlying rubric it enforced was genuinely well thought out. The problem wasn't the evaluation, it was everything surrounding it.

The most time-consuming parts of the process came from pulling examples and crafting emails, and had nothing to do with quality assessment. They were pure overhead that the tools imposed.

That framing shaped the entire design direction: keep the rubric, eliminate the overhead, and let the AI handle what it's good at so the human can focus on what only a human can do.
User flow of V1 MVP process doing chats only and not incorporating the schedule or email services.
User flow of ideal final review process completely housed in the singular tool.
Designing under real constraints
This was not a greenfield product. Every design decision had to work within real operational constraints:
Reviewer skepticism about AI — Asking QA agents to trust AI assessments of their colleagues' performance was a significant UX challenge. The design had to make the AI's role transparent and the human's authority unambiguous. Every AI suggestion had to be clearly framed as a suggestion, not a decision.

Rubric complexity — The evaluation rubric covered Critical Thinking, Professionalism, and Resolution across multiple sub-criteria, each with Yes/No responses and conditional dropdown options. The form logic was non-trivial and had to remain fully intact while becoming significantly faster to complete.

Closed case requirement — Interactions could only be reviewed if the associated case was closed and met minimum length thresholds. The filtering system had to enforce these requirements without making them feel like obstacles.
Multi-type interactions — Chats, phone calls, and emails each had different sources, different formats, and different qualifying criteria. V1 focused on chats, but the architecture had to account for all three from the start so that adding types later wouldn't require rebuilding the foundation.

AI training specificity — The AI needed to learn the company's specific QA standards, not generic customer service benchmarks. The feedback loop built into the design, where reviewers could flag incorrect AI assessments, was as much a training mechanism as a UX feature.

New database and custom structure — Since this was a new consolidated tool and all of the interactions needed to be housed or referenced here with the ability to bring in example links and transcripts, the database developers had a lot of restrictions for the design of the tables and filters.
The solution
What was clearly needed was a single place to house everything but since we needed to be agile and walk up small wins, the V1 MVP simply was meant to handle a table of the chats, the new form, and the first stage of AI integration which on release would reliably be able to fill in 4 yes/no radios and choose the appropriate reasoning for those fields in the select input.

To avoid a massive amount of work later, I also drew up the beginning layouts and functions of how it might work when housing all of the different interaction types and the schedule as well as messaging between the QA team and the customer service reps.
The review queue
Upon logging in, a QA agent sees a filterable table of all completed customer service chat interactions. The table has a default filter of the last 30 days as this is the date range that the QA agents are looking for. There are further filters for the customer service specialist's name, the case number, the case status, and a range for amount of messages that the CS specialist sent.

This replaced the manual process of cross-referencing Power BI tables to find appropriately average-length examples. The filtering happens in one place, in seconds.

If the customer left an indication as to whether or not the chat was helpful, that is included as well as sometimes that is also used in determining which interactions to review. Lastly, there is a button at the end of each row to open up the review form for that specific interaction, which speeds up the process immensely.
The review interface
Clicking Review on any interaction opens a two-panel layout: the full interaction transcript on the left, the QA evaluation form on the right.

The form opens pre-populated with identifying information pulled automatically such as the CS rep name, case number, date, interaction type, and link to the example (in this case a chat transcript).

Further down, AI-suggested responses are already filled into the relevant evaluation fields. The AI is trained to assess specific performance criteria across categories including Critical Thinking, Professionalism, and Resolution, suggesting Yes/No responses and pre-selecting dropdown reasoning options based on its reading of the transcript.

The QA agent's job shifts from filling out a form from scratch to filling out some of it, then reviewing and confirming what the AI has already assessed, correcting where it's wrong, confirming where it's right, and adding qualitative notes at the end.
New form for other interaction types
For phone calls and emails not yet covered by the AI pipeline in V1, agents can open a blank form from the main queue, manually enter the identifying information, and paste or link to the interaction example.

Once linked, the transcript or reference populates the left panel in the same layout as chat reviews maintaining a consistent experience across interaction types.

Agents also have the option to manually trigger the AI review on these forms, running the same assessment pipeline on demand.
Customer service representative view
When the customer service reps log in to the tool themselves, they can see all of their recent interactions in a tabbed table, with similar filter structure that persists across the tabs.

In addition, they also get a snapshot of ther stats at the top and I was working on an idea to include automatic filtering based on the stats for a future version of the tool.

Navigating to the Messaging tab at the top, they can see their correspondence with the QA review agents and send messages back. This was not instant messaging but close async like a delayed email.
Design Decisions Worth Calling Out
Two-panel layout as the core UX decision. One of the most important things this QA tool does is keep the transcript and the evaluation form visible simultaneously. In the old process, agents had to context-switch between applications constantly, reviewing an interaction in one window while filling out a form in another. The two-panel layout eliminates that entirely. The agent never loses their place in the transcript while filling out the form.

AI as an assistant, not an authority. Every AI-suggested field is editable and every assessment is visible with its reasoning pre-shown in the dropdown. The design is explicit that the AI is making a suggestion, not a decision. The feedback modal reinforces this as agents are expected to push back when the AI is wrong, and that feedback matters. This was an intentional framing choice to build trust with a QA team that was understandably skeptical about AI evaluating human performance.

Filtering as research replacement.
The filter bar in the queue isn't just a convenience feature, it's directly replacing the most time-consuming part of the old process. Designing it to support message-count and call-duration ranges specifically was a direct response to how QA agents manually hunted for average-length interactions in Power BI. The filter does that work automatically.
Final future version
V1 covered chat reviews with AI assistance. V2 was to include emails from the internal case management system. V3 was to include the phone calls component...
That's all well and nice but the end vision I was working toward included:
Aggregate performance dashboards showing rep trends over time across review cycles
Full AI pipeline for phone call and email review
Automated interaction selection based on statistical averages, removing the need for agents to manually identify representative examples
Scheduling and assignment directly within, replacing the Excel spreadsheet entirely
Results
When I left the company, this was at the v1 handoff stage, so post-launch metrics aren't available. What I can speak to is the scale of the inefficiency it was designed to eliminate and the projected impact based on direct observation of the existing process.
60–75 minutes per review → ~15–20 minutes for chat reviews in V1, based on workflow comparison before and after
Projected 4–5x increase in review capacity — moving from monthly spot-checks toward genuine continuous performance monitoring
5 tools consolidated into 1 — Excel, Power BI, Power Apps, Briefcase, and a third-party call system replaced by a single environment
AI pre-population of evaluation fields eliminating the most repetitive data-entry work from every review
Feedback loop built in from day one — the system gets more accurate over time rather than remaining static
4 - 5x
increased capacity for reviews by combining tools.
AI incorporation
to boost automation and increase ease of the review process.
5 systems → 1
streamlining the workflow and eliminating redundancy.
Built-in stats
so that the QA team, the CS reps, and leadership see performance in one place.
75% faster
review process and even faster as AI learns, making true assessment possible.
What I learned
This tool taught me that trust is a design problem.

The QA team wasn't resistant to a better tool — they were skeptical of an AI making judgments about their colleagues. That skepticism was reasonable and I had to design around it rather than dismiss it. Every decision about how the AI's suggestions were displayed, labeled, and overridable was a decision about whether the system felt like a collaborator or an authority. Getting that framing right mattered as much as the information architecture.

It also reinforced something I'd learned on the Returns project: the most important design work happens before any screens are made. The hours I spent watching the QA team navigate their existing process to understand not just what they did, but why the workarounds were in place, were what made it possible to design something that replaced the right things and kept the things that were actually working.

If I'd been able to see this tool through to launch and into the hands of the QA team, I think the feedback loop feature would have been the most interesting thing to study. Systems that learn from their users in real time are still relatively new in internal tooling, and watching that dynamic play out with a skeptical team would have been genuinely valuable.