Review every ElevenLabs voice agent call with a QA scorecard

By General Input

Open one queue each morning, score every voice agent call against your rubric, and let an assistant pre-fill each scorecard before you approve it.

Integrations

  • ElevenLabs
  • Google Sheets
  • Slack Bot
  • HubSpot

Type

App

Categories

  • Customer Support
  • Operations

Build me an app my support QA lead opens every morning to review the calls our ElevenLabs voice agents handled. The whole point is that a reviewer works through a queue of calls, scores each one against a fixed rubric, and signs off, with an assistant doing the first pass so nobody starts from a blank form.

The main view is a review queue. Use the ElevenLabs List Conversational AI Agents operation to get our agents, then List Conversations to pull recent calls across them. Each row shows the agent name, the call duration, the caller, and whether the call has been reviewed yet. Hide any call shorter than thirty seconds so accidental hangups do not clog the queue, and make that cutoff a setting rather than a hardcoded number. Let the reviewer filter by agent and by reviewed or unreviewed.

Selecting a row opens a detail pane. Fetch the full transcript with Get Conversation and show it turn by turn beside an inline audio player fed by Get Conversation Audio, so the reviewer can read and listen to the same moment. Keep the transcript and the player side by side rather than stacked, and make transcript lines clickable references the scorecard can quote.

Every call gets a scorecard against a fixed five line rubric: greeting, identity check, required disclosure, question actually answered, and clear next step. Each line takes a pass or fail plus a free text note. The scorecard also has a space for the drafted follow up wording. Store scorecards in the app's own storage so they persist across sessions and survive the original call being redacted upstream.

Add a "Review this call" button that kicks off a background agent. The agent reads the transcript for that conversation, pre-fills the scorecard with a suggested pass or fail on every rubric line, and quotes the exact transcript lines as evidence for each verdict. It should explicitly flag missed disclosures and questions the agent never actually answered, and draft the follow up wording. Its output lands back in the app as a draft scorecard the reviewer edits and approves. Nothing the agent produces is final on its own. Also add a "Sweep the backlog" button that runs the same agent across every unreviewed call currently in the queue, so a reviewer can arrive to pre-filled scorecards.

Approving a scorecard is the only thing that sends data anywhere. On approve, append a row to our Google Sheets QA log using Append Values, with the call id, agent name, caller, date, each rubric line's verdict and note, the overall pass or fail, and the reviewer's name. If the call failed, post it into a Slack review channel with the Slack Bot Send a Message operation, including which rubric lines failed and the supporting quotes. Then look the caller up in HubSpot with Search Contacts, and on the matching contact write a summary with Create Note plus a Create Task carrying the follow up wording and a due date. If Search Contacts finds no match, say so in the app and still write the spreadsheet row rather than failing the whole save.

Make it multi reviewer. Each reviewer sees their own queue of calls nobody has scored yet, and the name of whoever approves a scorecard is stamped on it and carried into the spreadsheet row, the Slack post, and the CRM note. A call another reviewer already signed off should drop out of my unreviewed queue.

Finally, add a trends tab showing pass rate per agent per week, built from the stored scorecards, so coaching conversations have history behind them. Show the per rubric line failure breakdown too, so it is obvious whether an agent keeps missing the disclosure specifically or is weak everywhere.

One important nuance: ElevenLabs conversation history can be redacted or retention limited depending on plan settings, so treat the Google Sheets log and the app's stored scorecards as the durable record of scores rather than assuming transcripts stay fetchable forever. Never make the trends tab depend on re-fetching old conversations.

Related prompts

Explore more prompts
Call overdue Xero customers with an AI collections agentLocal listing health board for every location you manageLet support send one-off Loops emails without an engineerA brand asset library your marketing team actually searchesTurn Mailjet email clicks into ranked HubSpot follow-upsClean out the Looker dashboards and Looks nobody opensStop cold emails to anyone with a live deal in PipedriveLiveKit live operations console for room moderationWake up dormant Keap leads with a researched reasoniMessage campaign console with pre-flight checks and delivery board