Your standards.
AI-powered evaluation.
You define what good performance looks like. Bluefrog's AI evaluates every conversation and process against that definition — and returns a score, the evidence behind it, an explanation in plain language, the coaching opportunity and the trend over time.
A rubric is your operating standard, written down.
Most companies already know what a good call sounds like. It lives in a training binder, in a manager's head, or in the way one veteran CSR handles a busy Monday. The problem has never been defining quality — it has been applying that definition consistently across thousands of conversations that nobody has time to listen to.
A Bluefrog rubric is a structured list of criteria your team writes: what must happen on the call, what earns full credit, what counts as partial, and how much each criterion is worth. The AI does not invent the standard and does not rank people against an outside benchmark. It applies your standard, the same way, to every interaction — the tenth call of the day and the four hundredth call of the week.
That is the difference between generative AI and Operational AI. One writes something new. The other reads what actually happened inside your business and measures it against the rules you set.
A full CSR rubric, scored.
Ten criteria a home services CSR rubric commonly contains. Yours can have five or fifty — the criteria, the weights and the language are configured to your operation.
Notice what the numbers do here. The call was booked — by outcome alone it looks like a win, and in most operations it would never be reviewed again. The rubric says something more useful: the front of the call was strong, the close was fine, and the middle was thin. Discovery and urgency are where revenue quietly leaks — the second problem that never got asked about, the customer who was never told how soon a technician could actually arrive.
Criteria are yours to weight. If membership conversion matters more than script adherence, weight it that way. If your dispatch rubric cares about promise-time accuracy above all else, that criterion carries the score. Bluefrog configures the rubric around how your business measures itself, then wires the results into the systems where the work happens — your CRM, ServiceTitan, and management reporting.
See how calls become transcription, intent and rubric input →
A number alone is not a review.
Every criterion returns five things. A score you cannot question is a score nobody will act on — so Bluefrog shows its work.
01 · Score
The criterion scored on your scale, with your weighting applied and the composite rolled up per call, per person, per team and per queue.
02 · Evidence
The specific moment in the transcript the score came from — quoted, timestamped and linked back to the recording so a manager can verify it in seconds.
03 · Explanation
Plain language describing why that evidence earned that score against the criterion you wrote. No black box, no unexplained deduction.
04 · Coaching Opportunity
What to do differently next time, phrased as something a supervisor can say in a one-on-one rather than an abstract performance note.
05 · Trend
The same criterion tracked over time, so you can tell a bad call from a bad habit — and confirm whether coaching actually changed the number.
The rule we don't break
AI does not decide who is a good employee. It evaluates defined behaviors against your written standard and hands the judgment to your managers, with the evidence attached.
One evaluation engine. Every role that matters.
Call handling is where most companies start. It is not where rubric analysis ends — any repeatable process with a record can be evaluated against a written standard.
| Rubric | Representative criteria | Evaluated from |
|---|---|---|
| CSR | Greeting, identification, reason for call, empathy, discovery, urgency, service explanation, appointment attempt, booking, closing | Inbound call transcripts |
| Sales | Needs analysis, option presentation, objection handling, financing discussion, ask for the sale, follow-up commitment | Sales calls, in-home visit notes |
| Dispatch | Capacity awareness, technician-to-job fit, promise-time accuracy, customer notification, schedule protection | Dispatch calls, CRM records |
| Customer Service | Issue ownership, resolution path, expectation setting, escalation judgment, recovery language | Calls, chats, tickets |
| Technicians | Arrival communication, diagnosis explanation, option presentation, documentation quality, customer education | Recorded calls, job notes |
| Lead Qualification | Qualifying questions asked, service-area check, intent classification, disposition accuracy | Calls, forms, chat |
| Estimate Follow-Up | Contact attempted, timeliness, value reinforcement, objection resolution, next step scheduled | CRM activity, calls, email |
| Management Processes | Review cadence, coaching documented, escalation handled, exception resolution, reporting completeness | System records, workflow logs |
| Quality Assurance | Sample coverage, scoring consistency, calibration variance, dispute resolution | Evaluation history |
Criteria shown are representative examples. Every rubric is written with the client.
Home services operations → · Customer intelligence → · Custom evaluation systems →
Coaching you can prove worked.
A single scored call is an anecdote. Eight weeks of the same criterion is management information. When a team works on discovery, the discovery line should move — and if it doesn't, the coaching approach is the thing to change, not the person.
Because evaluation runs on the same connected data as the rest of the platform, rubric scores sit next to what they actually produced: booking rate, job value, and revenue outcomes. That is the question worth answering — not "who scored highest," but "which criteria correlate with booked, completed, profitable work?"
Results roll into management dashboards and into the AI business analysis layer, so supervisors get the exceptions that need attention instead of another report to open.
Configured to your standard, not a template.
Bluefrog Intelligence Platform
AI Evaluation is a module of the Bluefrog Intelligence Platform — modular, configurable and integration-ready.
- Rubric builder: criteria, weights, pass conditions, examples
- Per-role, per-team, per-queue and per-campaign scoring
- Evidence-linked review queues and calibration tooling
- Scores written back to CRM and operational systems
Custom AI Systems Integration
If the platform solves 80% of the problem, we customize the remaining 20%. If the technology doesn't exist, we build it.
- Evaluation of processes with no off-the-shelf equivalent
- Custom scoring models, thresholds and alerting logic
- Integration with proprietary systems and internal databases
- Custom review interfaces for QA and operations teams
AI is easy to access. Making it useful is hard.
Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.
Technology development since 1997 · AI integration platforms since 2001