AI receptionist mistakes: what our voice agents got wrong in 3,358 graded calls

AI receptionist mistakes, counted: we graded 3,358 calls on two live voice agents. Six emergencies went down as routine and the grader caught none of them.

Published: 2026-09-23 · Author: Ahmed Heshmat · 10 min read

In short: We counted the AI receptionist mistakes on every call our two voice agents answered at one Toronto property management and brokerage operation between 1 June and 10 September 2026: 3,358 graded calls. At least one problem was flagged on 35.4% of property management calls and 17.1% of brokerage calls, and 707 of the 967 flagged calls were low severity. The costly failure was the caller who wanted a person and could not reach one live: 243 calls on the property management line. Nineteen calls carried real risk, six of them emergencies taken down as routine messages, and the automated grader's own fields flagged none of those six.

Key takeaways

  • A flag that changed the caller's outcome (medium or high severity) was on 9.8% of property management calls and 4.1% of brokerage calls.
  • The largest group of flags is about how the agent talks: script order, stacked questions, filler, talking over the caller, repeating itself. That was 400 flagged calls across both lines, all but three of them low severity.
  • The biggest single category is the caller who asked for a live person and did not get one: 243 property management calls, 11.3% of the line, 98 of them needing a call back or a chase.
  • Three of the five property management emergencies logged as routine were items on the operation's own written emergency list. The other two were a missing smoke alarm and a faulty fire alarm, which the list does not cover and should.

Why we are publishing the misses

When we asked Perplexity about the firm in September, it described our proof as "mostly self-reported results". That was fair. Like every vendor in this category, we had published mostly what the agents got right.

We built both agents and we run them, so every miss below is ours. What they are built to do is on our voice agents page. This is the other column: what kind of failure to expect, how often, and which failures carry risk.

How every call gets graded

A separate AI grader reads every call and writes a report: what went well, what didn't go well, and any caller questions left unanswered. We read every report in the window and sorted each problem by type and severity. Low is style or script order, with no effect on the caller's outcome. Medium means a person had to call back or chase. High is a real risk: an emergency delayed or missed, or wrong information that matters. Where a report showed several problems, we counted the one that did the most to the caller.

The property management line: 2,147 graded calls

At least one flag on 760 calls, 35.4%: 550 low, 192 medium, 18 high. So 210 calls, 9.8% of the line, were medium or high.

| Flag type | Calls | Share of calls | Low / medium / high |

|---|---|---|---|

| Caller wanted a person and could not reach one live | 243 | 11.3% | 143 / 98 / 2 |

| Script order, stacked questions, filler | 159 | 7.4% | 159 / 0 / 0 |

| Name, number, unit or the issue not captured | 110 | 5.1% | 27 / 80 / 3 |

| No rule for the situation | 100 | 4.7% | 86 / 8 / 6 |

| Talked over the caller | 71 | 3.3% | 71 / 0 / 0 |

| Looped or repeated itself | 45 | 2.1% | 43 / 2 / 0 |

| Misheard a name, number or email | 19 | 0.9% | 16 / 3 / 0 |

| Gave wrong information | 5 | 0.2% | 2 / 1 / 2 |

| Missed an emergency | 5 | 0.2% | 0 / 0 / 5 |

| Tagged a routine call as an emergency | 3 | 0.1% | 3 / 0 / 0 |

The brokerage line flagged at half the rate: 207 of 1,211 calls, 17.1%, with 157 low, 49 medium and 1 high. The biggest types were script order (62), information not captured (36, of which 24 medium) and looping (34). Its one high-severity call was an overflowing sink reported after a showing, taken as a routine message.

The largest group: how the agent talks

Two questions stacked into one turn, a phone number asked for before the reason for the call, filler, talking on over a caller who has started to speak. A caller on the end of that is irritated, and then gives the number. They are the faults you hear first in a demo recording and the cheapest to change, because they live in the script. They also tell the caller within seconds that the line is answered by software, and what that caller does next is the expensive part.

The costly one: a caller who wants a person

On the property management line, 243 callers asked for a live person or a transfer and did not get one. In 143 of them the caller left a full message anyway. In 98 somebody at the operation had to call back or chase. Two were emergencies, covered below.

Flags of any type, split by the reason for the call (the reason is our judgment), point the same way:

| Call reason | Calls | Share with at least one flag |

|---|---|---|

| Complaint | 48 | 65% |

| Request for a staff callback | 268 | 41% |

| Leasing | 254 | 39% |

| Rent and payment | 525 | 36% |

| Maintenance | 316 | 35% |

| Emergency | 167 | 29% |

Complaints and callback requests are calls where the person ringing already wanted a person. The other half of the problem is the "no rule for the situation" row: 100 calls, mostly a caller asking when someone would call back, with no timeframe the agent was allowed to give. What the agent says when asked for a person, and what callback time it can promise, are decisions the operation has to write down. Our property management answering service page sets out what the agent finishes and what it hands off.

The 19 calls that carried real risk

Eighteen on the property management line, sixteen of them on emergency calls, and one on the brokerage line.

Five emergencies logged as routine messages. An active laundry leak flooding a bedroom (July). A missing smoke alarm and a malfunctioning fire alarm (both August). A power outage at the property (August). A new tenant with no access to their keys (August). With the brokerage sink, that is six.

Two emergency callers who asked for a person (July). Both got the intake script instead, and no contact details were captured. One was reporting a fire.

Six callers ringing back about an open emergency. Each asked when help would arrive for something urgent they had already reported. The agent had no rule for a follow-up call on an open emergency, so it could only repeat its escalation line.

Three urgent calls closed without the details needed to act. In one, a caller reporting a flood left after the agent said it was automated, before any contact detail was taken.

Two wrong answers. The agent improvised what a fee covers. On another call it told a caller they could move out on any date with sixty days' notice. Under section 44 of the Residential Tenancies Act, a tenant's notice on a monthly tenancy must be given at least 60 days ahead and must end on the last day of a rental period. Section 37 lets a landlord and tenant agree to end a tenancy on another date, but that needs the landlord's agreement, and the agent offered it as a rule.

None of the six missed emergencies appeared in the grader's own fields. Its "what didn't go well" field was silent on each, and we found them in the call summaries. The grader is a model as well, and on the calls that mattered most it missed what a person reading the summary caught.

Checked against our own emergency list

The property management line's emergency list is published in our after-hours triage rules. It covers water still running, a tenant who cannot get in, no power to a unit and fire, among others.

Three of the five misses were on it: the laundry leak, the tenant without keys and the power outage. (The list escalates a unit without power and logs a dark street as the utility's problem; we classed this call as a miss from its summary.) Fire is on the list too, and one of the two callers who asked for a person was reporting one. On those calls the list existed and the agent did not apply it. That is the uncomfortable finding, and the reason to keep grading after a list is written.

The other two, a missing smoke alarm and a malfunctioning fire alarm, are not on it. Ontario's Fire Code, O. Reg. 213/07, makes the landlord of a rental suite responsible for smoke alarms, which "shall be maintained in operating condition" (Division B, Articles 6.3.3.2 and 6.3.3.3). On the evidence of these two calls, a missing or faulty smoke or fire alarm belongs on the list, and that is our recommendation to anyone writing one.

The list also fired three times when it should not have: a door handle repair, a summer heating repair and cleaning not done before a move in. We rated those low, since no caller was worse off, though each put a routine job on the on-call phone.

What to ask any voice agent vendor, us included

Do you grade every call, and who reads the grades? A grader that is itself a model misses things. On our lines it flagged none of the six missed emergencies.

Can I see the failures, and not just the transcripts? Ask for a monthly count for your own line by type and severity, like the first table above.

What happens when a caller asks for a person, at 2pm on a Tuesday and at 9pm on a Saturday? If it takes a message, ask what callback time it promises and who set it.

How is an emergency classified, and by whose list? It should be your list, written by whoever carries the on-call phone and tested against real calls after launch.

What does the agent say to someone ringing back about an emergency they already reported? Ours had no rule for that call, and six high-severity flags came out of it.

Method, and what this is not

Every call in the window, 1 June to 10 September 2026, has a grader's report: 3,358 in all. Of the 967 flagged calls, 840 were named in the grader's own fields ("what didn't go well" and the list of unanswered questions) and 127 were our judgment from the call summary where those fields said nothing. Another 83 reports were too thin or contradictory to call and are not counted as flags. Severity and type are our classification of the grader's text.

The grader's rubric changed around 10 June, from free prose to named failure labels, so month to month rates are not comparable and we have not presented a trend.

These are graded reports, not the call log, so the counts differ slightly from the call study, which counted 2,142 and 1,190 calls over the same window. Most of the brokerage difference is 20 reports with non-standard headers that the earlier extraction skipped. Human follow-up does not appear in the reports, so we cannot say how many flagged calls a person put right afterwards. No caller name, number, address or recording was kept. The tenancy and Fire Code references are an operator's reading, not legal advice, taken from e-Laws on 23 September 2026.

It covers one operation in Toronto over a single summer. No human-answered line was graded, so nothing here says whether a person would have done better or worse. How we handle call data is on our trust page, and every figure we publish is listed with its source on the proof page.

Frequently asked questions

What mistakes do AI receptionists make most often?

On our two lines, the largest group is about how the agent talks: script order, stacked questions, filler, talking over the caller, repeating itself. Nearly all were low severity. The largest single category was a caller who asked for a live person and did not get one, 11.3% of property management calls. Wrong information was rare, 5 calls in 2,147.

How often does an AI receptionist miss an emergency?

On our property management line, 5 of the 167 calls we classed as emergencies were logged as routine messages, and one more was missed on the brokerage line: six in 3,358 graded calls. Three of the five were items on the operation's written emergency list, and the automated grader flagged none of the six.

What happens when a caller asks an AI receptionist for a real person?

It depends on a rule the operation writes. On our property management line, 243 callers asked for a live person and did not get one: 143 left a full message anyway and 98 needed a call back or a chase. Decide what the agent says, and what callback time it can promise, before launch.

Is an AI receptionist better than a human answering service?

This data cannot say, because no human-answered version of these lines was graded. What it shows is what to expect from a voice agent: mostly low-severity faults, a steady share of callers who wanted a person, and a few emergency misses that justify grading every call.