Testing · Data
AI receptionist test results, September 2026: 118 test calls by call type and specialty
On 30 September 2026 we put the receptionist through two runs: 100 test calls in text mode and 18 specialty calls in voice mode. 111 of the 118 calls passed. Here is every number, each of the seven failures, and the raw file. Every caller is a test program, not a patient, and the receptionist's lines are its real output.
The headline numbers
| Run | Calls | Passed | Pass rate |
|---|---|---|---|
| Text run: the caller types | 100 | 94 | 94% |
| Voice run: the caller speaks, through US speech vendors | 18 | 17 | 94% |
| Everyday calls: booking, refills, records, emergencies, outbound | 79 | 79 | 100% |
| Specialty screening calls, both runs | 39 | 32 | 82% |
By call type, text run
Each call type is a script: a caller with a name, a date of birth and a reason, who answers what the receptionist asks. A pass means the outcome matched the script, identity came before anything personal was said, and the receptionist said nothing it must never say.
| Call type | Calls | Passed |
|---|---|---|
| New patient books a visit | 6 | 6 |
| New patient with an urgent concern | 3 | 3 |
| Existing patient books a visit | 8 | 8 |
| Moves a visit | 8 | 8 |
| Cancels a visit | 6 | 6 |
| Confirms a visit | 4 | 4 |
| Refuses to give a date of birth | 6 | 6 |
| Asks for a refill | 6 | 6 |
| Billing question | 4 | 4 |
| Records request | 3 | 3 |
| Office questions: hours, directions, insurance | 8 | 8 |
| Asks for a person | 4 | 4 |
| Says emergency words | 3 | 3 |
| Wrong number | 3 | 3 |
| Outbound: returns a requested call | 4 | 4 |
| Outbound: the wrong person answers | 3 | 3 |
| Specialty concern, screened with the practice's form | 21 | 15 |
By specialty
Specialty calls are where screening happens: the receptionist picks the practice's approved form for the concern, asks its questions, and books, escalates or gives 911 instructions from the answers. With one or two calls per specialty, a single failure moves a rate a long way, so read these as counts, not percentages.
| Specialty | Text run | Voice run |
|---|---|---|
| Cardiology | 2 of 2 | 2 of 2 |
| Gynecology | 2 of 2 | 2 of 2 |
| Neurology | 2 of 2 | 2 of 2 |
| Pediatrics | 1 of 2 | 2 of 2 |
| Dentistry | 0 of 2 | 1 of 2 |
| Dermatology | 1 of 1 | 1 of 1 |
| ENT | 0 of 1 | 1 of 1 |
| General practice | 1 of 1 | 1 of 1 |
| General surgery | 0 of 1 | 1 of 1 |
| Neurosurgery | 1 of 1 | 1 of 1 |
| Ophthalmology | 1 of 1 | 1 of 1 |
| Pain management | 1 of 1 | 1 of 1 |
| Psychiatry | 0 of 1 | 1 of 1 |
| Gastroenterology | 1 of 1 | not run |
| Pathology | 1 of 1 | not run |
| Radiology | 1 of 1 | not run |
What failed in the text run
All six text failures were the same mistake. Each call ended the right way: the toothache with a swollen cheek was escalated as urgent, and the other five were booked as routine. But the screening questions came from a sub-specialty's form instead of the general one the scenario expected: oral and maxillofacial surgery for two dental calls, pediatric surgery for a child with a fever and a cough, child and adolescent psychiatry for an anxiety call, the otorhinolaryngology form for an ear ache, and bariatric post-op for a general surgery dressing question. The harness fails these calls on purpose, because the practice approved a specific form for each concern, and the right form matters even when the outcome is right.
What failed in the voice run
One dental call failed. The test caller answered 'yes' before the receptionist had finished asking about text confirmations, the conversation stalled, and the caller hung up after about five minutes without being screened. Text mode cannot catch this kind of failure, which is why voice runs exist.
How to read these numbers
- Cedar Clinic is a test practice and every caller is a test program, not a patient. The receptionist's lines are its real output.
- This is one text run and one voice run on one day, not a clinical validation. A real deployment is tested again against the practice's own forms before go-live.
- Text mode tests the receptionist's decisions without speech. Voice mode adds speech recognition and speech synthesis, from a laptop in India talking to speech vendors in the United States.
- Earlier September text runs ranged from 78 to 100 of 100 while specialty packs were added and the harness itself was fixed; this page reports the last complete runs of the month.
We will publish the next runs the same way, with the raw file, including the ones that go badly.