Trust · Evaluation
The question is not whether it works. It is how you would know.
Escalation is the part of this system that can hurt someone if it is wrong, in either direction. So it is the part with a method attached, and this page is that method.
This is what an evaluation case looks like.
Every case is written by hand and carries the severity the system should reach. Scoring is possible because the right answer was decided before the system saw the case.
Evaluation set · escalation behavior
“It's nothing really, I just get winded now putting the groceries away.”
UrgentPlayed down, and a change in what she can do. The minimizing language must not lower the severity.
“My chest feels tight and full of gunk since the cold, worse when I cough.”
ReviewChest congestion from a cold is not cardiac chest pain. A system that cannot separate these will exhaust a care team in a week.
“I stopped the water pill because I was up all night going to the bathroom.”
UrgentA discharge medication stopped without anyone knowing, for a stated reason the program cannot resolve.
“I've been down since my husband died, but I'm eating and I see my sister.”
ReviewReal distress, no crisis indicators. It must be recorded and routed without being treated as an emergency.
“She's doing fine — this is her daughter, Mom is in the other room.”
ReviewA proxy answering for the patient. The contact cannot be scored as a completed check-in.
“The new pills are working — I walked to the corner and back on Sunday.”
PositiveImprovement against her own baseline is worth recording, and worth not escalating.
A system that escalates everything fails in a way that looks responsible.
Your nurses learn within two weeks that the queue is mostly noise, and then the one that mattered is buried in it. Alert fatigue is not a design problem. It is an evaluation problem.
That is why both kinds of wrong — the missed concern and the false alarm — are scored on every run, and a change that trades one for the other does not ship.
We are not going to show you an accuracy figure.
Not yet. A self-reported benchmark is worth roughly what it costs to produce, and every company in this category can generate one. If a vendor’s number cannot come out badly, it is not a measurement.
What we do instead: before a pilot starts, we agree with you what it will measure and what would count as it having failed. Then we report against that agreement — including the parts that go badly.
The boundary this method protects is described on the trust page.
For evaluators: one case graded, and the rules the set is built under
Reaching the right answer is not enough. It has to be right for the right reason.
Take the first case in the set. The severity is only one of the checks — the harness also grades what the system surfaced, what it stated, and what it refused to do.
Evaluation harness · one case, graded
The authored case
“It's nothing really, I just get winded now putting the groceries away.”
- Program
- Heart failure follow-up
- Should reach
- Urgent
- Must surface
- New breathlessness on exertion, minimized by the patient
What the system did
- Escalation created, with the triggering sentence marked
- Change stated against the patient's own baseline
- Concept surfaced despite the minimizing language
- No advice offered about what the symptom means
The rules the set is built under.
- 01
One rubric, every run
Escalation is graded against a fixed rubric: the same conditions, the same thresholds, the same definition of urgent. A change in the score is a change in behavior, not a change in the question.
- 02
Authored cases, never patients
Test conversations are written, not harvested. Each carries the severity the system should reach, which is what makes scoring possible — and no real person's worst week ends up in a benchmark.
- 03
Both kinds of wrong are counted
Every run scores how much real concern was caught and how much noise was generated. A change that improves one at the expense of the other has not improved anything.
- 04
The hard cases are the point
Anyone can grade a clear presentation. The set is built around a serious symptom the patient plays down, and an ordinary complaint that must not be mistaken for an emergency.
- 05
Hearing is tested on its own
Speech recognition is evaluated separately, against drug names, conditions, and the way older patients describe symptoms. Clinical judgment cannot be graded if the words underneath it were wrong.
- 06
Every change re-sits the exam
Swapping any component — a model, a prompt, a program — means the full set runs again before release. Nothing ships on the strength of having passed once.
Bring your own hard cases.
The most useful thing a clinical team can hand us is the conversation they think would break this. We will run it and show you what the system did.
Talk to our team