21 cases · 21 high or critical
When the tool fails, the agent’s honesty is the test
Scenario coverage for scheduling, CRM, lookup, and payment-adjacent tool outages.
Read the report →Scenario coverage reports
These reports are generated from the full 100-case product dataset at build time. Each one turns a failure pattern into a small set of checks your team can rerun.
Counts describe scenarios in the dataset. No vendor or customer agent was empirically tested to produce these reports.
21 cases · 21 high or critical
Scenario coverage for scheduling, CRM, lookup, and payment-adjacent tool outages.
Read the report →17 cases · 14 high or critical
The safety-focused report for gas, electrical, flooding, storm, and structural-risk calls.
Read the report →13 cases · 13 high or critical
A report on unsupported certainty, permission boundaries, and actions the system did not authorize.
Read the report →Run it yourself
The browser scorecard stores your review locally. It does not upload call transcripts or notes.
For agents and crawlers: agent-manifest.json · llms.txt · free GitHub sample