Evaluation Framework
Test continuously in simulation.
Benchmark Categories
● Multi-step tasks
● Tool calling accuracy
● Hallucination rate
● Prompt injection resistance
● Recovery after outage
● Cost per successful task
● Approval accuracy
● Latency SLAs
● Harness-vs-model attribution — isolate how much of a score change came from the harness vs. the underlying model
Key Metrics
Metric |
Meaning |
|
Task Success Rate |
Completed successfully |
|
Autonomy Success Rate |
Completed without human intervention or policy breach |
|
Override Frequency |
Human corrections |
|
Tool Error Rate |
Failed tool calls |
|
Avg Cost / Task |
Economics |
|
p95 Latency |
Reliability |
No comments to display
No comments to display