# Evaluation Framework

Test continuously in simulation.

## Benchmark Categories

● Multi-step tasks

● Tool calling accuracy

● Hallucination rate

● Prompt injection resistance

● Recovery after outage

● Cost per successful task

● Approval accuracy

● Latency SLAs

● Harness-vs-model attribution — isolate how much of a score change came from the harness vs. the underlying model

## Key Metrics

<div align="left" dir="ltr" id="bkmrk-metric-meaning-task-"><table><colgroup><col width="311"></col><col width="312"></col></colgroup><tbody><tr><td>### Metric

</td><td>### Meaning

</td></tr><tr><td>Task Success Rate

</td><td>Completed successfully

</td></tr><tr><td>Autonomy Success Rate

</td><td>Completed without human intervention or policy breach

</td></tr><tr><td>Override Frequency

</td><td>Human corrections

</td></tr><tr><td>Tool Error Rate

</td><td>Failed tool calls

</td></tr><tr><td>Avg Cost / Task

</td><td>Economics

</td></tr><tr><td>p95 Latency

</td><td>Reliability

</td></tr></tbody></table>

</div>