Skip to main content

Evaluation Framework

Test continuously in simulation.

Benchmark Categories

●       Multi-step tasks

●       Tool calling accuracy

●       Hallucination rate

●       Prompt injection resistance

●       Recovery after outage

●       Cost per successful task

●       Approval accuracy

●       Latency SLAs

●       Harness-vs-model attribution — isolate how much of a score change came from the harness vs. the underlying model

Key Metrics

Metric

Meaning

Task Success Rate

Completed successfully

Autonomy Success Rate

Completed without human intervention or policy breach

Override Frequency

Human corrections

Tool Error Rate

Failed tool calls

Avg Cost / Task

Economics

p95 Latency

Reliability