Evaluation is only useful if the negative results are published. This is what we measure and where we fall short.
Can the agent's account of the past be reconstructed from the log and verified against the merkle root? We measure divergence between narrated and recorded history.
Under budget exhaustion or missing evidence, does the system report incompletion instead of asserting success? Measured as fabricated-completion rate.
Do actions outside granted capabilities get blocked at the executor rather than at the prompt layer? Measured with adversarial intent suites.
Over multi-day tasks, how far does the working intent frame drift from the original objective before a human notices?