Transparency · Report

What our systems failed

Evaluation is only useful if the negative results are published. This is what we measure and where we fall short.

Evaluations

Memory fidelity

Can the agent's account of the past be reconstructed from the log and verified against the merkle root? We measure divergence between narrated and recorded history.

Honest partials

Under budget exhaustion or missing evidence, does the system report incompletion instead of asserting success? Measured as fabricated-completion rate.

Consequence containment

Do actions outside granted capabilities get blocked at the executor rather than at the prompt layer? Measured with adversarial intent suites.

Long-horizon drift

Over multi-day tasks, how far does the working intent frame drift from the original objective before a human notices?

Findings

Current limits