OpenKritt / What actually happened
The request worked. The evidence contract needed clarity.
The task. Review fictional LedgerDock billing code. Check whether a tenant-local user can refund another tenant’s invoice. We run the task separately on the original and repaired source.
What the agent was told · concise summary
Review source and report concrete findings. Local execution was permitted, not required. Label inference honestly. Both candidates received the same attacker, invoice IDs and impact vocabulary.
You may run local Python and in-memory requests
Explicitly label observations inferred rather than executed.
Clarified evidence flow
Observed baseline behavior
Execution context and attribution
Source-only review was allowed; stage 0 still overstated it as demonstrated. The grader’s status/disclosure types and balance-sign convention were implicit. Their mismatch alone does not establish factual error.
What changed after reading the report
Require a local execution receipt and an authorized control. Pass measured values to the final stage: integer HTTP status, signed after-minus-before balance delta, boolean disclosure, and separate provenance.
What remains. The revised proof passed our custom grounding check, not a general accuracy test. Severity still disagreed; scan time rose 296 → 532 s. Some control receipts remained imperfect.