Agent Eval Flow

Coding-agent research · September–October 2026

Inside the
coding-agent harness.

A final score gives us an outcome. Retained evidence lets us locate the dropped requirement, inspect the missing test and check whether the harness itself made the right decision.

Pi and OpenCode, the same model, real Python issues. This public edition separates the original frozen experiment, subsequent tuning, fresh validation and synthetic controls.

48original held-out workflows
8 issue families × 3 repeats × 2 agents
24separate follow-up attempts
16 development + 8 fresh validation
0new model calls to regrade
the retained original captures

From score to diagnosis

Three things the evidence made testable.

We evaluated a three-stage coding workflow: investigate without editing, repair the issue, then produce a handoff without further edits. Each agent used its own native tools and system behavior.

01 / FOLLOW THE REQUIREMENT

A lead disappeared between stages.

In a Django cleanup attempt, Pi explicitly flagged pagination for investigation. Its final handoff omitted that lead; the corresponding source stayed unchanged and two independent checks failed.

Comparing the explanation, patch and handoff located an observable omission. That gave subsequent development a concrete target: preserve unresolved obligations through the workflow.

Original held-out observation

Pi · django-async-close · repeat 1. This does not establish an internal memory mechanism.

02 / CHECK THE TEST'S PREMISE

Passing tests can miss an interaction.

On calibration, generated tests often compared two interfaces that took the same route after the patch. They did not establish equivalence between distinct probability and score representations.

On splines, missing inputs combined with a positive fitted range exposed behavior that passing self-tests did not cover. Both boundaries failed in all six original attempts for their respective issue.

Independent behavior checks

These are specific coverage gaps in two selected issue families.

03 / AUDIT THE HARNESS TOO

We corrected our own experimental harness.

Subsequent audits identified a mismatch between our validator and the identifier format our instructions allowed. We corrected both in our experimental harness, in the separate local tuning project. The native Pi and OpenCode packages were unchanged.

Two fresh synthetic controls under that correction passed their independent checks and process acceptance criteria. They establish a working control path; they do not establish better real-bug repair rates.

Later synthetic controls

v5.4 · one disposable addition-bug control per agent.

CaptureKeep original events, tool activity, final patches and phase status.
CheckGrade behavior independently of the agent's explanation.
DiagnoseConnect an outcome to the recorded workflow and test coverage.
RegradeValidate saved captures without another model invocation.

Agent Eval Flow provides the shared capture, result and regrading layer. Case-specific checks, adapters and analyst review supply the domain judgments; the library does not automatically infer these diagnoses from arbitrary logs.

Experiment A · frozen before inference

The complete original outcomes.

Eight held-out issue families in Django, SymPy and scikit-learn; three fresh attempts per agent and issue. All 48 candidate snapshots were independently scorable, including interrupted workflows. A complete workflow and a passing patch are different measurements.

Pi 0.99.1

All selected checks passed13 / 24
All three workflow stages completed14 / 24

OpenCode 1.18.33

All selected checks passed14 / 24
All three workflow stages completed17 / 24

Both used gpt-5.6-sol with low reasoning. Native system prompts, tools and context handling differ. These counts are repeated observations on eight issues and do not isolate an inherent framework effect.

How to read “passed”. Every selected frozen behavioral check passed. This does not prove full task coverage, compatibility with the entire upstream suite, or general correctness. Later exploratory checks are reported separately and never replace the original score.
48 attempts shown
Every original held-out attempt. No failed or interrupted attempt has been removed.
IssueAgentRepeatIndependent checksWorkflow
django-async-closeOpenCode112 / 12 passedComplete
django-async-closeOpenCode212 / 12 passedComplete
django-async-closeOpenCode310 / 12 passedComplete
django-ipaddressOpenCode19 / 9 passedComplete
django-ipaddressOpenCode29 / 9 passedComplete
django-ipaddressOpenCode39 / 9 passedComplete
sklearn-calibrationOpenCode110 / 14 passedComplete
sklearn-calibrationOpenCode210 / 14 passedComplete
sklearn-calibrationOpenCode310 / 14 passedIncomplete
sklearn-splineOpenCode116 / 17 passedComplete
sklearn-splineOpenCode216 / 17 passedComplete
sklearn-splineOpenCode316 / 17 passedIncomplete
sympy-function-rangeOpenCode16 / 6 passedComplete
sympy-function-rangeOpenCode26 / 6 passedComplete
sympy-function-rangeOpenCode36 / 6 passedComplete
sympy-orderOpenCode18 / 8 passedComplete
sympy-orderOpenCode28 / 8 passedComplete
sympy-orderOpenCode38 / 8 passedIncomplete
sympy-quaternionOpenCode15 / 5 passedIncomplete
sympy-quaternionOpenCode25 / 5 passedComplete
sympy-quaternionOpenCode35 / 5 passedComplete
sympy-singularitiesOpenCode113 / 17 passedIncomplete
sympy-singularitiesOpenCode215 / 17 passedIncomplete
sympy-singularitiesOpenCode312 / 17 passedIncomplete
django-async-closePi110 / 12 passedComplete
django-async-closePi28 / 12 passedComplete
django-async-closePi312 / 12 passedComplete
django-ipaddressPi19 / 9 passedComplete
django-ipaddressPi29 / 9 passedIncomplete
django-ipaddressPi39 / 9 passedComplete
sklearn-calibrationPi110 / 14 passedIncomplete
sklearn-calibrationPi210 / 14 passedIncomplete
sklearn-calibrationPi310 / 14 passedIncomplete
sklearn-splinePi114 / 17 passedComplete
sklearn-splinePi214 / 17 passedComplete
sklearn-splinePi316 / 17 passedComplete
sympy-function-rangePi16 / 6 passedComplete
sympy-function-rangePi26 / 6 passedComplete
sympy-function-rangePi36 / 6 passedIncomplete
sympy-orderPi18 / 8 passedComplete
sympy-orderPi28 / 8 passedComplete
sympy-orderPi38 / 8 passedIncomplete
sympy-quaternionPi15 / 5 passedIncomplete
sympy-quaternionPi25 / 5 passedComplete
sympy-quaternionPi35 / 5 passedComplete
sympy-singularitiesPi113 / 17 passedIncomplete
sympy-singularitiesPi212 / 17 passedIncomplete
sympy-singularitiesPi315 / 17 passedIncomplete
Workflow observations and capture qualifications

Two OpenCode attempts changed source during an explicit investigation-only stage: Django IP-address repeat 1 and spline repeat 3. A self-authored implementation plan did not grant permission. These observations are separate from patch correctness.

The capture audit flagged one OpenCode quaternion attempt for further review. A read-only database recovery found the third user prompt and an unfinished assistant record, but no completed handoff. The original incomplete status remains; differences between wall-clock and monotonic durations limit latency interpretation.

For Django cleanup, post-hoc probes exposed QuerySet cleanup gaps and callback behavior when cleanup raises. Those probes were selected after inspecting patches, are exploratory, and do not change the frozen 12-check scores.

The original report also qualifies representation-sensitive symbolic inequality assertions and ambiguity in interpreting the broad Django task scope. These limitations affect interpretation of selected-check scores; they do not silently revise the frozen grading.

Eleven earlier development attempts across four issue families, including an unscorable setup attempt, are retained separately in the downloadable JSON. They are not pooled into the 48 held-out workflows.

Population, protocol and revisions

What stayed fixed.

Before the original experiment

  • Four development issue families were separate from eight held-out families. Related variants stayed in the same split.
  • Task prompts, source snapshots, selected outcome checks, model settings, split assignments and time limits were frozen before held-out inference.
  • Every attempt began with a fresh source checkout and agent session. Maintainer fixes and hidden checks stayed outside the agent environment.
  • Three fixed turns: investigation (10 minutes), repair (15 minutes), handoff (4 minutes). No hidden-test feedback or interactive coaching.

What the checks establish

  • Buggy and maintainer-fixed revisions qualified the selected checks. Passing means the finite selected inventory passed.
  • Outcome checks, workflow completion, process observations and analyst judgments are recorded separately.
  • Manual judgments were unblinded and evidence-linked; unavailable judgments remain unknown.
  • Preserved captures were consolidated and regraded with unchanged grading. All 48 original captures retained their individual measurements with zero new model calls.

Frozen 2026-09-30T14:11:50.541602+00:00 · Python 3.12.7 · container image sha256:965dbc8b7fcb10384b4b252eb71a387021c1ab3e977db7c97777dd54f4c836ac

Original dataset-freeze SHA-256: a7632854b7f067bf8847955916c0858e1cce0609ab0cf32440b60f7e0490053f

All source projects, issue references and split assignments

The full before/after commit IDs and qualification counts are included in study-data.json. Fix dates are provenance, not evidence that the model had never encountered an issue.

IssueProjectSplitBefore → after
pytest-monkeypatchpytest-dev/pytestdevelopmentae777851d8b3 → 472b4de93d83
click-short-helppallets/clickdevelopmentd036881798e3 → 70689853e39c
pytest-walruspytest-dev/pytestdevelopmente984a9b94740 → 51e9a9f148cd
pytest-helppytest-dev/pytestdevelopmentfe9eef7ebc62 → a453d653f92b
django-async-closedjango/djangoheld_out4fab678a0739 → 9e06baf2e4dd
django-ipaddressdjango/djangoheld_out6ade6258480f → 581deb9402ed
sympy-singularitiessympy/sympyheld_out06fcaef06d1e → c59b9b67b0d5
sympy-ordersympy/sympyheld_outc59b9b67b0d5 → 3f58eb0ce808
sympy-quaternionsympy/sympyheld_outc169dc6c618c → a7c459d910e8
sympy-function-rangesympy/sympyheld_out379f910bbb8e → f14b5cd2d13c
sklearn-calibrationscikit-learn/scikit-learnheld_outed8f34a8cb1a → 857849927da6
sklearn-splinescikit-learn/scikit-learnheld_outd059607efd6d → a5046923f202

Experiment B · separate population

What the first follow-up verified.

Once the original failures informed an intervention, those inspected cases became development data. The follow-up compared a general instruction addendum with the same instructions plus native read-only tool restrictions during investigation and handoff. It retained the original tasks and model.

EVIDENCE INTEGRITY

24 captures round-tripped.

All follow-up captures were saved, loaded, validated and regraded through Agent Eval Flow without new model calls during regrading. Dataset freezes and policy locks stayed intact.

RECORDED PHASES

47 no-edit phases inspected.

No net source changes or edit-tool calls were observed in the 47 recorded no-edit phases across both variants. One handoff was absent after a timeout; it is excluded from this denominator.

TOOL CONTROL

Permissions were exercised.

Separate native controls demonstrated edits and shell execution during repair, with no source changes in restricted phases. This verifies the configured tool path, not an operating-system sandbox.

ConfigurationDevelopment: full passesFresh validation: full passesFresh workflows complete
Pi · instructions1 / 41 / 22 / 2
OpenCode · instructions2 / 41 / 21 / 2
Pi · instructions + read-only tools1 / 40 / 22 / 2
OpenCode · instructions + read-only tools2 / 42 / 22 / 2

Both development arms produced 3/8 full passes and 90/104 selected checks. Both fresh-validation arms produced 2/4 full passes. Instructions passed 31/36 checks with 3/4 complete workflows; restrictions passed 33/36 with 4/4 complete workflows. These small counts do not establish a reliable improvement or causal effect.

The fresh reserve contained two SymPy issues: symbol-independent relationals and hypergeometric numerical evaluation. Candidate hashes were locked before inspecting their answer keys. There was one attempt per agent, issue and arm, with no outcome feedback or changes between arms. SymPy also appeared in the earlier experiment.

Follow-up fresh-reserve SHA-256: 77ee4220f243ad1cb313b9cd87ad76890e8e135317ad19cf941b6d9f70422c66

A correct patch could accompany an incorrect explanation. Both OpenCode relational patches passed all 11 selected independent checks by normalizing the relational difference. Their initial explanation incorrectly claimed that independent relations bypassed the generic no-symbol early return. Patch outcomes and diagnosis accuracy were evaluated separately.

Limits that remain. Calibration and spline coverage gaps persisted. In fresh validation, one instruction-only OpenCode workflow timed out during investigation and never reached repair; its score describes unchanged buggy code. Passing a finite check inventory is not universal correctness. Source snapshots detect net changes, not every transient side effect.

Subsequent development · not new held-out results

We also evaluated the evaluator.

Later work added explicit requirement records, independent test authors, review and completion checks. These interventions ran on inspected cases. Their purpose was to diagnose and qualify changes; their results cannot be presented as untouched generalization.

The v5.4 identifier correction is local commit 3458382 in our separate experimental tuning project. It changes our orchestration and validation code, including the instructions that code supplies. It is not a change to Pi or OpenCode, and publishing this report does not publish that fix as an Agent Eval Flow library release.

A semantic coverage gap

50 passing generated checks.

In the v5.2 Pi spline development attempt, all 50 authored checks passed and the completion gate accepted the run. The patch still passed only 16/17 original checks and 0/10 separately qualified supplemental checks.

The missing-query example used a fitted range containing zero, hiding the invalid zero substitute outside a positive range. A second gap involved entirely missing fitted features with explicit nondegenerate knots.

The result identifies test premises to challenge. It does not show that simply generating more tests guarantees the requested behavior.

A corrected validator contract

Valid IDs survived the handoff.

In v5.3 synthetic controls, both patches passed 5/5 independent checks, but our validator silently required IDs matching I[0-9]+ although the prompt permitted stable IDs. It replaced valid requirement records and then rejected their final dispositions.

The v5.4 correction aligned the announced grammar and validation. Two fresh controls passed 5/5 each, retained the original IDs, passed completion and replay checks, and passed the measurement audit with no failed checks.

This is a verified control-path correction. Synthetic controls are excluded from real-issue success rates.

Retained failures, stopped work and current status
  • All three v3 screening arms produced 0/4 complete frozen-check repairs. v4 and v5.2 hard-development batches also produced 0/4. Missing and partial repairs remain in the denominators.
  • In v5.2, qualification stopped both OpenCode attempts before a solver patch. Their scores describe unchanged buggy code. One timeout after collection had been mislabeled as a changed test inventory; the audit records that measurement limitation without rewriting the raw result.
  • Earlier controls exposed validation preparation and collection timeouts. Retained-input replays helped qualify narrow validator changes; those replays are not new native-agent successes.
  • The user stopped the v5.4 hard-development batch. Pi calibration completed with 6/14 on unchanged code and no solver invocation after invalid author fixtures failed qualification. Two other attempts were interrupted; the fourth never started. Full-batch auditing of the partial attempts was not run.
  • A later four-issue fresh Test set remained unused. There is no automatic resume and no broad improvement claim. Original captures and frozen scores remain unchanged.

Read, inspect, qualify

Evidence and practical limits.

This public edition contains all 48 original held-out outcome rows, the 11 earlier development rows in its JSON, source revisions, split assignments, the first follow-up's 24-attempt evidence audit, and the later development/control summaries discussed above.

Native traces and complete experimental workspaces remain in the local research archive. The public export is curated; it is not a self-contained replay of every native run. Source-document hashes identify the retained records without publishing credentials, local filesystem paths, raw reasoning payloads or unused holdout answers.

What remains uncertain

  • The issue sample is small and purposive. Repeats on one issue are correlated; no population reliability estimate or agent ranking follows.
  • Public source and fixes may have been in training data. Recent merge dates do not establish uncontaminated discovery.
  • Code returned in a tool result establishes observable exposure. Understanding, attention and exact later context retention are unobserved.
  • Model/provider delays and shared host resources were not isolated. Token counters and elapsed times are not controlled billing or speed comparisons.

What was not tested

  • All coding tasks, programming languages, deployment environments or native-agent configurations.
  • Compatibility with every full upstream suite, or the absence of every possible defect.
  • A complete provider HTTP trace, all internal subprocesses, every transient write, or an OS-level sandbox guarantee.
  • Generalization of the later development gates. The remaining fresh Test set was not used.

Published study-data.json SHA-256: d1ff72dcda12d309e4b573f29f7a26ce0451d793362256aeee1b1f142742bde0

Retained source-document fingerprints

These fingerprints refer to archived source records, not files downloaded by this page. They support provenance within the retained research archive; they are not independent public replay evidence.

original-results
b137d89a4c38d3315b88b3ee237f6c9a32b3771c12d6dbe5781d8dfdb2638c18

original-freeze
a7632854b7f067bf8847955916c0858e1cce0609ab0cf32440b60f7e0490053f

original-regrade-audit
4c61d443ea3a94024b680047b027e722ba5b1c7e56e15c5e4a7c18211a55b549

findings-record
26168efa567d01bc8be261165b886d59319652fd8e7db979011c5e0ac3dcfead

followup-evidence-audit
b90de65afffda4eccbd6a7b3e96cd428ae98637e24e8b903aa5e1f6c3e4bcbed

development-v5.2-audit
fa7f60fe3c2b5e4587b1ccc3faff3929958222d70e929009c22355478a39bf36

control-v5.4-audit
338e3fdd3d74198542c42f9048435bf33307dab87497b5ddf1f56a7f6c6a4c5c

control-v5.4-regrade
9a503fe0a3af7c4a94c4566b19b21a9f72ba89639409a98907d82f5f0fee71a9

stopped-v5.4-batch
cc5173466626bf06bf9f773e151f6aedee3f76049bffecd37764b15f249c9fce