AGENT EVAL FLOW · SAVED EVALUATION

GPT Researcher · recorded pilot

Project gpt-researcher · Dataset gptr-selected-simpleqa · Suite recorded-simpleqa-pilot / 1

Created 2026-09-23 18:06:54.540027+00:00 · Result evaluation/cce8b41291724ed2ba4cc744746b669c

Study coverage

3Planned runs
3Completed
0Failed
0Pending
0Unavailable
3Unknown resources

3 task units · 1 repetition(s) per candidate. Diagnostic detail rows are evidence within a run.

Recorded candidate summaries

CandidateSummaryValue and basisSource coverageExplanation
gptr-codex-luna-recordedtarget_correct_mean0.6666666666666666estimated Complete planned population3 / 3 observed
0 missing · 0 errors · 0 not applicable
Complete planned population

Available subset: 0.6666666666666666 (descriptive)

gptr-codex-luna-recordedmodel_calls_mean3.0observed Complete planned population3 / 3 observed
0 missing · 0 errors · 0 not applicable
Complete planned population

Available subset: 3.0 (descriptive)

gptr-codex-luna-recordedsearch_calls_mean5.0observed Complete planned population3 / 3 observed
0 missing · 0 errors · 0 not applicable
Complete planned population

Available subset: 5.0 (descriptive)

gptr-codex-luna-recordedcaptured_pages_mean10.0observed Complete planned population3 / 3 observed
0 missing · 0 errors · 0 not applicable
Complete planned population

Available subset: 10.0 (descriptive)

gptr-codex-luna-recordedelapsed_seconds_mean134.2096201578776observed Complete planned population3 / 3 observed
0 missing · 0 errors · 0 not applicable
Complete planned population

Available subset: 134.2096201578776 (descriptive)

Grading resources

Shared grading activities are counted once. These charges are separate from agent execution costs below.

All retained and new grading

Cost USD · scope: model
unknownunknown The complete execution/grading inventory is not established
Input tokens
unknownunknown The complete execution/grading inventory is not established
Output tokens
unknownunknown The complete execution/grading inventory is not established
Human minutes
unknownunknown The complete execution/grading inventory is not established

Captured native grading inventory: unknownunknown Public artifact summarizes original grading receipts

Grading performed in this evaluation

Cost USD · scope: model
unknownunknown At least one cost_usd observation is unknown
Input tokens
29390observed Sum of exclusive recorded quantities
Output tokens
15observed Sum of exclusive recorded quantities
Human minutes
unknownunknown At least one human_minutes observation is unknown

Performed grading inventory: Trueobserved All contributing inventories are complete

15 recorded grading activities
[
  {
    "id": "grading/c40904f0e1f2456cbe55c342c31c26a8",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "b03c2eea605acc848623799235fcf1e81fdccd24383570a31eacd3ae3c454261",
    "run_ids": [
      "recorded/simpleqa-3810"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": null,
        "status": "unknown",
        "reason": "No dollar charge reported by Codex CLI",
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 9599,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 5,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": null,
        "status": "unknown",
        "reason": "Not timed",
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.523497+00:00",
    "ended_at": "2026-09-23 18:06:54.526276+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/9210b9f68d6549bd998c39d9e5d539a8",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "b03c2eea605acc848623799235fcf1e81fdccd24383570a31eacd3ae3c454261",
    "run_ids": [
      "recorded/simpleqa-3787"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": null,
        "status": "unknown",
        "reason": "No dollar charge reported by Codex CLI",
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 9843,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 5,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": null,
        "status": "unknown",
        "reason": "Not timed",
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.526628+00:00",
    "ended_at": "2026-09-23 18:06:54.526628+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/9d6c8c5f1fa24b0f897eb8202d8664ab",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "b03c2eea605acc848623799235fcf1e81fdccd24383570a31eacd3ae3c454261",
    "run_ids": [
      "recorded/simpleqa-1111"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": null,
        "status": "unknown",
        "reason": "No dollar charge reported by Codex CLI",
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 9948,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 5,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": null,
        "status": "unknown",
        "reason": "Not timed",
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.526628+00:00",
    "ended_at": "2026-09-23 18:06:54.527629+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/ff466308c12545a1a4e0c0220779d314",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "97cc2aa2d8df6cbea5faed9b3e6ce9e03e6b8729bb246b6272e05423fe6cfd16",
    "run_ids": [
      "recorded/simpleqa-3810"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.527629+00:00",
    "ended_at": "2026-09-23 18:06:54.528631+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/40ff57e6a2e047dc9df9edec648c9076",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "97cc2aa2d8df6cbea5faed9b3e6ce9e03e6b8729bb246b6272e05423fe6cfd16",
    "run_ids": [
      "recorded/simpleqa-3787"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.528631+00:00",
    "ended_at": "2026-09-23 18:06:54.528631+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/a9952fe836e94765a44b97b840ffeff5",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "97cc2aa2d8df6cbea5faed9b3e6ce9e03e6b8729bb246b6272e05423fe6cfd16",
    "run_ids": [
      "recorded/simpleqa-1111"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.528631+00:00",
    "ended_at": "2026-09-23 18:06:54.529630+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/c823b51e22754db095736e48062380bd",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "cbe115c9965c6423ef4bee2702fa0a055c23825ea606b33d09930682ebc48919",
    "run_ids": [
      "recorded/simpleqa-3810"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.530629+00:00",
    "ended_at": "2026-09-23 18:06:54.531796+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/015d7e064de14ea9a0d95c84e3916951",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "cbe115c9965c6423ef4bee2702fa0a055c23825ea606b33d09930682ebc48919",
    "run_ids": [
      "recorded/simpleqa-3787"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.531796+00:00",
    "ended_at": "2026-09-23 18:06:54.532466+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/93920bc5a78749cfb332afd0fe9f5746",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "cbe115c9965c6423ef4bee2702fa0a055c23825ea606b33d09930682ebc48919",
    "run_ids": [
      "recorded/simpleqa-1111"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.532670+00:00",
    "ended_at": "2026-09-23 18:06:54.532670+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/4f30fead876d4812a7671ee1aa720d74",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "43187ba6e818c8a3e6056b77160fa2f2625f8b46824ff1a94a60a1bf5f60e224",
    "run_ids": [
      "recorded/simpleqa-3810"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.532670+00:00",
    "ended_at": "2026-09-23 18:06:54.534262+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/284de19681fd4a25a34f9678fc82a057",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "43187ba6e818c8a3e6056b77160fa2f2625f8b46824ff1a94a60a1bf5f60e224",
    "run_ids": [
      "recorded/simpleqa-3787"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.534762+00:00",
    "ended_at": "2026-09-23 18:06:54.535008+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/1051ed71103b43bc84cf379a291a6ca3",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "43187ba6e818c8a3e6056b77160fa2f2625f8b46824ff1a94a60a1bf5f60e224",
    "run_ids": [
      "recorded/simpleqa-1111"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.535508+00:00",
    "ended_at": "2026-09-23 18:06:54.536007+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/611f72077b1b481f8c183cba88fdc549",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "fc9c16723bae0743de4bebb0b2523b2e2146cbe84e5d34a1653f7198700448a4",
    "run_ids": [
      "recorded/simpleqa-3810"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.537007+00:00",
    "ended_at": "2026-09-23 18:06:54.537007+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/100f3c2b6158446d8829f91d807dfca3",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "fc9c16723bae0743de4bebb0b2523b2e2146cbe84e5d34a1653f7198700448a4",
    "run_ids": [
      "recorded/simpleqa-3787"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.537007+00:00",
    "ended_at": "2026-09-23 18:06:54.538024+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  },
  {
    "id": "grading/9d77a32a887a4da8ba09c5e6a63fa6fd",
    "evaluator": {
      "name": "example.gpt-researcher-recorded-metrics",
      "revision": "1"
    },
    "config_fingerprint": "fc9c16723bae0743de4bebb0b2523b2e2146cbe84e5d34a1653f7198700448a4",
    "run_ids": [
      "recorded/simpleqa-1111"
    ],
    "status": "completed",
    "phase": "post_run",
    "resources": {
      "cost_usd": {
        "value": "0",
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 0,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": 0.0,
        "status": "observed",
        "reason": null,
        "evidence": []
      }
    },
    "started_at": "2026-09-23 18:06:54.538024+00:00",
    "ended_at": "2026-09-23 18:06:54.538024+00:00",
    "native_refs": {},
    "artifacts": {},
    "error": null
  }
]

Performed activity IDs

[
  "grading/c40904f0e1f2456cbe55c342c31c26a8",
  "grading/9210b9f68d6549bd998c39d9e5d539a8",
  "grading/9d6c8c5f1fa24b0f897eb8202d8664ab",
  "grading/ff466308c12545a1a4e0c0220779d314",
  "grading/40ff57e6a2e047dc9df9edec648c9076",
  "grading/a9952fe836e94765a44b97b840ffeff5",
  "grading/c823b51e22754db095736e48062380bd",
  "grading/015d7e064de14ea9a0d95c84e3916951",
  "grading/93920bc5a78749cfb332afd0fe9f5746",
  "grading/4f30fead876d4812a7671ee1aa720d74",
  "grading/284de19681fd4a25a34f9678fc82a057",
  "grading/1051ed71103b43bc84cf379a291a6ca3",
  "grading/611f72077b1b481f8c183cba88fdc549",
  "grading/100f3c2b6158446d8829f91d807dfca3",
  "grading/9d77a32a887a4da8ba09c5e6a63fa6fd"
]

Runs and evidence

Candidate / taskRepetitionRun statusScoreAcceptanceDetails
gptr-codex-luna-recorded
{ "case_id": "simpleqa-3810" }
0completedunknownunknown No rubric configuredpassInspect run
gptr-codex-luna-recorded
{ "case_id": "simpleqa-3787" }
0completedunknownunknown No rubric configuredfailInspect run
gptr-codex-luna-recorded
{ "case_id": "simpleqa-1111" }
0completedunknownunknown No rubric configuredpassInspect run

gptr-codex-luna-recorded · { "case_id": "simpleqa-3810" } · repetition 0

recorded/simpleqa-3810 · completed · output available

Elapsed seconds: 212.127206observed

Agent execution resources

Cost USD · scope: model
unknownunknown The complete execution/grading inventory is not established
Input tokens
unknownunknown The complete execution/grading inventory is not established
Output tokens
unknownunknown The complete execution/grading inventory is not established
Human minutes
unknownunknown The complete execution/grading inventory is not established
Delivered output
{
  "answer_summary": "Roberto Vázquez García",
  "grade": {
    "case_id": "simpleqa-3810",
    "grade": "CORRECT",
    "target_correct": 1,
    "expected_answer": "Roberto Vázquez García",
    "report_sha256": "aea4017e0c3e32e46dd1f0e3a945adc3c68a3ba8f36e44f57a27210c308747ca",
    "grader_model_requested": "gpt-5.6-luna",
    "rubric": "upstream SimpleQA GRADER_TEMPLATE",
    "limitation": "Same-model judge, not independent ground truth or complete report factuality."
  },
  "metrics": {
    "target_correct": 1,
    "model_calls": 3,
    "search_calls": 5,
    "captured_pages": 13,
    "elapsed_seconds": 212.127206325531
  },
  "judge_usage": {
    "input_tokens": 9599,
    "output_tokens": 5,
    "cached_input_tokens": 4864
  }
}

Measurements

Metric / detail keyValueStatus / basisWhyActivity IDs
target_correct1ok / estimatedRecorded same-model SimpleQA judgment of target only; not whole-report factualitygrading/c40904f0e1f2456cbe55c342c31c26a8
model_calls3ok / observedRecorded native count or timinggrading/ff466308c12545a1a4e0c0220779d314
search_calls5ok / observedRecorded native count or timinggrading/c823b51e22754db095736e48062380bd
captured_pages13ok / observedRecorded native count or timinggrading/4f30fead876d4812a7671ee1aa720d74
elapsed_seconds212.127206325531ok / observedRecorded native count or timinggrading/611f72077b1b481f8c183cba88fdc549
Score contributions and acceptance gates
{
  "run_id": "recorded/simpleqa-3810",
  "score": {
    "value": null,
    "status": "unknown",
    "reason": "No rubric configured",
    "evidence": []
  },
  "contributions": [],
  "acceptance": "pass",
  "acceptance_basis": "estimated",
  "gates": [
    {
      "threshold": {
        "metric": "target_correct",
        "op": "==",
        "value": 1
      },
      "decision": "pass",
      "basis": "estimated",
      "reason": "target_correct: 1 == 1 is pass"
    }
  ],
  "reason": "Applied configured rubric and acceptance rule"
}
1 executions · 3 events
[
  {
    "id": "recorded/simpleqa-3810/aggregate",
    "slot": "recorded",
    "retry_index": 0,
    "parent_id": null,
    "status": "completed",
    "started_at": "2026-09-23 15:48:17.944108+00:00",
    "ended_at": "2026-09-23 15:51:50.071314+00:00",
    "effective_config": {
      "value": {
        "model_requested": "gpt-5.6-luna",
        "resolved_model": "Not emitted by Codex CLI JSONL",
        "provider": "Experimental local Codex CLI transport",
        "reasoning": "low",
        "retriever": "DuckDuckGo",
        "embedding": "local BAAI/bge-small-en-v1.5",
        "max_iterations": 3,
        "results_per_query": 5,
        "report_min_words": 600,
        "limitations": "Native temperature and max_tokens settings cannot be enforced by this CLI transport; cost in dollars is unknown."
      },
      "status": "observed",
      "reason": null,
      "evidence": []
    },
    "resources": {
      "cost_usd": {
        "value": null,
        "status": "unknown",
        "reason": "No dollar charge reported by Codex CLI",
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 29910,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 1791,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": null,
        "status": "unknown",
        "reason": "Not timed",
        "evidence": []
      }
    },
    "role": "main",
    "native_refs": {},
    "error": null
  }
]
[
  {
    "id": "recorded/simpleqa-3810/aggregate/call/1",
    "execution_id": "recorded/simpleqa-3810/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 1,
      "stage": "agent selection",
      "request_sha256": "0efe710fd67a268ad624df65c0ac65ac88180ff17d716d1756482f0f4807a16c",
      "response_sha256": "e30b8fb52b972b322f8a693ec6b315cc31dc40d5f0662d7cc9853824bc445eca"
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/0/trace/0",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  },
  {
    "id": "recorded/simpleqa-3810/aggregate/call/2",
    "execution_id": "recorded/simpleqa-3810/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 2,
      "stage": "query planning",
      "request_sha256": "4716f855f1afa48ca275b1d23316e1301b733556136869ef303feaf32ba51cd0",
      "response_sha256": "7b7f80d778af3ed662fdc0eb361823e6ec8db550212c45001865b61066e8109a"
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/0/trace/1",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  },
  {
    "id": "recorded/simpleqa-3810/aggregate/call/3",
    "execution_id": "recorded/simpleqa-3810/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 3,
      "stage": "report writing",
      "request_sha256": "5fb0faa1d372a0a7e957eb28c48b46c88e38681935eb3179a7ed7dd125e8852a",
      "response_sha256": "aea4017e0c3e32e46dd1f0e3a945adc3c68a3ba8f36e44f57a27210c308747ca"
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/0/trace/2",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  }
]
Environment and native references
{
  "value": {
    "os": "Windows",
    "repo": "https://github.com/assafelovic/gpt-researcher",
    "commit": "6f998577d547b1e54ec662dac63583aa11e3b84b"
  },
  "status": "observed",
  "reason": null,
  "evidence": []
}
{
  "case_id": "simpleqa-3810",
  "row_index": "0",
  "public_capture_sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
}

Evidence references

gptr-codex-luna-recorded · { "case_id": "simpleqa-3787" } · repetition 0

recorded/simpleqa-3787 · completed · output available

Elapsed seconds: 91.899897observed

Agent execution resources

Cost USD · scope: model
unknownunknown The complete execution/grading inventory is not established
Input tokens
unknownunknown The complete execution/grading inventory is not established
Output tokens
unknownunknown The complete execution/grading inventory is not established
Human minutes
unknownunknown The complete execution/grading inventory is not established
Delivered output
{
  "answer_summary": "Chirik",
  "grade": {
    "case_id": "simpleqa-3787",
    "grade": "INCORRECT",
    "target_correct": 0,
    "expected_answer": "Anastas",
    "report_sha256": "a5bdaaf44937066dd04b195d7e0e87c049d12a32a35d340b187fcf36e5c2e684",
    "grader_model_requested": "gpt-5.6-luna",
    "rubric": "upstream SimpleQA GRADER_TEMPLATE",
    "limitation": "Same-model judge, not independent ground truth or complete report factuality."
  },
  "metrics": {
    "target_correct": 0,
    "model_calls": 3,
    "search_calls": 5,
    "captured_pages": 7,
    "elapsed_seconds": 91.90039706230164
  },
  "judge_usage": {
    "input_tokens": 9843,
    "output_tokens": 5,
    "cached_input_tokens": 0
  }
}

Measurements

Metric / detail keyValueStatus / basisWhyActivity IDs
target_correct0ok / estimatedRecorded same-model SimpleQA judgment of target only; not whole-report factualitygrading/9210b9f68d6549bd998c39d9e5d539a8
model_calls3ok / observedRecorded native count or timinggrading/40ff57e6a2e047dc9df9edec648c9076
search_calls5ok / observedRecorded native count or timinggrading/015d7e064de14ea9a0d95c84e3916951
captured_pages7ok / observedRecorded native count or timinggrading/284de19681fd4a25a34f9678fc82a057
elapsed_seconds91.90039706230164ok / observedRecorded native count or timinggrading/100f3c2b6158446d8829f91d807dfca3
Score contributions and acceptance gates
{
  "run_id": "recorded/simpleqa-3787",
  "score": {
    "value": null,
    "status": "unknown",
    "reason": "No rubric configured",
    "evidence": []
  },
  "contributions": [],
  "acceptance": "fail",
  "acceptance_basis": "estimated",
  "gates": [
    {
      "threshold": {
        "metric": "target_correct",
        "op": "==",
        "value": 1
      },
      "decision": "fail",
      "basis": "estimated",
      "reason": "target_correct: 0 == 1 is fail"
    }
  ],
  "reason": "Applied configured rubric and acceptance rule"
}
1 executions · 3 events
[
  {
    "id": "recorded/simpleqa-3787/aggregate",
    "slot": "recorded",
    "retry_index": 0,
    "parent_id": null,
    "status": "completed",
    "started_at": "2026-09-23 15:53:56.189033+00:00",
    "ended_at": "2026-09-23 15:55:28.088930+00:00",
    "effective_config": {
      "value": {
        "model_requested": "gpt-5.6-luna",
        "resolved_model": "Not emitted by Codex CLI JSONL",
        "provider": "Experimental local Codex CLI transport",
        "reasoning": "low",
        "retriever": "DuckDuckGo",
        "embedding": "local BAAI/bge-small-en-v1.5",
        "max_iterations": 3,
        "results_per_query": 5,
        "report_min_words": 600,
        "limitations": "Native temperature and max_tokens settings cannot be enforced by this CLI transport; cost in dollars is unknown."
      },
      "status": "observed",
      "reason": null,
      "evidence": []
    },
    "resources": {
      "cost_usd": {
        "value": null,
        "status": "unknown",
        "reason": "No dollar charge reported by Codex CLI",
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 24869,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 1990,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": null,
        "status": "unknown",
        "reason": "Not timed",
        "evidence": []
      }
    },
    "role": "main",
    "native_refs": {},
    "error": null
  }
]
[
  {
    "id": "recorded/simpleqa-3787/aggregate/call/1",
    "execution_id": "recorded/simpleqa-3787/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 1,
      "stage": "agent selection",
      "request_sha256": "abf5c7db245f64989bbc739fbc30612f58e492d80792831e22f6a1473c74f06d",
      "response_sha256": "b9d009e159e6569484f0474060a7621364b909c1a0d1d35e34d3403c82e2cb00",
      "observation": "Selected a scientific research role; did not choose an award winner.",
      "prompt_term_counts": {
        "Anastas": 0,
        "Yale": 0,
        "Royal Society of Chemistry": 0,
        "Chirik": 0,
        "DNSError": 0
      }
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/2/trace/0",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  },
  {
    "id": "recorded/simpleqa-3787/aggregate/call/2",
    "execution_id": "recorded/simpleqa-3787/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 2,
      "stage": "query planning",
      "request_sha256": "cdca9d014253d98e4c66a48b47470a43a1d21e6b740466e6e20cc4029118bf99",
      "response_sha256": "82676acf90ae2e6a5284e4c5ddc72739b33878361087c4f91a6deaa47ff6d0d6",
      "observation": "Saw Yale and RSC leads for Paul Anastas and generated a targeted verification query.",
      "prompt_term_counts": {
        "Anastas": 3,
        "Yale": 2,
        "Royal Society of Chemistry": 2,
        "Chirik": 1,
        "DNSError": 0
      },
      "generated_queries": [
        "2016 Green Chemistry Award winner surname",
        "Paul Anastas 2016 Royal Society of Chemistry Green Chemistry Award",
        "who won the Green Chemistry Award in 2016"
      ]
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/2/trace/1",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  },
  {
    "id": "recorded/simpleqa-3787/aggregate/call/3",
    "execution_id": "recorded/simpleqa-3787/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 3,
      "stage": "report writing",
      "request_sha256": "19286c29e4f542185fd30a8c7d8b9df9b385f067e7643507abc34d85f4dbda8f",
      "response_sha256": "a5bdaaf44937066dd04b195d7e0e87c049d12a32a35d340b187fcf36e5c2e684",
      "observation": "Answered Chirik. The supplied prompt had no Anastas, Yale, Royal Society of Chemistry, or retrieval error notice.",
      "prompt_term_counts": {
        "Anastas": 0,
        "Yale": 0,
        "Royal Society of Chemistry": 0,
        "Chirik": 5,
        "DNSError": 0
      }
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/2/trace/2",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  }
]
Environment and native references
{
  "value": {
    "os": "Windows",
    "repo": "https://github.com/assafelovic/gpt-researcher",
    "commit": "6f998577d547b1e54ec662dac63583aa11e3b84b"
  },
  "status": "observed",
  "reason": null,
  "evidence": []
}
{
  "case_id": "simpleqa-3787",
  "row_index": "2",
  "public_capture_sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
}

Evidence references

gptr-codex-luna-recorded · { "case_id": "simpleqa-1111" } · repetition 0

recorded/simpleqa-1111 · completed · output available

Elapsed seconds: 98.601257observed

Agent execution resources

Cost USD · scope: model
unknownunknown The complete execution/grading inventory is not established
Input tokens
unknownunknown The complete execution/grading inventory is not established
Output tokens
unknownunknown The complete execution/grading inventory is not established
Human minutes
unknownunknown The complete execution/grading inventory is not established
Delivered output
{
  "answer_summary": "Olga Polverino",
  "grade": {
    "case_id": "simpleqa-1111",
    "grade": "CORRECT",
    "target_correct": 1,
    "expected_answer": "Olga Polverino",
    "report_sha256": "561247a4bba652acd6540c96dca9c5ae7d9e7806af689ace1eb51ad225350793",
    "grader_model_requested": "gpt-5.6-luna",
    "rubric": "upstream SimpleQA GRADER_TEMPLATE",
    "limitation": "Same-model judge, not independent ground truth or complete report factuality."
  },
  "metrics": {
    "target_correct": 1,
    "model_calls": 3,
    "search_calls": 5,
    "captured_pages": 10,
    "elapsed_seconds": 98.60125708580017
  },
  "judge_usage": {
    "input_tokens": 9948,
    "output_tokens": 5,
    "cached_input_tokens": 5888
  }
}

Measurements

Metric / detail keyValueStatus / basisWhyActivity IDs
target_correct1ok / estimatedRecorded same-model SimpleQA judgment of target only; not whole-report factualitygrading/9d6c8c5f1fa24b0f897eb8202d8664ab
model_calls3ok / observedRecorded native count or timinggrading/a9952fe836e94765a44b97b840ffeff5
search_calls5ok / observedRecorded native count or timinggrading/93920bc5a78749cfb332afd0fe9f5746
captured_pages10ok / observedRecorded native count or timinggrading/1051ed71103b43bc84cf379a291a6ca3
elapsed_seconds98.60125708580017ok / observedRecorded native count or timinggrading/9d77a32a887a4da8ba09c5e6a63fa6fd
Score contributions and acceptance gates
{
  "run_id": "recorded/simpleqa-1111",
  "score": {
    "value": null,
    "status": "unknown",
    "reason": "No rubric configured",
    "evidence": []
  },
  "contributions": [],
  "acceptance": "pass",
  "acceptance_basis": "estimated",
  "gates": [
    {
      "threshold": {
        "metric": "target_correct",
        "op": "==",
        "value": 1
      },
      "decision": "pass",
      "basis": "estimated",
      "reason": "target_correct: 1 == 1 is pass"
    }
  ],
  "reason": "Applied configured rubric and acceptance rule"
}
1 executions · 3 events
[
  {
    "id": "recorded/simpleqa-1111/aggregate",
    "slot": "recorded",
    "retry_index": 0,
    "parent_id": null,
    "status": "completed",
    "started_at": "2026-09-23 15:52:17.585940+00:00",
    "ended_at": "2026-09-23 15:53:56.187197+00:00",
    "effective_config": {
      "value": {
        "model_requested": "gpt-5.6-luna",
        "resolved_model": "Not emitted by Codex CLI JSONL",
        "provider": "Experimental local Codex CLI transport",
        "reasoning": "low",
        "retriever": "DuckDuckGo",
        "embedding": "local BAAI/bge-small-en-v1.5",
        "max_iterations": 3,
        "results_per_query": 5,
        "report_min_words": 600,
        "limitations": "Native temperature and max_tokens settings cannot be enforced by this CLI transport; cost in dollars is unknown."
      },
      "status": "observed",
      "reason": null,
      "evidence": []
    },
    "resources": {
      "cost_usd": {
        "value": null,
        "status": "unknown",
        "reason": "No dollar charge reported by Codex CLI",
        "evidence": []
      },
      "cost_scope": [
        "model"
      ],
      "input_tokens": {
        "value": 41335,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "output_tokens": {
        "value": 2345,
        "status": "observed",
        "reason": null,
        "evidence": []
      },
      "human_minutes": {
        "value": null,
        "status": "unknown",
        "reason": "Not timed",
        "evidence": []
      }
    },
    "role": "main",
    "native_refs": {},
    "error": null
  }
]
[
  {
    "id": "recorded/simpleqa-1111/aggregate/call/1",
    "execution_id": "recorded/simpleqa-1111/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 1,
      "stage": "agent selection",
      "request_sha256": "1d994772a770a03afd2061ded1e98c31e1b7f589f02296d3af5aa42555f6d199",
      "response_sha256": "0a949c1f5dfa3ef9c59951ea827304ce7cd7923628af9d5ff631eb4d741b2751"
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/1/trace/0",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  },
  {
    "id": "recorded/simpleqa-1111/aggregate/call/2",
    "execution_id": "recorded/simpleqa-1111/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 2,
      "stage": "query planning",
      "request_sha256": "3d5f19de7562ade0062225dc85f84354823053e991079a80402b45c61d217c68",
      "response_sha256": "63b6da82b92901f98862a1114f0117a1ee1c152ceb7b1a6554dc74f3dd6c1909"
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/1/trace/1",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  },
  {
    "id": "recorded/simpleqa-1111/aggregate/call/3",
    "execution_id": "recorded/simpleqa-1111/aggregate",
    "kind": "llm.call",
    "at": null,
    "fields": {
      "kind": "llm.call",
      "call_number": 3,
      "stage": "report writing",
      "request_sha256": "b818440c0376c0f42db2f52db957b7db7fa79384d13d397ccc338cc9e741cc95",
      "response_sha256": "561247a4bba652acd6540c96dca9c5ae7d9e7806af689ace1eb51ad225350793"
    },
    "inputs": [],
    "outputs": [],
    "source": {
      "artifact": {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      },
      "locator": "json:/runs/1/trace/2",
      "description": "Curated summary of retained native evidence; full captures omitted"
    }
  }
]
Environment and native references
{
  "value": {
    "os": "Windows",
    "repo": "https://github.com/assafelovic/gpt-researcher",
    "commit": "6f998577d547b1e54ec662dac63583aa11e3b84b"
  },
  "status": "observed",
  "reason": null,
  "evidence": []
}
{
  "case_id": "simpleqa-1111",
  "row_index": "1",
  "public_capture_sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
}

Evidence references

Declared configurations and provenance

gptr-codex-luna-recorded · curated-capture-import / 1
{
  "id": "gptr-codex-luna-recorded",
  "backend": {
    "name": "curated-capture-import",
    "revision": "1"
  },
  "components": {
    "model": {
      "kind": "model",
      "ref": {
        "name": "gpt-5.6-luna",
        "revision": null
      },
      "params": {},
      "content": null
    }
  },
  "settings": {
    "model_requested": "gpt-5.6-luna",
    "resolved_model": "Not emitted by Codex CLI JSONL",
    "provider": "Experimental local Codex CLI transport",
    "reasoning": "low",
    "retriever": "DuckDuckGo",
    "embedding": "local BAAI/bge-small-en-v1.5",
    "max_iterations": 3,
    "results_per_query": 5,
    "report_min_words": 600,
    "limitations": "Native temperature and max_tokens settings cannot be enforced by this CLI transport; cost in dollars is unknown."
  },
  "description": "Three historical failures selected before inference; one trial each; not representative or ranked hardest.",
  "native": null
}
Native jobs and capture projections
[]
[
  {
    "mapper": {
      "name": "example.gpt-researcher-curated",
      "revision": "1"
    },
    "source_format": {
      "name": "curated-gpt-researcher-pilot",
      "revision": "1"
    },
    "sources": [
      {
        "uri": "https://raw.githubusercontent.com/guybass/agent-eval-flow/ba0dd007c082b82f649f8b3c638d641612b42194/examples/data/gpt-researcher/capture.json",
        "media_type": "application/json",
        "sha256": "3d8b2a4839737ceab1122c1861959c9ae59e32bf3636fa39c1347c9e297a7c94"
      }
    ],
    "omitted_fields": [
      "Full prompts and reports",
      "Scraped page bodies",
      "CLI stdout and session identifiers",
      "Local paths",
      "Detailed grading receipts"
    ],
    "issues": []
  }
]

This report displays saved measurements and declared rules. A trace alone does not establish causal attribution. No confidence interval method is recorded in contract 0.4. Evidence links retain their original locations; external availability is not checked or downloaded by this report.