External Benchmarks
BMO can participate in third-party coding-agent repair benchmarks as a black-box patch-producing CLI.
Benchmark-specific task materialization, prediction files, official harness
execution, grading normalization, and comparison reports should live in an
external project space. BMO’s product contract is the generic agent process:
bmo run, provider/model selection, tool policy, logs, and bounded run
evidence.
External harnesses should record model, budget, task set, allowed tools, BMO version, harness version, patch artifacts, stdout/stderr, and grading provenance. If a harness uses workspace-intelligence or stigmergic comparison lanes, it must inject only bounded non-oracular trace summaries and must not feed gold patches or official grading verdicts back into live attempts.
Use these reports for comparison context only until a reviewed external
benchmark baseline is promoted. Deterministic product confidence remains owned
by task test:eval:readiness.
For maintainer details, see the quality topic for external benchmarks.