Skip to content

External Benchmarks

BMO can participate in third-party coding-agent repair benchmarks as a black-box patch-producing CLI.

Benchmark-specific task materialization, prediction files, official harness execution, grading normalization, and comparison reports should live in an external project space. BMO’s product contract is the generic agent process: bmo run, provider/model selection, tool policy, logs, and bounded run evidence.

External harnesses should record model, budget, task set, allowed tools, BMO version, harness version, patch artifacts, stdout/stderr, and grading provenance. If a harness uses workspace-intelligence or stigmergic comparison lanes, it must inject only bounded non-oracular trace summaries and must not feed gold patches or official grading verdicts back into live attempts.

Use these reports for comparison context only until a reviewed external benchmark baseline is promoted. Deterministic product confidence remains owned by task test:eval:readiness.

For maintainer details, see the quality topic for external benchmarks.