Benchmark Methodology
A benchmark that compares AI coding platforms is only worth publishing if it is fair, reproducible and honest about its limits. Vibrr's rules: every platform receives an identical base specification and acceptance criteria; every run records the platform, model, compiler, knowledge and audit versions; comparisons need at least three valid runs per platform; results are descriptive, never a ranking; only synthetic or internal data is used; and a human reviews before anything is published.
The fairness contract
- Same task. One base specification and one set of acceptance criteria per benchmark, verified by comparing a hash of both across every platform's prompt.
- Adaptation only where sourced. The prompt differs by platform only in an adaptation section built from documented or validated knowledge. Nothing is added per platform on a hunch.
- Comparable starting point. Each run starts from the platform's default empty project.
- Everything recorded. Platform version, model and model version where available, prompt, compiler version, knowledge version, audit version, environment and timestamp.
- Iterations and fixes recorded. Corrective prompts and failed executions are counted, not hidden.
Sample size and statistics
A comparison requires at least three valid runs on every platform involved, all using the same base specification and the same audit version. Below that, or if any of those differ, the honest output is that no conclusion has been established. Results are descriptive: the number of runs and the minimum, median and maximum of each metric. Vibrr does not compute significance and does not name a winner, because with few runs a p-value would suggest a precision the data does not have.
Metrics are defined before they are used
Every metric states its definition, formula, inputs and limitations, and a metric without all four cannot be registered. The current definitions, including the ones the first benchmark will use, are on the benchmark index. Scores that exist only to make a page look scientific are not allowed.
What data may be used
Benchmark runs use synthetic or internal projects only. A user's project is never research data by default: it would need an explicit opt-in with a recorded consent identifier, and even then it is excluded from a benchmark run. Text is scanned for credential-shaped strings and instruction-shaped content before it can be published.
Before anything is published
- The promotion gate refuses any claim whose validator is the same as its extractor, that has no source or experiment behind it, or that falls below the confidence threshold.
- A finding must be seen in repeated distinct runs inside a versioned experiment before it can become durable platform knowledge.
- A human reviews the draft; low-confidence findings stay unpublished.
- Documentation changes flag affected pages for review; dates move only when the content materially changes.
Back to Vibrr research or the benchmark index.
Known limitations of this approach
- Platforms differ in kind — a hosted builder, an editor agent and a terminal agent — so the benchmark measures delivered output, not the experience of using each.
- Platform and model versions change; a result describes the recorded versions only.
- The audit only finds what it checks; a clean audit is not proof of correctness.
- Small samples: results are indicative at best.