AI Coding Platform Benchmark: Design and Status
This index lists Vibrr's benchmarks for comparing AI coding platforms on the same specification. The first, a multi-role SaaS build, is designed but has not been run: there are no results yet. What is published is everything needed to judge the design — the base specification, the enumerated requirements, and every metric's definition, formula, inputs and limitations.
Status: infrastructure ready, no results
Benchmark designs
Multi-role SaaS build
Status: designed Valid runs: 0 Conclusion: none established
Research question. When one base specification for a small multi-role SaaS application is compiled for each platform, how do the delivered projects differ in requirements coverage, build effort, authorization correctness and audit findings?
Base specification. Build a multi-tenant task-tracking SaaS with email/password authentication, three roles (owner, member, viewer) with role-based authorization, a PostgreSQL-backed data model, a dashboard, CRUD for projects and tasks, a validated JSON API, responsive UI with loading/empty/error states, automated tests, accessible markup, crawlable public pages, and deployment configuration.
Requirements
| ID | Requirement | How it is checked |
|---|---|---|
| R1 | Users can sign up, sign in and sign out. | End-to-end scripted flow. |
| R2 | Owner, member and viewer roles have distinct permissions enforced by the backend. | Authorization matrix run against the API. |
| R3 | Data is stored in PostgreSQL with a documented schema. | Schema inspection plus a migration run. |
| R4 | Projects and tasks support create, read, update and delete with validation. | API contract tests. |
| R5 | Every data view has loading, empty and error states. | Forced-state screenshots reviewed against a checklist. |
| R6 | The project ships automated tests that run with one documented command. | Run the command on a clean clone. |
| R7 | The project deploys to a preview from a clean clone following its README. | Deploy log. |
Metrics
| Metric | Definition | Formula | Limitations |
|---|---|---|---|
| Requirements coverage | Share of the benchmark's enumerated requirements that a reviewer confirms are implemented and working. | requirementsPassed / requirementsTotal | Depends on how requirements are enumerated and on reviewer judgment for partially working features; requirements are equally weighted. |
| Iterations to a passing build | Number of follow-up prompts or corrections needed before the project builds and its own checks run green. | count of corrective prompts after the initial prompt until build+checks pass | Counts interventions, not effort; a long corrective prompt counts as one. Depends on the operator's judgment of when to stop. |
| Failed build runs | Number of build or test executions that failed before the first fully passing one. | count of failed build/test executions before first pass | Only counts executions that were observable; a platform that hides its internal retries under-reports. |
| Authorization matrix pass rate | Share of role × resource × operation cases in the benchmark's authorization matrix that are correctly allowed or denied by the backend. | casesCorrect / casesTotal | Measures only the enumerated matrix; requires the benchmark to define roles and resources up front. |
| Critical audit findings | Findings of critical severity reported by the pinned Vibrr audit version against the final output. | count(findings where severity = critical) | Bounded by what the audit checks; a clean audit is not proof of absence of defects. Audit version must match across compared runs. |
| High-severity audit findings | Findings of high severity reported by the pinned Vibrr audit version. | count(findings where severity = high) | Same as critical findings; severity is assigned by the audit rubric, not by the platform. |
| Deployment from a clean checkout | 1 if the output builds and deploys to a preview from a clean clone following only its README; otherwise 0. | 1 if deploy succeeds else 0 | Binary, and sensitive to the target host and the operator's environment. |
Design limitations
- Platforms differ in whether they run a model, an editor or a hosted builder; the benchmark measures delivered output, not the experience of using each.
- Prompts differ only where platform adaptation is justified by sourced knowledge; some difference in prompt structure is unavoidable.
- Small samples: results are descriptive and are not a ranking.
- Platform and model versions change; results describe the recorded versions only.
Results
No runs have been executed. Benchmark infrastructure is ready; results will appear here after execution. No experimental conclusion has been established.
How results will be reported when they exist
- Per platform: number of valid runs, and the minimum, median and maximum of each metric — descriptive statistics only.
- No statistical significance testing. With a handful of runs per platform it would be decoration.
- Every result page links the exact benchmark version, platform versions, compiler version, knowledge version and audit version.
- Runs that lack any reproducibility field are excluded and listed as excluded.
- If platforms were not comparable, the page says so instead of comparing them.
The rules behind this are on the methodology page. The wider context is in Vibrr research.