Skip to content

AI Coding Platform Benchmark: Design and Status

This index lists Vibrr's benchmarks for comparing AI coding platforms on the same specification. The first, a multi-role SaaS build, is designed but has not been run: there are no results yet. What is published is everything needed to judge the design — the base specification, the enumerated requirements, and every metric's definition, formula, inputs and limitations.

Written by Vibrr Engineering · Published · Claims last verified · Not independently reviewed

Status: infrastructure ready, no results

Benchmark designs

Multi-role SaaS build

Status: designed Valid runs: 0 Conclusion: none established

Research question. When one base specification for a small multi-role SaaS application is compiled for each platform, how do the delivered projects differ in requirements coverage, build effort, authorization correctness and audit findings?

Base specification. Build a multi-tenant task-tracking SaaS with email/password authentication, three roles (owner, member, viewer) with role-based authorization, a PostgreSQL-backed data model, a dashboard, CRUD for projects and tasks, a validated JSON API, responsive UI with loading/empty/error states, automated tests, accessible markup, crawlable public pages, and deployment configuration.

Requirements

Requirements every platform's output is checked against
IDRequirementHow it is checked
R1Users can sign up, sign in and sign out.End-to-end scripted flow.
R2Owner, member and viewer roles have distinct permissions enforced by the backend.Authorization matrix run against the API.
R3Data is stored in PostgreSQL with a documented schema.Schema inspection plus a migration run.
R4Projects and tasks support create, read, update and delete with validation.API contract tests.
R5Every data view has loading, empty and error states.Forced-state screenshots reviewed against a checklist.
R6The project ships automated tests that run with one documented command.Run the command on a clean clone.
R7The project deploys to a preview from a clean clone following its README.Deploy log.

Metrics

Every metric states its definition, formula and limitations
MetricDefinitionFormulaLimitations
Requirements coverageShare of the benchmark's enumerated requirements that a reviewer confirms are implemented and working.requirementsPassed / requirementsTotalDepends on how requirements are enumerated and on reviewer judgment for partially working features; requirements are equally weighted.
Iterations to a passing buildNumber of follow-up prompts or corrections needed before the project builds and its own checks run green.count of corrective prompts after the initial prompt until build+checks passCounts interventions, not effort; a long corrective prompt counts as one. Depends on the operator's judgment of when to stop.
Failed build runsNumber of build or test executions that failed before the first fully passing one.count of failed build/test executions before first passOnly counts executions that were observable; a platform that hides its internal retries under-reports.
Authorization matrix pass rateShare of role × resource × operation cases in the benchmark's authorization matrix that are correctly allowed or denied by the backend.casesCorrect / casesTotalMeasures only the enumerated matrix; requires the benchmark to define roles and resources up front.
Critical audit findingsFindings of critical severity reported by the pinned Vibrr audit version against the final output.count(findings where severity = critical)Bounded by what the audit checks; a clean audit is not proof of absence of defects. Audit version must match across compared runs.
High-severity audit findingsFindings of high severity reported by the pinned Vibrr audit version.count(findings where severity = high)Same as critical findings; severity is assigned by the audit rubric, not by the platform.
Deployment from a clean checkout1 if the output builds and deploys to a preview from a clean clone following only its README; otherwise 0.1 if deploy succeeds else 0Binary, and sensitive to the target host and the operator's environment.

Design limitations

  • Platforms differ in whether they run a model, an editor or a hosted builder; the benchmark measures delivered output, not the experience of using each.
  • Prompts differ only where platform adaptation is justified by sourced knowledge; some difference in prompt structure is unavoidable.
  • Small samples: results are descriptive and are not a ranking.
  • Platform and model versions change; results describe the recorded versions only.

Results

No runs have been executed. Benchmark infrastructure is ready; results will appear here after execution. No experimental conclusion has been established.

How results will be reported when they exist

  • Per platform: number of valid runs, and the minimum, median and maximum of each metric — descriptive statistics only.
  • No statistical significance testing. With a handful of runs per platform it would be decoration.
  • Every result page links the exact benchmark version, platform versions, compiler version, knowledge version and audit version.
  • Runs that lack any reproducibility field are excluded and listed as excluded.
  • If platforms were not comparable, the page says so instead of comparing them.

The rules behind this are on the methodology page. The wider context is in Vibrr research.