We Built the Same Website With ChatGPT, Claude and Gemini — Here Are the Results
Updated: 2026-08-13 · Author: VITON13 Research · Category: AI Development · Status: Open protocol — no results yet
Direct answer
This page is an open VITON13 research protocol, not a completed result. It fixes the question, sample, controls, metrics and publication threshold before collection. No model winner, conversion lift or performance effect has been measured yet.
Key findings
- No result or winner has been measured.
- The original element is fixed in advance: One frozen brief, retained outputs and an auditable 100-point scorecard.
- Every run must retain inputs, outputs, timestamps and exclusions.
- The page remains noindex until the result package passes review.
Research question and information gain
Research question: How do ChatGPT, Claude and Gemini compare when given one frozen website brief and the same test environment?
Primary intent: Product comparison.
Original contribution: One frozen brief, retained outputs and an auditable 100-point scorecard.
| Research object | Required evidence | Publication rule |
|---|---|---|
| Question | How do ChatGPT, Claude and Gemini compare when given one frozen website brief and the same test environment? | Frozen before collection |
| Original element | One frozen brief, retained outputs and an auditable 100-point scorecard. | Retained with the article |
| Result | Pending first-party collection | Noindex until complete |
Source: VITON13 Research synthesis; individual evidence sources are linked in context.
Methodology
The research unit was defined before drafting: claim, source class, observation date, evidence status, limitation and reviewer note. Product and technical capabilities use primary documentation. VITON13 implementation statements are verified against shipped routes and code; business outcomes are not inferred from feature availability. The experiment remains unexecuted, so the page reports method only.
Dates: research and source review completed 2026-08-13. Vendor features and prices require rechecking at the point of purchase or implementation.
Exclusions: affiliate rankings, unattributed statistics, invented quotations, synthetic user outcomes and undisclosed paid claims.
1. The frozen website brief
This section is pre-registered for definition and controlled setup. The team will record exact inputs, environment versions, timestamps, failures and exclusions before any comparative claim is written. A missing observation remains missing; it is never imputed as a model win or user outcome.
The publication threshold is simple: the retained evidence must support the wording, another reviewer must be able to reproduce the calculation, and the limitations must remain visible beside the result.
2. Scoring rubric
This section is pre-registered for measurement and review. The team will record exact inputs, environment versions, timestamps, failures and exclusions before any comparative claim is written. A missing observation remains missing; it is never imputed as a model win or user outcome.
The publication threshold is simple: the retained evidence must support the wording, another reviewer must be able to reproduce the calculation, and the limitations must remain visible beside the result.
3. Environment and controls
This section is pre-registered for measurement and review. The team will record exact inputs, environment versions, timestamps, failures and exclusions before any comparative claim is written. A missing observation remains missing; it is never imputed as a model win or user outcome.
The publication threshold is simple: the retained evidence must support the wording, another reviewer must be able to reproduce the calculation, and the limitations must remain visible beside the result.
4. Result capture
This section is pre-registered for measurement and review. The team will record exact inputs, environment versions, timestamps, failures and exclusions before any comparative claim is written. A missing observation remains missing; it is never imputed as a model win or user outcome.
The publication threshold is simple: the retained evidence must support the wording, another reviewer must be able to reproduce the calculation, and the limitations must remain visible beside the result.
5. What would count as a winner
This section is pre-registered for measurement and review. The team will record exact inputs, environment versions, timestamps, failures and exclusions before any comparative claim is written. A missing observation remains missing; it is never imputed as a model win or user outcome.
The publication threshold is simple: the retained evidence must support the wording, another reviewer must be able to reproduce the calculation, and the limitations must remain visible beside the result.
Data required from VITON13
- Execute the frozen protocol across every declared condition.
- Retain raw inputs, generated outputs, screenshots or traces, and timestamps.
- Record exclusions before analysis.
- Complete independent result review.
No result has been measured yet. This protocol must not be summarized as a completed experiment.
Limitations
- Vendor documentation establishes supported behavior, not universal outcomes.
- VITON13 implementation evidence describes this codebase and may not generalize to other organizations.
- Rapidly changing models, prices and private previews can make dated details obsolete.
- No comparative or causal conclusion is allowed before data collection.
- English is the primary research language for this programme.
Practical checklist
- Define one decision the page or system must support.
- Link each material claim to the closest primary source.
- Separate shipped capability, observation, interpretation and forecast.
- Keep permissions narrow and reversible.
- Test keyboard, mobile, error and reduced-motion states where interfaces are involved.
- Record dates and update triggers.
Frequently asked questions
What is the direct answer to AI website builder benchmark?
This page is an open VITON13 research protocol, not a completed result. It fixes the question, sample, controls, metrics and publication threshold before collection. No model winner, conversion lift or performance effect has been measured yet.
What evidence does this VITON13 page add?
One frozen brief, retained outputs and an auditable 100-point scorecard.
What has not been proven yet?
The page does not prove hidden ranking factors, universal conversion effects or outcomes outside its stated evidence.
How should a small team use this framework?
Start with the smallest verifiable layer, assign an owner, add an audit trail and test a representative task before scaling.
When will this page be updated?
VITON13 records updates in the manifest and changes the page date only when the evidence or implementation materially changes.
Sources & methodology
- OpenAI — Models and current API capabilities, accessed 2026-08-13.
- OpenAI — Model guidance, accessed 2026-08-13.
- Google AI for Developers — Gemini models, accessed 2026-08-13.
- Model Context Protocol — Specification and trust boundaries, accessed 2026-08-13.
- VITON13 production code and public routes, inspected 2026-08-13.
Editorial disclosure
VITON13 is both the publisher and, for product case studies, the system operator. That conflict is disclosed rather than hidden. No placement in this research programme is sold, and no unfinished result is converted into a marketing claim.
