Sign in
Benchmarks
Frontend
Dated

There is no trustworthy current leaderboard, so the most recent reliable figure is already months old.

Design2Code

Design2Code tests whether a model can look at a picture of a webpage and reproduce it as working front-end code (HTML and CSS), so the output renders to look like the original.

What it measures

The hard part is not writing valid code but matching the visual design faithfully, layout, spacing, colors, and content, from an image alone.

Example task

A model is shown a screenshot of a real webpage and must generate the HTML and CSS that recreates it; the result is scored on how closely the rendered page matches the original, both overall visual similarity and whether individual elements (text, images, blocks) are present and placed correctly.

Why you should care

Turning a design mockup into working front-end code is a concrete, fast-growing use of AI for anyone building apps or sites, and Design2Code is the named benchmark for it. But its public numbers come from aggregators and labs, not an independent referee, so read them as claims.

How scoring works

A visual-and-structural similarity score (higher is better) over a set of real webpage screenshots; the original benchmark is from a 2024 Stanford-led paper. There is no single maintainer-run live leaderboard; current figures are collected by third-party aggregators from self-reported results.

How to read the numbers

Scores are a similarity percentage, not a pass or fail, so a high number means close to the original, not perfect. Current leaders sit in the 90s while many capable models trail in the 70s, but because the metric rewards rough visual closeness, small differences near the top are not very meaningful.

What to watch

There is no trustworthy current leaderboard for screenshot-to-code, which is why this panel shows a dated anchor instead of a live ranking. The figures that circulate come from aggregators republishing self-reported results, not independent re-evaluations, so treat any Design2Code number you see as a claim. The similarity metric is also a proxy: a page can score well while being subtly wrong in ways the metric misses, and the underlying benchmark dates to 2024, so it may not reflect how current models handle modern, interactive front-ends.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders

No reliable current leaderboard. The most cited figure is a dated anchor:

No current public leaderboard exists for screenshot-to-code. The original Design2Code benchmark (Stanford, 2024) introduced automatic visual-similarity and element-matching metrics over a set of real webpage screenshots; in that study the strongest system was GPT-4V. Every figure circulating since is self-reported through aggregators with no independent referee, so there is no trustworthy current ranking.

StatusDated
View the board