Fuse Intelligence · Community

Fuse Benchmarks

Real capability tests on models and providers — not vibes. Each run stores the HTML, CSS, and JavaScript the model produced, plus a hand-assigned Fuse Score from 0–1000. Models with multiple tests show an average score.

1 models shown
0–1000 Fuse Score scale
Live output HTML/CSS/JS preserved
Clear

Leaderboard

Sorted by average Fuse Score across all recorded showcase tests.

  1. Cursor Closed source 761 Fuse Score

    Composer 2.5 Fast

    composer-2.5-fast

    Agentic coding model benchmarked on the same widget-native card-game prompt.

    #cursor #closed-source #tool-use
    1 test Strong tier

What is a Fuse Score?

A 0–1000 ranking assigned after reviewing each test run — code quality, correctness, polish, and how well the output matches the prompt. It is not an automated metric; it reflects real harness evaluation.

Why store the output?

Every benchmark keeps the model’s HTML, CSS, and JavaScript so you can inspect what it actually built — not just a number on a leaderboard.

More tests coming

Card-game UI is the first harness. Widget layout, agent tool-use, and marketplace widget generation tests will appear here as they are run.