Email Bench
Evaluates language models on producing production-ready emails from real briefs, in HTML and React Email.
200 tasks · 100 HTML · 100 React Email
| Model | ||||
|---|---|---|---|---|
| 1 | GPT-6 Astra | 50.5% | $0.64 | 124,773 |
| 2 | GPT-6 Sol | 47.5% | $0.12 | 109,496 |
| 3 | Claude Opus 5.5 | 36.5% | $0.44 | 260,365 |
| 4 | GPT-6 Luna | 36.0% | $0.006 | 117,770 |
| 5 | Grok 4.7 | 33.5% | $1.20 | 1,653,584 |
| 6 | Claude Fable 5.1 | 25.0% | $0.99 | 261,027 |
| 7 | Muse Spark 1.3 | 18.0% | $0.15 | 197,504 |
| 8 | Gemini 3.8 Flash | 15.5% | $0.82 | 3,368,863 |
| 9 | DeepSeek V4.1 Flash | 13.0% | $0.02 | 733,785 |
| 10 | GLM 5.3 | 6.0% | $0.37 | 925,640 |
| 11 | GLM 5.3 Flash | 4.0% | $0.02 | 283,080 |
Every model runs each task once through the Pi coding agent, at high reasoning effort. A run that produces no email counts as a failure. Cost and tokens are averages per task.