Same standard of work. About 11× the energy.
GPT-5.4 scores higher than Claude Opus 4.1 on Epoch's Capabilities Index — 156.86 against 144.11, across 58 benchmarks covering coding, agents, long context and writing — while using about 11× less energy per answer. In blind head-to-head voting, people rate them 1470 and 1418 respectively.
| GPT-5.4 | Claude Opus 4.1 | |
|---|---|---|
| Energy per answer | 7.64 Wh | 86.7 Wh |
| Same carbon as driving | 14 metres | 160 metres |
| Energy rank | 200 of 266 | 257 of 266 |
| Capability (58 benchmarks) | 156.9 | 144.1 |
| Human vote rating | 1470 | 1418 |
| Votes cast | 60,537 | 75,913 |
| Confidence | Day-one estimate | Day-one estimate |
| Released | 2026-03-05 | 2025-08-05 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is Epoch AI's Capabilities Index, one score combining 58 separate benchmarks; the vote rating comes from blind human comparisons. The full method is here.