Same standard of work. About 6.7× the energy.
GPT-5.4 scores higher than Kimi K2.5 on Epoch's Capabilities Index — 156.4 against 148.19, across 58 benchmarks covering coding, agents, long context and writing — while using about 6.7× less energy per answer. In blind head-to-head voting, people rate them 1470 and 1446 respectively.
| GPT-5.4 | Kimi K2.5 | |
|---|---|---|
| Energy per answer | 7.94 Wh | 52.8 Wh |
| Same carbon as driving | 15 m | 97 m |
| Energy rank | 91 of 134 | 118 of 134 |
| Capability (science questions) | 93.3 | 87.6 |
| Epoch Capabilities Index | 156.4 | 148.19 |
| Human vote rating | 1470 | 1446 |
| Votes cast | 60,614 | 70,610 |
| Confidence | Day-one estimate | Day-one estimate |
| Released | 2026-03-05 | 2026-01-27 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is a science-question test; the Capabilities Index is a composite over 58 benchmarks; the vote rating comes from blind human comparisons. The full method is here.