Same standard of work. About 5.1× the energy.
GPT-5.4 scores higher than GLM-5.1 on Epoch's Capabilities Index — 156.86 against 149.71, across 58 benchmarks covering coding, agents, long context and writing — while using about 5.1× less energy per answer. In blind head-to-head voting, people rate them 1470 and 1462 respectively.
| GPT-5.4 | GLM-5.1 | |
|---|---|---|
| Energy per answer | 7.94 Wh | 40.8 Wh |
| Same carbon as driving | 15 metres | 75 metres |
| Energy rank | 92 of 136 | 117 of 136 |
| Capability (science questions) | 93.3 | 89.9 |
| Epoch Capabilities Index | 156.86 | 149.71 |
| Human vote rating | 1470 | 1462 |
| Votes cast | 60,537 | 48,901 |
| Confidence | Day-one estimate | Day-one estimate |
| Released | 2026-03-05 | 2026-04-07 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is a science-question test; the Capabilities Index is a composite over 58 benchmarks; the vote rating comes from blind human comparisons. The full method is here.