Same standard of work. About 20× the energy.
Qwen3.6 27B scores higher than GPT-4.1 on Epoch's Capabilities Index — 147.47 against 137.52, across 58 benchmarks covering coding, agents, long context and writing — while using about 20× less energy per answer. Qwen3.6 27B has not faced human head-to-head voting, where GPT-4.1 carries 49,920 votes — worth weighing.
| Qwen3.6 27B | GPT-4.1 | |
|---|---|---|
| Energy per answer | 0.351 Wh | 6.90 Wh |
| Same carbon as driving | under 1 m | 13 m |
| Energy rank | 38 of 134 | 85 of 134 |
| Capability (science questions) | 84.8 | 66.9 |
| Epoch Capabilities Index | 147.47 | 137.52 |
| Human vote rating | — | 1383 |
| Votes cast | — | 49,920 |
| Confidence | Calculated | Day-one estimate |
| Released | 2026-04-22 | 2025-04-14 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is a science-question test; the Capabilities Index is a composite over 58 benchmarks; the vote rating comes from blind human comparisons. The full method is here.