Same standard of work. About 5.2× the energy.
Qwen3.6 27B scores higher than Llama 3.1-70B on Epoch's Capabilities Index — 147.47 against 125.71, across 58 benchmarks covering coding, agents, long context and writing — while using about 5.2× less energy per answer. Qwen3.6 27B has not faced human head-to-head voting, where Llama 3.1-70B carries 55,240 votes — worth weighing.
| Qwen3.6 27B | Llama 3.1-70B | |
|---|---|---|
| Energy per answer | 0.351 Wh | 1.84 Wh |
| Same carbon as driving | under 1 m | 3 m |
| Energy rank | 38 of 134 | 64 of 134 |
| Capability (science questions) | 84.8 | 44.2 |
| Epoch Capabilities Index | 147.47 | 125.71 |
| Human vote rating | — | 1261 |
| Votes cast | — | 55,240 |
| Confidence | Calculated | Calculated |
| Released | 2026-04-22 | 2024-07-23 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is a science-question test; the Capabilities Index is a composite over 58 benchmarks; the vote rating comes from blind human comparisons. The full method is here.