Same standard of work. About 5.3× the energy.
Inkling-Small scores higher than GPT-4 Turbo (Nov 2023) on Epoch's Capabilities Index — 150.19 against 126.2, across 58 benchmarks covering coding, agents, long context and writing — while using about 5.3× less energy per answer. In blind head-to-head voting, people rate them 1414 and 1271 respectively.
| Inkling-Small | GPT-4 Turbo (Nov 2023) | |
|---|---|---|
| Energy per answer | 4.97 Wh | 26.5 Wh |
| Same carbon as driving | 9 m | 49 m |
| Energy rank | 78 of 134 | 114 of 134 |
| Capability (science questions) | 88.5 | 42.4 |
| Epoch Capabilities Index | 150.19 | 126.2 |
| Human vote rating | 1414 | 1271 |
| Votes cast | 13,438 | 98,114 |
| Confidence | Calculated | Day-one estimate |
| Released | 2026-07-15 | 2024-01-25 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is a science-question test; the Capabilities Index is a composite over 58 benchmarks; the vote rating comes from blind human comparisons. The full method is here.