Same standard of work. About 7.0× the energy.
GPT-5.3 Codex scores higher than Gemini 3.1 Pro on Epoch's Capabilities Index — 156.58 against 155.0, across 58 benchmarks covering coding, agents, long context and writing — while using about 7.0× less energy per answer. GPT-5.3 Codex has not faced human head-to-head voting, where Gemini 3.1 Pro carries 106,951 votes — worth weighing.
| GPT-5.3 Codex | Gemini 3.1 Pro | |
|---|---|---|
| Energy per answer | 7.11 Wh | 49.8 Wh |
| Same carbon as driving | 13 metres | 92 metres |
| Energy rank | 195 of 266 | 245 of 266 |
| Capability (58 benchmarks) | 156.6 | 155.0 |
| Human vote rating | — | 1480 |
| Votes cast | — | 106,951 |
| Confidence | Day-one estimate | Day-one estimate |
| Released | 2026-02-05 | 2026-02-19 |
Energy is calculated the same way for every model on this site, from public data and calibrated against laboratory measurement — accurate to roughly a factor of two to three, which is why we only ever state a difference of 5× or more. Capability is Epoch AI's Capabilities Index, one score combining 58 separate benchmarks; the vote rating comes from blind human comparisons. The full method is here.