---
title: "Benchmarking tokens per watt: how AI inference energy efficiency is measured"
slug: tokens-per-watt-benchmarks
question: What benchmarks and measurement methodologies exist for quantifying AI inference energy efficiency in tokens per watt (or per joule), from accelerator to data-centre level, and what do they establish?
status: published
depth: deep
created: 2026-08-17
updated: 2026-08-17
summary: "The literature on measuring AI inference efficiency in tokens-per-watt terms is young (mostly 2023-2026) and fragmented: one consortium standard exists at the system level (MLPerf Power), but the benchmarks that actually report tokens per watt or joules per token are research tools with incompatible measurement boundaries. Measured numbers span orders of magnitude — roughly 3-4 joules per output token for a 65B model on A100s, 0.002 to 2.9 kWh per 1,000 inferences depending on task, a 65x spread across models in commercial data centres, and a proposed 1/W law under which tokens per watt halves each time the context window doubles. No retrieved benchmark measures tokens per watt at the data-centre (facility) level; facility efficiency is still expressed as PUE, so wall-level tokens per watt is derived, not measured. Confidence is moderate: 20 of 65 sources were read in full text and several prominent items were unreachable in-session."
disciplines: ["computer systems", "energy systems", "machine learning"]
tags: ["tokens per watt", "energy per token", "MLPerf Power", "LLM inference energy", "power measurement", "PUE", "carbon per query", "benchmarking"]
source_count: 65
year_range: [2009, 2026]
confidence: moderate
search:
  databases: ["OpenAlex", "Crossref", "arXiv API", "arXiv abs pages", "Semantic Scholar (best-effort)", "Unpaywall", "ACL Anthology", "Zenodo and GitHub (AI Energy Star probe)"]
  queries:
    - "tokens per watt"
    - "tokens per joule"
    - "tokens per kilowatt hour"
    - "energy per token large language model"
    - "LLM inference energy benchmark"
    - "MLPerf power energy efficiency machine learning"
    - "AI energy star benchmark"
    - "benchmark energy consumption LLM inference"
    - "energy efficiency large language model inference benchmark"
    - "inference energy measurement transformer"
    - "GPU power consumption measurement deep learning inference"
    - "power metering neural network inference"
    - "energy accounting GPU fine-grained"
    - "power disaggregation data center application"
    - "wattmeter deep learning training measurement"
    - "AI data center energy efficiency PUE"
    - "generative AI data center energy consumption"
    - "machine learning inference energy data center"
    - "data center energy efficiency metric AI workload"
    - "carbon footprint generative AI inference energy"
    - "ChatGPT energy consumption estimate"
    - "energy consumption generative AI models"
    - "sustainable machine learning inference energy"
    - "energy efficient transformer inference"
    - "power measurement large language model"
    - "speculative decoding energy efficiency"
    - "LLM inference energy scaling"
    - "green AI benchmark"
    - "MLPerf Inference Benchmark"
    - "MLPerf Tiny Benchmark"
    - "Power Hungry Processing: Watts Driving the Cost of AI Deployment"
    - "Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model"
    - "The growing energy footprint of artificial intelligence"
    - "Carbon Emissions and Large Neural Network Training"
    - "Energy and Policy Considerations for Deep Learning in NLP"
    - "Green AI"
    - "Sustainable AI: Environmental Implications, Challenges and Opportunities"
    - "LLMCarbon: Modeling the end-to-end carbon footprint of large language models"
    - "A Survey on Green Deep Learning"
    - "EcoServe: Designing Carbon-Efficient Inference Serving for Large Language Models"
    - "Identifying Shades of Green: The SPECpower Benchmarks"
    - "Sustainable AI: Environmental Implications, Challenges and Opportunities (Crossref/arXiv/OpenAlex)"
    - "EcoServe (arXiv ti)"
    - "Green500 (Crossref/arXiv/OpenAlex)"
    - "AI Energy Star (Zenodo/arXiv/OpenAlex/GitHub — no records)"
    - "Weight-Only Quantization Does Not Always Save Energy (arXiv ti — none)"
    - "Power Hungry Processing Watts Driving (arXiv ti)"
    - "DynamoLLM (arXiv ti)"
    - "The 1/W Law (pool resolution)"
  last_run: 2026-08-17
---


## Summary

The literature on measuring AI inference efficiency in tokens-per-watt terms is young and fragmented. One consortium standard exists — MLPerf Power, which collected 1,841 measurements across 60 systems from microwatts to megawatts [1] — but its workload set predates the generative-LLM era, and the benchmarks that actually report tokens per watt or joules per token are research tools with incompatible boundaries: per-GPU, per-node, or per-request; with or without power meters; over different workloads [8][18][41]. The measured numbers span orders of magnitude — about 3–4 joules per output token for a 65B model on A100s [18], 0.002 to 2.9 kWh per 1,000 inferences depending on task [41], a 65× spread across models in commercial data centres [31], and a proposed "1/W law" under which tokens per watt halves every time the context window doubles [64]. No retrieved benchmark measures tokens per watt at the data-centre level; facility efficiency is still expressed as PUE, so wall-level tokens per watt is derived from server numbers, not measured [37]. Confidence is moderate: 20 of 65 sources were read in full text, and several prominent items — including the widely cited Joule estimate paper and two SSRN preprints — were unreachable in-session.

## Why this question

Tokens per watt has become the efficiency currency of AI infrastructure: accelerator purchases, serving-stack choices, power-capping decisions, and data-centre designs are argued in these units. But a unit is not a measurement. What the metric denotes (one GPU or one node or one facility; decode-only or end-to-end; one request length or another), how it is taken (a calibrated wattmeter, a software power model, or vendor telemetry), and on which workload it is taken all change the number by orders of magnitude. An architect comparing H100, B200, and TPU efficiency figures, or deciding whether power capping pays for a decode-heavy workload, is comparing numbers produced by different metrologies — and the published benchmarks are the only place those metrologies are described well enough to be compared.

What turns on the answer is expensive. The prior review in this series (model-to-grid AI datacenter efficiency) established what the efficiency *levers* are and how strong the evidence for each is. This review asks the measurement question underneath it: when a number like "tokens per watt" is reported — in a paper, a submission, or a vendor slide — what was actually measured, and can it be compared with any other number?

## Scope and methods

**Question.** What benchmarks, metrics, and measurement methodologies exist for quantifying AI inference energy efficiency in tokens per watt (or per joule), from accelerator to data-centre level, and what do they establish about how inference efficiency should be measured and compared?

**Inclusion criteria.** Works published or posted 2019–2026 (the modern inference-efficiency literature), plus two metric-antecedent sources from 2009 and 2018 (SPECpower, Green500) because the modern metric descends from them; English; peer-reviewed papers and arXiv preprints (this fast-moving field lives partly in preprints); benchmark descriptions, measurement studies, energy- and carbon-per-inference quantifications, and metric or methodology critiques. **Exclusion criteria.** Training-only efficiency studies except where they define the metric landscape (Zeus, Patterson, Strubell are included for the measurement lineage, not as tokens/watt evidence); power-grid engineering; efficiency levers (power capping, scheduling, serving configuration) as such — already covered by the model-to-grid review — except where they report tokens/watt measurements; microarchitecture-only work; works unreachable through any open archive (see below).

**Databases and dates.** OpenAlex, Crossref, the arXiv API, arXiv abs pages, Semantic Scholar (best-effort; persistent 429s), Unpaywall for OA checks, the ACL Anthology for abstract rescue, and Zenodo/GitHub for a vendor-benchmark probe. Forty-eight query strings (listed in frontmatter) plus a 13-title landmark-resolution ladder (Crossref `query.bibliographic` → arXiv `ti:`) and citation-graph follow-ups. Searches run 2026-08-17.

**Screening counts.** 1,948 raw records retrieved (1,142 from OpenAlex, 806 from arXiv/Crossref); ~1,566 after DOI/title deduplication; 210 passed an automated topic-and-venue screen; 13 landmark titles were resolved through the resolution ladder (12 resolved cleanly, one resolved to a different but real paper that was kept on its own merits); 66 records were curated into the final selection. One source was dropped at retrieval — the Joule commentary "The growing energy footprint of artificial intelligence" (de Vries) — because every legitimate channel (cell.com, ScienceDirect, the institutional repository, OpenAlex, Crossref, Semantic Scholar) is bot-walled or abstract-less in-session; per this review's protocol, unfetchable sources are not cited. Two SSRN preprints were dropped at screening (a weight-only-quantization energy study and "Green My LLM") because SSRN blocks scripted access. A probe for the vendor benchmark "AI Energy Star" found no paper, DOI, Zenodo record, or GitHub repository — it exists only as marketing material, which is not citable under this review's rules; the absence is noted in Gaps. **65 sources included**, all retrieved this session: 20 read in full text, 45 at abstract level (flagged `abstract-only` and hedged accordingly). Every DOI was verified against Crossref (title similarity 1.00 on all checks; no retraction notices) or, for arXiv DOIs, by abs-page resolution; every `10.48550/arxiv.*` DOI was confirmed to resolve.

**What this review deliberately does not cover.** Training-energy benchmarking as a topic (covered elsewhere in this series); the efficiency-lever evidence base (power capping, scheduling — see the model-to-grid review); vendor marketing figures with no citable artifact.

## The landscape

The literature has three neighbourhoods that barely cite each other. The **consortium-benchmark neighbourhood** (MLPerf family, SPEC, Green500) defines what "efficiency benchmark" means institutionally: rules, submissions, comparability, and rankings in FLOPS/watt or queries-per-watt. The **research-measurement neighbourhood** (2023–2026) is where tokens-per-watt-style numbers actually live: wattmeter and telemetry studies of LLM inference on specific GPUs, serving engines, and cloud APIs. The **carbon-accounting neighbourhood** (FAccT, lifecycle-analysis venues) measures the same physical quantity but reports it in grams of CO₂ per query or per day, folding in grid carbon intensity and embodied manufacturing.

A vocabulary finding shapes everything below: the exact phrase "tokens per watt" appears rarely in the peer-reviewed literature. The field speaks in joules per token [8][26], watt-hours per request [29], kWh per 1,000 inferences [41], and tokens per watt in the few papers that use the unit directly [64]. These are different quantities — a per-GPU decode-phase figure is not comparable with an end-to-end per-request figure — and converting between them requires request lengths, batch sizes, and measurement boundaries that most papers report incompletely. The unit heterogeneity is itself a comparability barrier, and no benchmark in the retrieved set standardises the conversion.

## Theme 1: The metric lineage — from FLOPS per watt to energy per token

The efficiency metric did not start with language models. Green500 institutionalised FLOPS/watt rankings for supercomputers, with measurement-qualification rules that labs must satisfy to be listed [63]. SPEC established SPECpower_ssj2008 as the first industry-standard server power-and-performance benchmark [6]. The "Green AI" agenda then made efficiency a first-class evaluation criterion for machine learning, documenting an estimated 300,000× growth in deep-learning compute from 2012 to 2018 [65]. Strubell, Ganesh and McCallum supplied the first systematic energy accounting for NLP — including inference, not just training — and the policy recommendations that followed [44].

The MLPerf family carries this lineage into machine learning. The training suite was characterised in a peer-reviewed study of its workloads versus earlier benchmarks such as DAWNBench and DeepBench [2]; the inference benchmark spans three orders of magnitude in system power and five in performance, and its closed division enforces latency constraints precisely because throughput-only comparisons mislead [3]. The mobile and tiny variants extended the family down the power axis — MLPerf Tiny made energy a first-class reported metric alongside accuracy and latency [4][5].

Against this background, the token-based unit crystallised in 2024–2026. A workshop paper argued that Energy-per-Token should complement accuracy benchmarks, showing that test-time-compute strategies can trade enormous energy for accuracy (chain-of-thought on MATH raised accuracy by 281% at a reported +15,132% energy cost) and that smaller models with controlled reasoning beat larger models on an energy-accuracy frontier (a 40–60% energy saving at similar or better accuracy) [26]. TokenPowerBench formalised joules-per-token as its headline metric [8], and the 1/W law paper defined tokens/watt explicitly, decomposing it into single-GPU and fleet-level quantities [64]. A survey of green deep learning catalogued the same landscape from the training side and flagged the recurring gap between FLOPs-based and runtime-based efficiency claims [49], and an analysis of inference-energy trends argued that inference energy follows different laws than the performance-versus-parameter scaling curves [27].

## Theme 2: The benchmarks that measure inference efficiency

**MLPerf Power** is the closest thing to a standard. Developed by a consortium of more than 20 organisations, it establishes rules and best practices for measuring ML-system energy efficiency from microwatts to megawatts, and reports 1,841 reproducible measurements across 60 systems [1]. Its significance is institutional: it is the only benchmark in this review whose numbers are produced under agreed rules rather than per-lab conventions. Its limitation is scope: its workloads are the classic MLPerf inference set, not generative-LLM serving, so tokens/watt for chat workloads is outside its remit.

The MLPerf-based evaluations that do exist are instructive about comparability. One study measured four accelerators (NVIDIA A100, Intel Gaudi, Graphcore Bow-Pod64, Groq LPU) on two MLPerf workloads to a common target accuracy: which accelerator is most efficient depends on the workload — Gaudi2 had the lowest energy on ResNet50, the A100 on BERT-Large inference [11]. A 2026 study of 739 MLPerf Inference v5.0 submissions found no significant public-versus-private-cloud efficiency difference but large accelerator-model differences (the H200 showing roughly 45% higher median throughput-per-TDP efficiency), concluding efficiency scales multiplicatively across accelerator generations [10]. Both results cut against any single "best accelerator" narrative.

The research benchmarks report the tokens-per-watt-style numbers. **TokenPowerBench** measures GPU-, node-, and system-level power without specialised meters, attributes energy to prefill and decode phases per request, and reports joules per token across four model families from 1B to 405B parameters, with batch size, context length, parallelism, and quantisation as configurable dimensions [8]. **GPU-NEST** established an energy-efficiency characterisation methodology for multi-GPU inference servers and showed power as a first-order constraint in them [7], and **WattWiser** targeted power- and resource-efficient scheduling for multi-model multi-GPU inference servers [25]. **From Words to Watts** benchmarked LLaMA models on V100 and A100 clusters: a 65B model consumed roughly 3–4 joules per output token at 512-token generation, and power-capping the GPUs from 250 W to 175 W cut total energy by 23.21% at a 6.7% time penalty [18]. **How Hungry is AI** benchmarked 30 models in commercial data centres through their APIs, combining public performance data with environmental multipliers: the most energy-intensive models exceeded 29 Wh per long prompt — over 65× the most efficient — and a 0.42 Wh short query, scaled to 700 million queries per day, aggregates to annual electricity comparable to 35,000 US homes [31]. A NAACL 2025 study benchmarked inference energy across NLP tasks and found energy correlates strongly with output-token length and response time, with quantisation, optimal batch sizes, and prompt phrasing as significant levers [30]. An IEEE Access 2024 profiling study reached a similar conclusion from the architecture side: model size, layer count, parallelised attention, and vocabulary size drive inference energy, while batch size and quantisation level are the tunable levers [19]. A cluster-redesign study for GPT-Neo reported substantial throughput, latency, and energy improvements from a novel interconnect and architecture, though without retrievable magnitudes [20].

A second strand benchmarks the serving engines rather than the models. One study decomposed LLM inference engines (vLLM, TensorRT-LLM, DeepSpeed) on a two-H100 node into setup and token-generation stages to locate energy bottlenecks [9]. Energy-cost models extend this: a workload-based model fitted energy and runtime across heterogeneous CPU-GPU systems with R² above 0.96 [22], and a token-count-aware allocation strategy across heterogeneous clusters saved 7.5% of energy versus a workload-unaware baseline [59]. On the edge, 28 quantised LLMs were measured on a Raspberry Pi 4 [32], an energy-aware DVFS-plus-speculative-decoding scheme reported a 52.4% energy reduction on-device [24], and measurements on Apple Silicon found energy per token grows monotonically but weakly sublinearly with prompt length [13] — with Apple Silicon reported at 4–7× less energy per request than NVIDIA-based servers on interactive workloads in one comparison [61]. Model-level energy estimates round out the set: aggressive weight-only quantisation cut Llama-2-7B inference energy by 4.6× and moved the bottleneck to attention [23], and pre-storing attention matrices cut energy 1.45–2.83× against baselines [62].

## Theme 3: The metrology — how the numbers are taken

Underneath the benchmarks sits a measurement-methodology literature, and it is where the comparability problems live. Physical power meters are accurate but expensive and cannot attribute energy to individual services; software-based power meters (power models plus vendor interfaces) are deployable at scale, and a 2023 study compared the leading tools for CPU and GPU measurement, concluding the choice of tool materially changes what you can claim [14]. A review of energy-estimation approaches catalogued the same trade-off for machine learning generally [52], and a 2025 MLOps study applied software-based power measurement across discriminative and generative pipelines, concluding architecture and hardware choices dominate pipeline energy [53]. An energy-consumption-index proposal aims to standardise cross-architecture DL comparison, though it has been validated on CNNs rather than LLMs [54]. On the attribution side, WattScope estimates application-level power from a server's aggregate draw without OS access, exploiting the low-variability, periodic power signatures of data-centre workloads [16], and a EuroSys 2026 paper tackles the harder version — attributing GPU power to individual jobs in shared cloud settings [15].

Phase attribution is the emerging standard for LLM-specific measurement. TokenPowerBench aligns metrics to prefill versus decode per request [8]; the engine study separates setup from token generation [9]; and the CCGrid 2026 characterisation shows why phase matters — decode-phase frequency scaling from 2842 MHz to 180 MHz saves 42% of energy at only 1–6% latency cost, while prefill behaves differently [12]. Scale matters for methodology too: the BUTTER-E dataset, 63,527 watt-metered runs, found GPU training used a median 9.47 mJ per datum versus 6.16 mJ on CPU in its setting, and that a fitted energy model predicts GPU energy within ±2.74% but CPU energy only within ±20.2% [17].

Reporting practice lags all of this. A 2020 audit found that only about 1 in 100 sampled NeurIPS papers reported energy at all, and proposed a systematic reporting standard; the audit's own experiments consumed 24.344 kWh and 8.021 kg CO₂eq [48]. The configuration space is the final confound: power capping is simultaneously an optimisation lever and a measurement choice. The same GPU measured at different caps produces different tokens-per-watt — 250 W versus 175 W changed total energy by 23.21% on one workload [18], and running at 1.6 GHz instead of 2.0 GHz held throughput constant at under 80% of the energy [28] — so an efficiency number without its power configuration is almost meaningless. Idle power compounds this: the 1/W law analysis shows the B200's 430 W idle draw becomes a dominant fraction of the bill at long context, where KV-cache limits cut concurrency [64].

## Theme 4: What the measurements show

The verified numbers cluster into four findings.

**First, per-token energy at the server level is in the single-digit joules on current data-centre GPUs, with a wide spread.** A 65B LLaMA on A100s: roughly 3–4 J per output token [18]. Llama-2-70B serving: 0.77 to 13.21 Wh per request depending on request class, tensor parallelism, and frequency [29]. Per query, 2.9 J for a 1B model rising to 21.0 J for a 32B model — a 7.2× model-size spread [12].

**Second, the task and workload spread dwarfs the hardware spread.** Per 1,000 inferences, energy ranges from 0.002 kWh for text classification to 2.9 kWh for image generation — a factor above 1,450 — with multi-purpose generative models orders of magnitude costlier than task-specific ones [41]. In commercial data centres, the model spread is 65× on long prompts [31]. Anyone comparing "tokens per watt" across studies is comparing numbers that vary more with workload than with hardware generation.

**Third, inference, not training, dominates steady-state energy — which is the entire justification for the metric.** Industry reports cited by TokenPowerBench put inference above 90% of total power consumption [8]; a comparison study estimated GPT-4's training at ~9,450 MWh versus inference exceeding 500 MWh per day [51]; BLOOM's API deployment was estimated at ~19 kg CO₂eq per day versus 24.7–50.5 tCO₂eq for training [42]; and the serving phase has now surpassed training in energy terms per one 2024 assessment [58].

**Third-and-a-half, model size and accuracy do not move together on energy.** Across 14 open-source LLMs (6B–34B) on an A100, larger models often consumed substantially more energy without proportional accuracy gains, with mid-sized models matching larger ones at much lower energy [33]. Prompt-engineering choices also change the number: prompt phrasing measurably altered Llama 3's inference energy on code generation [55].

**Fourth, the measured levers are large but regime-dependent.** Power/frequency capping: 23.21% energy saved at 6.7% time cost [18], ~20% power reduction with zero latency impact [28], 42% via decode-phase DVFS [12], and near-zero impact at a supercomputing centre where jobs were not power-bound [36]. Routing small queries to small models: 80% energy saving at a 6.8-point quality cost, 88% combined with DVFS [12]. Cluster-level design: 53% energy, 38% operational carbon, and 61% customer-cost reduction under latency SLOs in DynamoLLM's evaluated design [29].

A related but distinct literature reports carbon per query, which is the same measurement problem with a grid-intensity multiplier: LLMCarbon models end-to-end carbon within ≤8.2% of measured footprints [46]; EcoServe's carbon-aware design cut emissions by up to 47% with embodied carbon exceeding 50% of lifetime in its analysis [56]; Clover reduced inference-service carbon under SLA constraints [57]; and cloud-side estimates argued that provider-reported software carbon intensity is the missing prerequisite for mitigation [47]. Patterson et al. quantified the upstream lever space — efficiency and siting together can reduce training carbon by 100–1,000× [43] — and a longitudinal analysis found embodied carbon roughly equal to operational footprint and ~20% operational-power reduction per six months of hardware improvement, at an assumed PUE of 1.1 [45]. A workload model of ChatGPT-class inference compared request-routing policies — local, balanced, and carbon-minimising — for their power and carbon impacts [34], and a life-cycle analysis of LLM-powered chatbots identified eight energy- and carbon-relevant phases with three strategic mitigation pathways [50]. At the serving-system level, a cloud-scale characterisation of ML serving reported up to 28.3% energy savings from hierarchical GPU power management across 105 servers [21].

## Theme 5: The facility-level gap — tokens per watt for data centres

The review's central negative finding: **no retrieved benchmark or measurement study reports tokens per watt at the data-centre level.** The facility literature measures efficiency in PUE — the ratio of total facility energy to IT energy. Studies tune PUE with sensor-plus-ML methodology [39], model PUE statistically across locations [38], measure a small data centre end-to-end including cooling and demand response [40], and — most tellingly — an integrative review asks whether PUE is still fit for purpose for AI infrastructure at all [37]. Production-scale power trends are tracked at the facility level (NERSC's machines stayed well below provisioned power even after GPU transitions [35]), and supercomputing centres manage power draw through capping [36].

The bridge between tokens and facilities exists only in fragments. MLPerf Power spans microwatts to megawatts but measures systems, not facilities [1]. One 2026 preprint treats tokens served as the dispatchable unit for data-centre demand response, cutting operating cost 34.3% without curtailing token volume [60]. And the facility metric that would complete the chain — tokens per watt at the wall, i.e. server-level tokens/watt divided by facility overheads — appears nowhere as a measured quantity. Every published number in this review stops at the server or node boundary; the data-centre boundary is reached only by arithmetic.

## Where the evidence disagrees

**What the metric should optimise.** The energy-per-token advocates argue efficiency belongs alongside accuracy in benchmark scores [26]; the MLPerf tradition treats efficiency as one dimension of a comparability regime built for performance [1]; the carbon literature argues neither energy nor performance is the right objective — grid intensity and embodied carbon are [56][57][47]. These are disagreements about objectives, not numbers, but they change which benchmark you build.

**Does quantisation save energy?** Yes, measured 4.6× on Llama-2-7B with weight-only quantisation [23], and quantised edge models are the default for efficiency [32]. But the one peer-reviewed-adjacent counter-evidence — an empirical study arguing weight-only quantisation does not always save energy across NVIDIA platforms — exists only as an SSRN preprint that was unreachable in-session and is therefore not cited here. The tension is real in the field and unresolved in this review.

**Which engine or accelerator wins?** Rankings do not transfer across workloads: Gaudi2 beats A100 on ResNet50 energy but loses on BERT inference [11]; engine efficiency depends on whether you measure setup or generation phase [9]; and the accelerator-generation effect dominates any cloud-type effect in the MLPerf submission record [10]. Any single ranking is an artefact of its workload.

**Power capping: free lunch or nothing?** Large savings on some workloads [18][28], 42% on decode with 1–6% latency cost [12], and minimal impact at a centre whose jobs were not power-bound [36]. The disagreement tracks workload memory-boundedness — the same regime dependence the model-to-grid review found — and reinforces that tokens-per-watt numbers are configuration-relative.

**Energy versus carbon.** Optimising energy is not optimising carbon: the grid's marginal intensity, not the GPU's watt draw, decides emissions [56][57][43][47]. Studies that report both (DynamoLLM's 53% energy versus 38% carbon savings [29]) show the gap explicitly.

## Gaps and open questions

- **No facility-level tokens/watt benchmark exists.** Nothing in the retrieved literature measures tokens per watt at the data-centre boundary. Settling this requires a consortium-standard benchmark that measures LLM serving at node and facility level with a fixed boundary, workload mix, and phase attribution — the MLPerf Power model extended to generative serving and to the wall.
- **No consortium benchmark covers generative-LLM serving energy.** MLPerf Power's workloads predate it; TokenPowerBench is research-grade; the commercial-DC numbers come from API-side estimation, not independent measurement [31]. The field is benchmarking around the edges of the workload that matters.
- **No independent cross-cloud measurement.** The largest comparative dataset (739 submissions) is vendor-submitted [10]; the cloud-carbon literature argues providers must publish software carbon intensity for anyone to act [47].
- **Measurement boundaries are unstandardised.** GPU-only versus node versus system; idle power included or excluded; phase attribution present or absent; power configuration reported or not. Henderson's proposed reporting standard [48] has not been adopted by any major benchmark.
- **Vendor benchmarks are uncitable.** The "AI Energy Star"-style figures circulate in marketing with no paper, DOI, or repository; this review can neither verify nor cite them, and the ecosystem's reliance on them is itself a gap.
- **Idle power and long-context behaviour are under-measured.** The 1/W law [64] suggests the efficiency cliff at long context is physical (KV-cache concurrency), but it is one preprint's analytical model — H100-calibrated, B200-scaled — awaiting independent measurement.

## Confidence and limitations

Confidence is moderate. The 20 full-text sources carry the quantitative claims in Themes 2 and 4, and every load-bearing number was re-verified against the retrieved full text or abstract before citation. The 45 abstract-only sources (mostly paywalled IEEE and ACM items) are cited for qualitative claims only. Specific limitations: (1) the de Vries Joule estimate — probably the most-cited per-request energy figure in public discourse — could not be retrieved through any legitimate channel and is not cited; (2) two SSRN preprints (weight-only-quantisation energy study, "Green My LLM") are excluded for the same reason, leaving a known counter-voice out of the quantisation disagreement; (3) the "AI Energy Star" probe found no citable artifact, which is evidence about the vendor-benchmark ecosystem but not proof of absence; (4) English-only sources, 2009–2026 with the two metric-antecedent exceptions noted; (5) the facility-level gap is an absence finding — no retrieved source measures it, which is strong evidence of a gap but not proof none exists.

## Evidence table

| key | design | sample | measure | finding | limitations | confidence | access | note |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ali2026assessing | benchmark | 14 open-source LLMs (6B-34B params) on 8 benchmark datasets spanning code generation, summarization, mathematical reasoning, and QA; NVIDIA A100 (80 GB) GPU with fine-grained GPU-level power sampling | Inference-time energy consumption vs accuracy (energy-accuracy trade-off) | Across 14 open-source LLMs (6B-34B) on an A100 80GB GPU, increasing model size often raises inference energy consumption significantly without proportional accuracy gains, mid-sized models achieve comparable or superior accuracy at substantially lower energy, and summarization is the most energy-intensive task domain while QA is the least (no absolute energy numbers reported in the abstract). | Single GPU platform (A100 80GB) and controlled lab conditions; trade-offs may not generalize to other hardware, serving stacks, or batch configurations; abstract-only access limits methodological detail. | moderate | abstract-only | Empirical energy-accuracy trade-off mapping for open LLM inference at GPU level - a method template for energy-aware model selection. |
| aneli2025modelling | simulation | One existing data center (experimental survey) with a numerical energy model built in TRNSYS; demand-response scenario exploiting indoor air temperature/humidity fluctuations | Data-center electricity consumption, specifically cooling-load reduction under demand response | The proposed demand-response flexibility scenario reduces the data center's electricity needs for cooling by about 30% in the calibrated and validated TRNSYS model of an existing data center. | Single existing data center; results are model-based (TRNSYS) rather than field-measured DR events; the 30% figure applies to cooling load, not whole-facility electricity. | moderate | abstract-only | Facility-level angle: data-center cooling demand response modeled in energy terms, linking facility efficiency to grid flexibility rather than tokens/watt. |
| aquino2025energy | benchmark | AlexNet, ResNet18, VGG16, EfficientNet-B3, ConvNeXt-T, and Swin Transformer trained on Imagenette; TITAN XP and GTX 1080 GPUs; OpenZmeter v2, CarbonTracker v1.2.5, CodeCarbon v2.4.1 | Newly developed energy consumption index for DL models over training and inference; energy efficiency across architectures and GPUs | Using sensor-based (OpenZmeter) and software-based (CarbonTracker, CodeCarbon) tools on TITAN XP and GTX 1080 GPUs, the study reports significant differences in energy efficiency across architectures and GPUs, but no specific energy numbers are given in the abstract. | Older GPU generation (TITAN XP/GTX 1080) and small Imagenette dataset; applicability of the index to LLM-scale workloads unshown; no absolute energy values in abstract. | moderate | abstract-only | Proposes a standardized, sensor-validated energy-efficiency index covering both training and inference across architectures - relevant as a measurement-methodology template. |
| argerich2024measuring | benchmark | Several state-of-the-art LLMs deployed for inference; varied model size, layer count, parallelized attention, and vocabulary size; input batch size and quantization levels as optimization levers | Energy consumption of LLM inference, inference energy efficiency, and latency | Profiling several state-of-the-art LLMs during inference shows that model size, layer count, parallelized attention, and vocabulary size drive inference energy consumption, while input batch size and quantization levels can be tuned to improve inference energy efficiency and latency (no numeric values reported in the abstract). | Abstract gives no numeric results or hardware details; measurement scope (which components are counted) and profiling methodology are not described in the abstract. | moderate | abstract-only | Inference-focused LLM energy profiling methodology with quantization and batching as efficiency levers - complements training-centric carbon studies. |
| banbury2021mlperftiny | benchmark | Four TinyML benchmarks: keyword spotting (Speech Commands v2, DS-CNN with 38.6K params), visual wake words (MSCOCO-2014, MobileNetV1 325KB), image classification (CIFAR-10, ResNet 96KB), anomaly detection (ToyADMOS/DCASE2020, FC-AutoEncoder 270KB); reference implementations on NUCLEO-L4R5ZI with TFLM; EMON energy monitor with electrical-isolation proxy | Accuracy, latency, and energy of ML inference on ultra-low-power (sub-milliwatt) TinyML systems | MLPerf Tiny v0.5 measures accuracy, latency, and energy for four ultra-low-power inference benchmarks with quality targets of 90%/80%/85% top-1 accuracy and 0.85 AUC (reference models reaching e.g. 92.2% KWS accuracy), using an EMON energy monitor on systems drawing under a milliwatt, but no aggregate tokens/watt or absolute energy figures are reported. | Authors flag benchmark evolution and long-term stability as limitations; measurement-scope questions (what counts in the power measurement, data-path and pre-processing variability) and exclusion of feature extraction from KWS measurement remain open methodology issues. | high | full-text | Foundational energy-inclusive ML benchmark: makes energy a first-class metric alongside latency and accuracy for ultra-low-power inference - a key reference for energy-in-benchmark methodology. |
| beitsayadeh2026benchmarking | cohort | 739 standardized submissions from MLPerf Inference v5.0 across public and private cloud infrastructures and multiple accelerator models (incl. NVIDIA H200-SXM-141GB); efficiency estimated as throughput normalized by vendor-reported TDP | Inference throughput per watt (TDP-normalized efficiency) across clouds and accelerators | In a cross-sectional secondary analysis of 739 MLPerf Inference v5.0 submissions, throughput-per-TDP efficiency showed no significant cloud-type effect but significant accelerator-model differences - NVIDIA H200-SXM-141GB had approximately 45% higher median efficiency - with a log-linear specification fitting best, implying multiplicative efficiency scaling across accelerator generations. | Efficiency is a proxy (throughput / vendor-reported TDP), not measured power; MLPerf submissions are self-selected; no absolute power or energy-per-token values. | moderate | abstract-only | Demonstrates a reproducible performance-per-watt benchmarking framework built entirely on open MLPerf data - directly relevant to tokens-per-watt standardization. |
| billah2026ispue | survey | n/a - integrative review of data-center metrics, centered on Power Usage Effectiveness (PUE), for assessing environmental impact of modern AI infrastructure | Fitness-for-purpose assessment of PUE and alternative data-center environmental metrics | No finding extractable - the available abstract is only a teaser for an integrative review questioning whether Power Usage Effectiveness (PUE) is still fit for purpose as a metric for modern AI infrastructure; no numbers are reported. | Fetched abstract is a truncated teaser with no substantive content; method and conclusions cannot be assessed. | low | abstract-only | Facility-level metric critique (PUE) relevant to the facility-level efficiency measurement debate for AI infrastructure. |
| chen2026oneoverw | theoretical | Llama-3.1-70B, Llama-3.1-405B, Qwen3-235B-A22B, DeepSeek-V3 on H100-SXM5 and B200-SXM (TP=8, fp16); Azure LLM Inference Trace and LMSYS-Chat-1M workloads; inference-fleet-sim queueing framework, logistic GPU power model, AIConfigurator roofline | Tokens per watt (tok/W) at single-GPU and fleet level | The derived 1/W law states tokens per watt halves each time the serving context window doubles - e.g., Llama-3.1-70B on H100 drops from 35.0 tok/W at 2K context to 1.5 tok/W at 64K (roughly 40x spread across 2K-128K) - while two-pool context-length routing (FleetOpt) delivers about 2.5x fleet tok/W (14.08 vs 5.58 tok/W on the Azure trace), H100-to-B200 upgrade about 1.7x, combined 4.25x, and Qwen3-235B-A22B reaches about 37.8 tok/W at 8K on H100 (5.1x Llama-3.1-70B, an upper bound). | Authors flag: no empirical validation on B200/H200 (analytical projections with +/-20% uncertainty), MoE dispatch overhead excluded (upper bounds), steady-state traffic assumption, two-pool topology only, static fleet sizing, single workload CDF, and output-only energy accounting. | high | full-text | Core theoretical contribution for the review: context window as the dominant tok/W lever (1/W law) plus quantitative decomposition of routing-topology vs hardware-generation energy gains. |
| chien2023reducing | simulation | Workload model of ChatGPT-style generative AI inference; request-direction policies: Local, Balance, CarbonMin | Power use and carbon impacts of generative AI inference under different request-direction approaches | The ChatGPT-exemplar workload model compares request-direction approaches (Local, Balance, CarbonMin) assessing their power use and carbon impacts, but the abstract reports no numeric results. | Very short abstract with no results; workload-model validity and assumptions unverifiable from abstract alone. | low | abstract-only | Early carbon-per-query request-routing comparison (Local vs Balance vs CarbonMin) for generative AI inference - carbon-aware direction angle. |
| csikai2026scaling | benchmark | Apple M4 Pro system with a consistent local serving stack and macOS power telemetry; four open-source model families spanning 3B-70B parameters; six prompt categories from 128 to 4096 input tokens | Token-normalized energy per generated token (CPU and GPU contributions), power traces sampled at fixed interval and integrated over the model-reported evaluation window | Energy per token increases monotonically with input prompt length but with weak, sublinear growth over the evaluated range, and log-log fits yield compact scaling parameters separating context-length effects from model-specific baseline efficiency (no absolute energy values reported in the abstract). | Single hardware platform (M4 Pro) and runtime path; absolute energy-per-token numbers not available in the abstract; controlled lab conditions may not reflect production serving. | moderate | abstract-only | Contributes an Apple Silicon-appropriate measurement workflow and empirical scaling laws for energy per token vs context length - relevant to on-device tokens-per-watt benchmarking. |
| desislavov2023trends | theoretical | Relevant computer-vision and NLP models, comparing first implementations against consolidated versions one to two years later, on newer higher-FLOPS hardware with energy-efficiency optimizations | Growth of inference energy consumption relative to performance gains | For consolidated (1-2 years post-breakthrough) CV and NLP models on newer, more efficient hardware, inference energy grows much more softly than previously anticipated for sustained performance increases - contradicting exponential energy-growth expectations - with the caveat that the multiplicative factor of pervasive AI adoption could still drive large totals (no specific numbers in the abstract). | Focus on consolidated implementations may understate the energy of first implementations; hardware-efficiency assumptions may not hold for LLM-scale serving; no quantitative data in abstract. | moderate | abstract-only | Challenges the exponential energy-growth narrative for inference by accounting for hardware efficiency and implementation maturity. |
| ding2024sustainablellm | survey | n/a - LLM development/deployment lifecycle spanning training and serving phases | n/a - identifies challenges and research directions for reducing the carbon footprint of LLM serving | The paper reports that the energy consumption of LLM serving has now surpassed that of training, identifies key challenges and outlines research directions for reducing the carbon footprint of LLM serving, but the abstract gives no numeric values. | Position/vision paper without new empirical measurements; the claim that serving energy surpasses training rests on cited prior work. | moderate | abstract-only | Positions sustainable LLM serving (inference-phase carbon) as a research agenda - motivational framing for tokens-per-watt benchmarking. |
| dodge2022measuring | framework | n/a - cloud computing and machine-learning workloads; argument targeting cloud providers and data scientists | Availability of software carbon-intensity information to users | The paper argues that cloud providers presenting software carbon intensity information to users is a fundamental stepping stone towards minimizing AI/ML emissions, because data scientists today lack easy and reliable access to such measurements (no numeric values in the abstract). | Argument/position paper without measurements; feasibility and accuracy of software carbon-intensity reporting are not evaluated. | moderate | abstract-only | Measurement-availability argument: carbon-intensity reporting by cloud providers as the prerequisite for emissions reduction - a measurement-infrastructure contribution. |
| du2026tokens | simulation | Representative multi-campus LLM inference data-center system; model-quantization configurations FP16/INT8/W4A8/W4A4 (model architecture metadata from model reports, NVIDIA-spec hardware); gold/silver/bronze request tiers; MILP co-optimization with on-site turbines, BESS, PV, and grid price/carbon signals | Total data-center operating cost under demand response; per-token energy and token throughput as dispatchable scheduling parameters | In multi-campus case studies, the quantization-enabled demand-response framework (model-instance switching, request routing, precision selection) reduces total data-center operating cost by 34.3% without curtailing served token volume, while the share of on-site generation rises from 7.1% to 12.6% and BESS charge-discharge throughput increases by 47.1%. | Operates at the 15-min grid-dispatch timescale rather than millisecond serving dynamics; throughput, per-token energy, and QoS-degradation parameters are calibrated from public benchmarks and reported measurements rather than measured in-house; no absolute energy-per-token values given in the text. | high | full-text | Bridges tokens and grid energy: a quantization-to-power mapping with per-token energy (J/token) and token throughput as dispatchable flexibility parameters for demand response. |
| faiz2024llmcarbon | framework | Dense and MoE LLMs (e.g., GPT-3 175B, OPT family) trained on V100/A100 GPUs; validation on GPT-3 175B inference on 16 A100 GPUs (batch 32, 128-token input); benchmarked against mlco2 | Predicted vs actual operational carbon footprint (kg CO2eq) across training, inference, experimentation, and storage phases; hardware efficiency under parallelism configurations | LLMCarbon's operational carbon footprint projections for LLM training show disparities of <=8.2% vs actual data, while mlco2's training estimates suffer disparities of more than 69%; predicted inference latency for GPT-3 was 3.1s vs 3s actual, with inference carbon prediction error not exceeding +3.3% (assuming PUE 1.1, carbon intensity 0.429 kg CO2eq/kWh). | Higher margin of error for MoE training footprints due to architectural intricacy; accuracy depends on parallelism-configuration assumptions (e.g., hardware efficiency 39%-19.7% for suboptimal settings); embodied carbon modeling remains approximate | high | full-text | Pre-training carbon projection model covering training/inference/storage/embodied phases for dense and MoE LLMs; relevant as an estimation methodology for carbon-per-token style accounting |
| ferdaus2025evaluating | benchmark | Four AI accelerators (Nvidia A100 GPU, Intel Habana Gaudi2 HPU, Graphcore Bow-Pod64 IPU, GroqRack LPU) evaluated on MLPerf BERT-Large and ResNet50 benchmarks | Energy consumption (Wh) and energy efficiency (throughput per watt) for training and inference to reach common MLPerf-specified target accuracy | For ResNet50, Intel Gaudi2 delivered the lowest energy consumption for both training and inference and the highest inference energy efficiency, while Graphcore showed the highest training energy efficiency; for BERT-Large inference, Nvidia A100 achieved the lowest energy consumption (no absolute Wh figures in abstract). | Initial study limited to 4 accelerators and 2 benchmarks; relies on vendor-provided power monitoring tools and vendor-optimized models; no absolute energy numbers in the abstract | moderate | abstract-only | Head-to-head accelerator energy-efficiency comparison using MLPerf workloads; directly relevant to energy-per-benchmark measurement methodology |
| floresmartin2025improving | framework | Real data center use case with integrated sensor monitoring of key operational variables; machine learning analysis of sensor data | Power Usage Effectiveness (PUE) and data center energy consumption | No quantitative results reported in the abstract; the step-by-step sensor-plus-ML methodology was validated on a real use case and demonstrates potential to optimize PUE and reduce energy consumption. | No numbers available in the abstract; validation limited to a single use case; methodology outcomes depend on sensor coverage and data quality | low | abstract-only | Facility-level PUE optimization via sensors and ML; provides context for facility-level efficiency beyond per-token metrics |
| garcia2019estimation | survey | Literature on energy estimation approaches in computer architecture and machine learning; survey of latest software tools for energy estimation plus two ML use cases | Review of methods and tools for estimating energy consumption of ML algorithms | No quantitative findings in the abstract; the paper reviews energy-estimation approaches and software tools and presents two use cases to guide ML practitioners in measuring energy consumption. | Survey/guidance paper; no experimental numbers in the abstract; tools reviewed are dated (2019) | low | abstract-only | Foundational review of energy-estimation approaches and tools for ML; grounds the measurement-methodology side of the review |
| geens2024energycost | simulation | Llama2-7B inference on a representative hardware architecture using a PyTorch-based generalized LLM workload template and extended ZigZag design-space exploration framework | Energy cost (J) of prefill and decode stages; energy bottleneck attribution (memory-bound compute, weight fetching, attention) | Aggressive weight-only quantization reduces Llama2-7B inference energy cost by 4.6x and shifts the bottleneck from weight fetching to the attention mechanism; memory-bound compute in the decode stage is detrimental to both latency and energy, and prefill's relative energy share grows in edge scenarios. | Simulation-based results on a single representative architecture and one model; simulation speedups come at a 'negligible' but nonzero loss of accuracy; no absolute Joules numbers in the abstract | moderate | abstract-only | Design-space exploration framework for early identification of LLM inference energy bottlenecks; supports energy-per-token modeling before hardware exists |
| guan2024wattscope | framework | Production datacenter workload; server- and rack-level aggregate power measurements already available in datacenters (no OS or application access) | Normalized mean absolute error of per-application power disaggregation from aggregate server power | WattScope disaggregates application-level power from external aggregate measurements with high accuracy, often <~10% normalized mean absolute error on a production workload. | Relies on datacenter workload power characteristics (low variability, low magnitude, high periodicity) being amenable to disaggregation; machine-learning-based disaggregation may not generalize to workloads violating these assumptions; abstract reports accuracy on a single production workload | moderate | abstract-only | Non-intrusive per-application power measurement from aggregate server power; key methodology for attributing datacenter power to specific AI workloads |
| henderson2020systematic | framework | experiment-impact-tracker framework applied to RL algorithms (Deep RL Energy Leaderboard), image classification and machine translation inference case studies; survey of 100 randomly sampled NeurIPS 2019 papers | Real-time energy consumption (kWh) and carbon emissions (kg CO2eq) per experiment; carbon intensity of energy grids | Only 1 of 100 sampled NeurIPS 2019 papers measured energy (45 measured runtime); the authors' own experiments contributed 8.021 kg CO2eq and 24.344 kWh of electricity; running jobs in carbon-efficient regions can cut emissions by up to 30x (Quebec vs Estonia, 2017 averages). | Carbon intensities are region-averaged estimates (electricitymap.org-based); embodied/manufacturing emissions excluded; adoption depends on voluntary self-reporting by researchers | high | full-text | Foundational standardized energy/carbon reporting framework (experiment-impact-tracker) and RL energy leaderboard; defines the reporting conventions that tokens-per-watt benchmarks should adopt |
| hisaharo2024optimizing | case-study | Redesigned inference cluster architecture (advanced interconnects, high-bandwidth memory, energy-efficient power management) running a modified GPT-Neo model vs baseline | Throughput, latency, and energy consumption of inference | No quantitative results in the abstract; the redesigned cluster and modified GPT-Neo model reportedly achieved substantial improvements in throughput, latency, and energy consumption over baseline. | No numbers available in the abstract; TechRxiv preprint (not peer-reviewed); details of cluster configuration and measurement methodology not visible | low | abstract-only | Cluster- and model-level redesign for energy-efficient inference; illustrates hardware/software co-optimization levers for tokens-per-watt |
| husom2025sustainable | benchmark | 28 quantized LLMs from the Ollama library (default PTQ and weight-only quantization) deployed on a Raspberry Pi 4 (4 GB RAM), benchmarked on CommonsenseQA, BIG-Bench Hard, TruthfulQA, GSM8K, HumanEval with a high-resolution hardware-based energy measurement tool | Energy efficiency (measured power/energy), inference performance (speed), and output accuracy across quantization levels and task types | No absolute energy numbers in the abstract; the study reveals trade-offs between energy efficiency, inference speed, and accuracy across quantization settings, identifying configurations that optimize LLM deployment on resource-constrained devices. | No quantitative figures in the abstract; single edge device (Raspberry Pi 4); limited to Ollama's default quantization schemes | moderate | abstract-only | Hardware-level energy profiling of quantized LLMs on edge hardware; direct example of energy-per-task measurement methodology at the edge |
| jacquet2026untangling | theoretical | n/a - not described in abstract (presumably GPU workloads in hyperscale data centers leased to diverse clients) | GPU energy consumption in hyperscale data centers (attribution presumed from title/context) | No findings available; the abstract only frames the problem of GPU energy consumption under scrutiny in hyperscale data centers where accelerators are centralized and leased to diverse clients. | Abstract is a single sentence with no methods, results, or numbers; cannot verify design, sample, or findings from available material | low | abstract-only | Presumably addresses attributing GPU energy to tenants/clients in hyperscale data centers; relevant to per-workload power attribution if full text becomes available |
| jahanshahi2020gpunest | framework | Multi-GPU cloud inference systems; case studies on multi-GPU scaling, inference scheduling, and non-GPU bottlenecks | Energy efficiency (e.g., performance per watt) of multi-GPU inference systems | Inference scheduling improves the energy efficiency of multi-GPU inference systems by as much as 40%. | Case-study based on the systems examined; no absolute energy numbers in the abstract; findings may not generalize across scheduling policies and hardware | moderate | abstract-only | GPU-NEST characterization methodology plus scheduling insight for multi-GPU inference; system-level tokens-per-watt optimization evidence |
| jahanshahi2023wattwiser | framework | Multi-GPU ML inference serving systems with per-request latency Service-Level Objectives (SLOs); load consolidation to a subset of GPUs (specific systems not stated in abstract) | Power consumption minimization subject to SLO latency bounds | No results in the abstract; the work motivates consolidating inference load onto a subset of GPUs (and potentially sharing GPUs) to minimize power consumption without violating SLO. | Abstract is truncated and contains no methods or results; no numbers available; relationship to the 2020 GPU-NEST work suggests shared lineage but is unverifiable from this text | low | abstract-only | GPU load consolidation for power reduction in SLO-bound inference serving; relevant to power-per-request optimization under latency constraints |
| jay2023software | benchmark | Several software-based power meters for CPU- and GPU-based infrastructures evaluated against high-precision physical power meters under various intensive workloads | Accuracy of software-based power measurement (power models and vendor internal interfaces) vs physical meters at node, application, and service levels | No quantitative results in the abstract; the empirical comparison highlights the strengths and limitations of each software-based power meter and shows that choosing the right tool for a given need is difficult. | No numbers in the abstract; scope limited to the meters and CPU/GPU platforms tested; software meters trade accuracy for deployability vs physical meters | moderate | abstract-only | Validation of software power meters against physical meters; critical evidence for the accuracy limits of measurement tooling used in tokens-per-watt studies |
| jegham2025howhungry | benchmark | 30 state-of-the-art LLMs in commercial datacenters; public API performance data, company-specific multipliers (PUE, WUE, CIF), statistical inference of hardware configurations; case studies on GPT-4o scaled usage and GPT-5 adaptive routing | Wh per prompt (short/long), water and carbon per query, annualized footprints, cross-efficiency DEA ranking of performance vs environmental cost | The most energy-intensive models exceed 29 Wh per long prompt, over 65x the most efficient systems; a 0.42 Wh short GPT-4o query scaled to 700M queries/day equals annual electricity of ~35,000 US homes, and GPT-5 ranges from 0.67 Wh (short, minimal reasoning) to 33.8 Wh (long, high reasoning), with the framework's GPT-4o estimate within 19% of OpenAI's reported 0.34 Wh/query. | Hardware configurations statistically inferred rather than observed; depends on company-reported PUE/WUE/CIF; excludes idle power of unutilized GPUs in partially loaded nodes; proprietary model sizes classified from API performance; per-prompt figures aggregate wide variance | high | full-text | Prompt-level Wh/query benchmark across 30 models with facility-level multipliers and DEA efficiency ranking; the closest source to a 'tokens-per-watt'-style comparative benchmark |
| jiang2024preventing | framework | LLM-powered intelligent chatbots (e.g., ChatGPT-class systems); life-cycle and interaction analysis across eight development/deployment phases (training, fine-tuning, updating, hardware manufacturing, operations, data management, recycling) | Life-cycle energy consumption and carbon emissions of LLM chatbot services (conceptual framework) | No quantitative results in the abstract; the paper identifies eight life-cycle phases with energy and carbon implications and proposes a system-level solution with three strategic mitigation pathways. | No numbers in the abstract; conceptual life-cycle framing without empirical measurement; mitigation pathways not quantitatively evaluated here | low | abstract-only | Life-cycle (training + inference + embodied) energy/carbon framing for LLM chatbots; broadens scope beyond inference-only energy benchmarks |
| kaneko2025comparing | case-study | Bitcoin and Ethereum blockchains, GPT-4, Visa payment network, and web search (secondary/estimated consumption data for cloud-based services) | Electricity consumption (TWh, MWh) at system-wide and per-use level; energy per transaction and per inference | GPT-4 training required ~9,450 MWh while daily inference exceeded 500 MWh (inference often exceeding training energy), and Bitcoin consumes ~121 TWh (~0.43% of global electricity) with per-transaction energy 720,000x that of Visa, while Ethereum's move to PoS cut energy by 99.988%. | Estimates for closed cloud services rely on secondary/external data (author-flagged Scope 3 measurement barrier); abstract-only access; blockchain vs GenAI figures are not directly comparable. | moderate | abstract-only | Provides system-level and per-use electricity numbers showing LLM inference energy (GPT-4 daily >500 MWh) can exceed training, motivating inference energy measurement. |
| lange2009specpower | benchmark | SPECpower_ssj2008: industry-standard server systems under a Java server-side workload (n/a for AI) | Power and performance characteristics of computer systems (performance-per-watt) | SPEC established SPECpower_ssj2008 as the first industry-standard benchmark for measuring power and performance of computer systems; the one-sentence abstract reports no quantitative results. | Abstract-only (single sentence); benchmark targets general servers, not AI/LLM inference, so transferability to tokens-per-watt is indirect. | moderate | abstract-only | Lineage source: the first industry-standard performance-per-watt benchmark (SPECpower_ssj2008), a conceptual ancestor of tokens-per-watt. |
| lei2020statistical | framework | 17 hyperscale data centers (HDCs) operated by Google and Facebook, modeled with thermodynamics-based PUE models, representative economizer choices, climate variables and energy-system parameters | Predicted PUE vs reported PUE; Sobol' total-order sensitivity indices of modeling parameters; minimum achievable PUE via differential evolution | Climate variables and uninterruptible power supply (UPS) efficiencies are the most important PUE model parameters, and predictions verified against reported PUE values of 17 HDCs capture regional and seasonal PUE variations and support point estimates for macro-level data center energy models (no specific PUE values in the abstract). | Macro-level point estimations with uncertainty in energy-system parameters and economizer choices; no numeric PUE results available in the abstract; facility-level scope rather than per-inference measurement. | moderate | abstract-only | Statistical PUE-prediction framework relevant to converting inference energy into facility-level efficiency metrics and to PUE target-setting for AI data centers. |
| li2023clover | framework | ML inference services using mixed-quality model pools with GPU resource partitioning (models and hardware not specified in the abstract) | Carbon emissions of inference vs accuracy and service-level agreement (SLA) compliance | Clover, a carbon-friendly ML inference runtime using mixed-quality models and GPU partitioning, substantially reduces carbon emissions while maintaining high accuracy and meeting SLA targets; the abstract reports no numeric results. | Abstract reports no quantitative results, and evaluation scope (models, GPU types, workloads) is not stated in the abstract. | moderate | abstract-only | Early carbon-SLA co-optimization for inference serving (quality mixing + GPU partitioning) rather than a measurement benchmark. |
| li2025ecoserve | framework | Two Generative AI services at a major cloud provider (traces); Gemma 27B on NVIDIA A100 vs H100; Intel SPR CPU with llama.cpp baseline; Watttime/GreenSKU carbon-intensity traces (261 gCO2/kWh mid-level) | Total carbon (operational + embodied) of LLM serving under performance targets and SLOs; decode throughput | EcoServe lowers total carbon emissions by up to 47% versus performance-, energy-, and cost-optimized design points while meeting SLOs, based on findings that offline batch inference accounts for up to 55% of serving capacity and embodied carbon can exceed 50% of lifetime emissions. | Findings derive from modeling plus traces of two services at one cloud provider; embodied-carbon and carbon-intensity estimates carry assumption error; CPU decode gains (up to 4.03x, avg 1.34x vs llama.cpp) are specific to Intel SPR. | high | full-text | Shows carbon-optimal serving differs from energy-optimal; 4R (Reduce/Reuse/Rightsize/Recycle) framework with cross-stack ILP co-design. |
| luccioni2022bloom | case-study | BLOOM 176B parameter LM; NVIDIA A100 SXM4 80GB (TDP 400 W); Jean Zay cluster; GCP us-central1 API deployment tracked over ~18 days | tCO2eq, kWh, and gCO2eq/kWh across the full training lifecycle and real-time API inference deployment | BLOOM's final training emitted ~24.7 tCO2eq (dynamic power only) or 50.5 tCO2eq full lifecycle (433,195 kWh at ~57 gCO2eq/kWh, plus 256,646 kWh idle), and API inference emitted ~19 kg CO2eq/day (340 kg total) with ~75% of deployment energy spent just keeping the model in memory. | Authors could not track real-time power (TDP-based estimates), excluded CPU power (~40x less than GPUs), and grid carbon intensity varies by time and location; deployment tracked on a single GCP instance. | high | full-text | Canonical lifecycle carbon accounting of a 176B LLM, quantifying memory-bound, idle-dominated inference energy and advocating granular energy/carbon-intensity/PUE reporting. |
| luccioni2024powerhungry | benchmark | 88 models (80 task-specific finetuned + 8 multi-purpose zero-shot: Flan-T5 base/large/xl/xxl, BLOOMz-560M/1B/3B/7B) across 10 tasks and 30 datasets (text/image classification, QA, MLM, token classification, text generation, summarization, captioning, object detection, image generation) run sequentially (no batching) 10x on 8x NVIDIA A100-SXM4-80GB GPUs (AWS us-west-2, 297.6 g CO2eq/kWh), energy/carbon measured with CodeCarbon | Energy (kWh) and carbon (g CO2eq) per 1,000 inferences; training-vs-inference cost parity (number of inferences) | Per-task mean energy per 1,000 inferences ranges from 0.002 kWh (text classification) to 2.9 kWh (image generation, median 1.35 kWh) - a >1,450x spread - with multi-purpose generative models orders of magnitude more energy-intensive than task-specific ones (e.g., 0.3-0.7 g CO2eq per 1,000 inferences for task-specific QA/sentiment vs 2.34-10 g for zero-shot models), total study consumption 754.66 kWh / 178.97 kg CO2eq, and BLOOMz training-inference cost parity reached at ~205M-593M inferences per model. | Single hardware platform and region (A100, us-west-2); sequential non-batched inference (reflects in-situ deployment but not batching gains); idle power of other GPUs included in measurements; open-source models only; authors note study is not representative of all deployment contexts and that proprietary-model transparency is lacking. | high | full-text | Foundational inference-phase energy benchmark establishing kWh (and g CO2eq) per 1,000 inferences as a comparison unit and quantifying the energy penalty of multi-purpose generative models - a key precursor to tokens-per-watt benchmarking. |
| maliakel2026characterizing | benchmark | Llama-1B/3B/8B and Qwen-14B/32B (five decoder-only LLMs) on a single NVIDIA RTX PRO 6000 (Blackwell) GPU; BoolQ, HellaSwag, TruthfulQA, NarrativeQA; NVML power sampling at 10 ms | Energy per query (J), end-to-end latency, output quality, and energy-performance tradeoffs under GPU SM DVFS (7 frequency levels, 180-2842 MHz) | Reducing SM frequency from 2842 to 180 MHz achieves ~42% average energy savings with only 1-6% latency increase (decode dominates 77-91% of time and is frequency-insensitive), and combining workload-aware model routing (32B->3B) with DVFS cuts per-query energy from 20.97 J to 2.52 J (88% saving) in the upper-bound use case. | Single-GPU offline setup (no multi-GPU or live serving); routing-based savings trade quality (83.8% -> 77.0% in the use case) and assume accurate difficulty prediction (input length is a weak predictor, r=0.002). | high | full-text | Phase-aware DVFS plus workload-aware routing evidence in joules-per-query terms, with a quantified energy-quality frontier. |
| niu2025energy | benchmark | vLLM, TensorRT-LLM, and DeepSpeed inference engines on one GPU node with 2x H100 GPUs | Power (W) and energy decomposed by inference stage (setup: initialization + model loading vs token generation) and by component (GPU, CPU, DRAM) | Provides a fine-grained power benchmark of LLM inference engines on a 2x H100 node, decomposing the inference lifecycle into setup vs token-generation stages and GPU/CPU/DRAM components to identify energy bottlenecks; the abstract reports no numeric results. | Abstract-only; single node with 2 H100 GPUs, no multi-node or production serving deployment, and engine/version details not in the abstract. | moderate | abstract-only | Stage- and component-level power breakdown of mainstream inference engines - evidence for where inference energy actually goes. |
| niu2026tokenpowerbench | benchmark | Llama, Falcon, Qwen, and Mistral model series from 1B up to Llama3-405B; declarative config over model/prompt set/inference engine; GPU-, node-, and system-level measurement without specialized power meters | Joules per token and other energy-efficiency metrics, with phase-aligned attribution of energy to prefill and decode stages per request | TokenPowerBench, the first lightweight benchmark for LLM-inference power studies, captures GPU/node/system-level power without specialized meters and attributes energy to prefill vs decode per request (joules per token) across models from 1B to Llama3-405B; the abstract reports no numeric results. | Abstract-only; measurement fidelity without specialized meters is an inherent design tradeoff, and coverage is limited to four model families and their engines. | moderate | abstract-only | Directly delivers a tokens-per-watt-style benchmark (joules per token, phase-aligned prefill/decode) motivated by inference being >90% of total LLM power per industry reports. |
| patterson2021carbon | case-study | T5, Meena, GShard, Switch Transformer, GPT-3, and Evolved Transformer (NAS); TPU v2/v3, P100, V100; Google datacenters at multiple locations | Energy use (kWh/MWh) and CO2e (tCO2e) of training and inference, plus average system power (W) per processor | Sparsely activated DNNs consume <1/10th the energy of dense DNNs at equal accuracy, cloud datacenters are ~1.4-2x and ML accelerators ~2-5x more efficient, and combined DNN/datacenter/processor choices cut carbon footprint up to ~100-1000x, with training CO2e of the studied models ranging 4-552 tCO2e. | Estimates rely on measured average system power (TPU v2 221 W, TPU v3 283 W, P100 271 W, V100 325 W) and grid carbon-intensity assumptions; retrospective estimates can be off by 18.7x (average org) to 88x (efficient org) as shown for the NAS case. | high | full-text | Foundational argument that energy use and CO2e should be first-class ML evaluation metrics; authors collaborated with MLPerf to include energy during training and inference. |
| poddar2025towards | benchmark | LLM inference across a wide range of NLP tasks; multiple models, tasks, prompts, and system-related factors | Inference energy across tasks, models, and system configurations | First broad benchmark of LLM inference energy across NLP tasks: inference energy correlates strongly with output token length and response time, and quantization, optimal batch sizes, and targeted prompt phrasing significantly reduce energy use. | Abstract-only; exact magnitudes not reported in the abstract. | moderate | abstract-only | Academic benchmark of LLM inference energy across tasks, models, and prompts. |
| reddi2020mlperfinference | benchmark | MLPerf Inference (v0.5): ResNet-50 v1.5, MobileNet-v1, GNMT (NMT) among other workloads; 30+ systems spanning embedded to datacenter; 600+ measurements from 14 organizations | Inference latency, throughput (latency-bounded), and accuracy across four scenarios (server, single-stream, multistream, offline); system power/performance range | ML inference systems span three orders of magnitude in power consumption and five in performance, and latency constraints cut server-scenario throughput by 39-55% for NMT versus 3-35% (avg ~20%) for ResNet-50 v1.5 and under 10% for MobileNet-v1. | The first-round benchmark did not include energy or power as a submission metric (latency/accuracy only), and results are self-reported by submitters subject to audit. | high | full-text | Foundational MLPerf Inference methodology (scenarios, LoadGen, quality targets, 99% tail-latency confidence bounds) that later gained a Power/energy division. |
| reddi2020mlpermobile | benchmark | Mobile devices with diverse SoCs (e.g., MediaTek Dimensity 1100, Snapdragon) and software stacks (NNAPI, TFLite, vendor SDKs, OpenVINO); CV and NLP tasks | Latency, throughput, and accuracy of on-device ML inference (no energy metric in this version) | Across the first two benchmark rounds within six months, offline throughput improved 3x and latency reduced by up to 12x on mobile devices, while an optimized framework delegate (Neuron vs generic NNAPI) delivered over 10% performance difference. | Benchmark measures latency/throughput/accuracy only - no energy or power metric in this version; results reflect early rounds (v0.7/v1.0 era) and specific devices. | high | full-text | Industry-standard mobile ML benchmark methodology (LoadGen-based) showing rapid stack-level improvements; its lack of an energy metric motivates power-aware mobile benchmarking. |
| rrapaj2024power | case-study | Cori and Perlmutter supercomputers at NERSC; six months of production HPC workload power measurements across the CPU-to-GPU transition | Power draw (W) vs peak provisioned power and TDP over time and across applications/users | Power usage varied considerably but stayed consistently well below peak provisioned power on both machines, and after the GPU transition production power demands did not grow as fast as peak capabilities (further lowering the fraction of TDP used), suggesting machines could be power-capped well below TDP; no numeric values in the abstract. | Abstract-only; findings cover HPC workloads at one site (NERSC), not AI/LLM inference, and no numeric values are available in the abstract. | moderate | abstract-only | Facility-level evidence that production power sits far below TDP/peak provisioning - supports power-capped, over-provisioned designs relevant to inference datacenters. |
| rubei2025prompt | benchmark | Llama 3 on the CodeXGLUE code-generation benchmark, evaluated in an isolated testing environment | Energy consumption and accuracy of generated code during LLM inference (carbon/energy impact of prompt engineering techniques) | Initial results show that using specific tags to distinguish prompt parts can reduce Llama 3's energy consumption during inference without compromising code-generation performance, though no quantitative energy figures are reported in the abstract. | Authors flag the results as initial and requiring more in-depth evaluation; single model and task; no numbers disclosed in the abstract. | moderate | abstract-only | Shows prompt engineering itself as a lever on inference energy (relevant to energy-per-token optimization), but lacks quantitative measurements. |
| samsi2023words | benchmark | LLaMA 7B/13B/65B on NVIDIA V100 and A100 GPUs with model sharding across up to 32 GPUs, batch sizes 64-512, on Alpaca and GSM8K datasets (4,096 sampled inputs per dataset) | Inference energy costs: Joules per output token, energy per response, energy per second (Watts), token rate, and effect of GPU power capping | LLaMA 65B inference consumes roughly 3-4 Joules per output token at ~300 W to ~1 kW across 8-32 shards, and power capping A100s from 250 W to 175 W reduces total energy by ~23.21% at only +6.7% average inference time (a 150 W cap costs +19.49% time). | Energy figures are estimates from power draw assumptions on only two GPU generations and one model family (LLaMA); limited datasets; authors note broader power-capping recommendations need additional experimentation. | high | full-text | Landmark 'from words to watts' study establishing per-output-token Joule figures and power-capping trade-offs for LLM inference benchmarking. |
| sanchez2025greenmlops | benchmark | Discriminative models (various architectures and hyperparameters) and generative LLMs of different sizes across multiple hardware setups in real-world MLOps pipelines | Energy consumption during training and inference via software-based power measurements; correlations with model size, reasoning complexity, and request-handling capacity | For discriminative models, optimizing architecture, hyperparameters, and hardware significantly reduces energy without sacrificing performance, and for LLMs larger models do not necessarily consume more energy when utilization is low; no numeric figures are given in the abstract. | Software-based power measurement rather than watt-meters; abstract provides no quantitative results or model/hardware enumeration. | moderate | abstract-only | Positions itself as a benchmark for estimating total AI energy use across model types in MLOps, emphasizing utilization-dependent LLM efficiency. |
| schwartz2020greenai | theoretical | n/a (position paper; cites aggregate trend: deep learning computation doubling every few months, ~300,000x increase from 2012 to 2018) | Proposes efficiency as an evaluation criterion alongside accuracy and reporting of financial cost ('price tag') for developing, training and running models | Argues deep learning's computations have an estimated 300,000x increase from 2012 to 2018 with a surprisingly large carbon footprint, and advocates making efficiency a standard evaluation criterion and reporting cost baselines to enable greener, more inclusive AI research (no new empirical measurements). | Position paper without new measurements; efficiency metrics proposed qualitatively rather than operationalized as a benchmark; no numeric per-model efficiency data. | high | abstract-only | Provides the normative argument ('Green AI') for treating energy efficiency as a first-class evaluation criterion - the conceptual basis for tokens-per-watt benchmarks. |
| sejourne2026sovereign | benchmark | Heterogeneous sovereign (SecNumCloud-compliant) infrastructure: NVIDIA GPUs (L40S, A100), legacy hardware, and Apple Silicon architectures; interactive inference workloads; carbon model integrating operational energy and embodied hardware carbon; custom simulator for hardware-renewal vs legacy-extension arbitrage | Energy per request across hardware; total carbon cost (operational + embodied); scale-to-zero capability | Apple Silicon architectures consume 4x-7x less energy per request for interactive workloads than traditional NVIDIA-based servers while high-end GPUs offer superior efficiency at saturation, and a custom simulator optimizes total carbon cost by arbitrating between hardware renewal and extending legacy equipment lifespans (measured values in Fig. 2 and Table II, not in abstract). | Constrained industrial setting (SecNumCloud compliance) and specific interactive workloads; quantitative energy-per-request values referenced to figures/tables not available in the abstract. | moderate | abstract-only | Extends per-request energy benchmarking to include embodied carbon and hardware-heterogeneity ('digital sobriety') decisions in sovereign AI infrastructure. |
| stojkovic2024towards | benchmark | Llama-2 70B served with vLLM on an NVIDIA DGX-H100 (H100 GPUs, frequencies 800-1980 MHz, tensor parallelism degrees 2/4/8, batch sizes up to 64) under latency SLOs of 5x solo TTFT/TBT | GPU power draw, total energy, latency (TTFT/TBT), and throughput under frequency capping, batching, and model-parallelism knobs | GPU frequency capping achieves ~20% lower power for most workload configurations with no latency or throughput impact, running at 1.6 GHz instead of 2.0 GHz yields about the same throughput at under 80% of the energy, and reducing maximum batch size during low-throughput phases cuts energy by up to 15%. | Single model (Llama-2 70B) and single-node DGX-H100; GPUs offer only GPU-wide frequency control (no fine-grained power knobs); results are workload- and SLO-dependent. | high | full-text | Characterizes frequency, batching, and parallelism levers for energy-efficient LLM serving under performance SLOs - core input for tokens-per-watt operating-point selection. |
| stojkovic2025dynamollm | framework | LLM inference clusters of 8xH100 DGX servers (12 servers baseline; 11 in 24h run) running Llama2-70B (primary) plus Llama2-13B, Llama3-70B, Mixtral-8x7B, Mixtral-8x22B, Falcon-180B on vLLM; workloads from Azure production traces (Coding, Conversation; 1h open-source, 1-day and 1-week traces); requests bucketed into 9 input/output length classes (SS..LL) with TTFT/TBT SLOs at 5x isolated latency | Energy in Watt-hours (Wh) per request/configuration and per cluster under latency SLOs; operational CO2 emissions; customer cost; GPU power (W) | At service level DynamoLLM conserves 53% energy, 38% operational carbon emissions (5.0 vs 3.1 t CO2/week on CAISO), and 61% customer cost (40 -> 24.6 GPU servers, $1,362.7/h saved) while meeting latency SLOs; measured per-request energy for Llama2-70B ranges ~0.77-13.21 Wh depending on request length, tensor parallelism and GPU frequency, and cluster energy is reduced 35% (1h trace), 42% (1-day trace), 23-51% across load levels, and 47-56% in week-long simulations. | Evaluated on open-source models and Azure traces only; week-long results come from a discrete-time simulator; considers tensor parallelism only (not pipeline parallelism); reports Wh per request and cluster energy rather than a per-token (J/token) metric; SLO assumptions (5x isolated latency) may not generalize. | high | full-text | First energy-management framework for LLM inference clusters, quantifying how GPU frequency scaling, tensor parallelism and instance scaling trade off against energy-per-request under SLOs - direct evidence for cluster-level tokens-per-watt optimization (energy savings of 23-56%). |
| strubell2019energy | benchmark | A variety of recently successful neural network models for NLP (the paper's measurements covered transformer-based architectures of the era, e.g., Transformer base/large, ELMo, BERT base/large, GPT-2, and NAS) trained on GPU hardware | Approximate financial cost (hardware/cloud), energy (kWh), and carbon emissions (CO2) of model training | Quantifies the approximate financial and environmental costs of training recently successful NLP models, showing accuracy gains depend on substantial energy consumption, and proposes actionable recommendations to reduce costs (no numeric values in the abstract; the paper's well-known figures are training-cost estimates in kWh/lbs CO2). | Estimates tied to specific hardware, electricity prices and grid carbon intensity of the period; training-phase focus only (inference not measured); abstract provides no numbers. | moderate | abstract-only | Seminal training-phase energy/carbon quantification for NLP models that motivated subsequent inference-phase tokens-per-watt benchmarking. |
| tripp2024measuring | benchmark | BUTTER-E dataset: 63,527 runs / 30,582 configurations of fully connected networks (13 datasets, 20 sizes, 8 shapes, 14 depths) on NREL Eagle HPC CPU nodes (Xeon Gold 6154) and GPU nodes (2x V100), measured with node-level watt-meters (HPE iLO) | Real-world energy consumption (Joules per training datum/epoch, power time series) and accuracy of a proposed hardware-informed energy model | GPU training consumed a median ~9.47 mJ per training datum vs ~6.16 mJ on CPU (GPUs less energy-efficient in this setting), energy per datum rises non-linearly with parameter count until ~2^20 params, and the proposed energy model predicts GPU energy within +/-2.74% and CPU energy within +/-20.2%. | Fully connected networks only (authors explicitly call for extending to LLMs, CNNs, GNNs); training rather than inference focus; single HPC datacenter; 163 runs filtered as system artifacts. | high | full-text | Provides a rigorous watt-meter-based measurement methodology and public dataset (BUTTER-E) relevant to energy benchmarking practice, and challenges FLOP/parameter-count assumptions about efficiency. |
| tschand2025mlperf | benchmark | 60 systems spanning edge devices to cloud datacenters running representative workloads from the MLPerf benchmark suite; 1,841 reproducible measurements | Energy efficiency of ML systems across power levels from microwatts to megawatts, via the MLPerf Power methodology (rules and best practices) | 1,841 reproducible measurements from 60 systems reveal trade-offs between performance, complexity, and energy efficiency across ML deployment scales; no per-measurement energy figures are given in the abstract. | Abstract-only; as a methodology/standards paper, comparability across heterogeneous platforms is itself a stated challenge the rules must address. | moderate | abstract-only | MLPerf Power defines the industry-standard benchmarking methodology for ML energy efficiency from edge to datacenter - the direct benchmark lineage for tokens-per-watt. |
| verma2020demystifying | benchmark | MLPerf benchmark suite workloads compared against DAWNBench and DeepBench on multi-GPU training systems | Compute rate, memory transactions per second, scaling efficiency, host CPU utilization, and mixed-precision training effects | MLPerf benchmarks exhibit moderately high memory transactions per second and compute rates (vs DAWNBench's high-compute/low-memory and DeepBench's low-compute profiles), with scaling-efficiency variation across models and quantified gains from mixed-precision Tensor Core training; no energy numbers are reported in the abstract. | Pre-LLM-era training workloads; characterization targets performance and bottlenecks, not energy efficiency. | moderate | abstract-only | Background on what MLPerf workloads measure (compute/memory behavior), showing why energy efficiency needs a separate measurement dimension (later supplied by MLPerf Power). |
| wang2025storellm | framework | LLM inference with permanently pre-stored attention matrices, evaluated against state-of-the-art LazyLLM, plus StoreLLM-MoE and StoreLLM-PTQ variants | Energy consumption of LLM inference (primary outcome) and inference delay | StoreLLM outperforms LazyLLM by 1.45x in energy consumption with a sacrifice of only 5.05% in delays, and the StoreLLM-MoE and StoreLLM-PTQ variants achieve 2.64x and 2.83x energy reductions respectively versus state-of-the-art LLM systems. | Abstract-only; relies on the observation that attention matrices stay largely unchanged across inferences; storage costs of pre-stored attention matrices are not quantified in the abstract. | moderate | abstract-only | Radical energy-vs-compute substitution idea (replace attention-matrix computation with storage access) claiming large inference energy reductions. |
| wilhelm2025beyond | benchmark | Llama 3.2 1B and 8B (plus 7B-class LLMs for MT-Bench single-token tests) on a single NVIDIA L40S GPU, batch size 1, no parallelism; MMLU (57 categories clustered into 8 domains) and MT-Bench; NVML-based power measurement | Energy-per-Token (Joule) = W_consumed x time / tokens processed, alongside accuracy on MMLU and MT-Bench | Routing to Llama 8B instead of CoT on Llama 1B offers similar or better accuracy at 40-60% lower energy, while CoT boosts Math accuracy by 281% at a ~15,132% energy increase (category baselines of 76-84 MJ), motivating their proposed Energy-per-Token metric and operating-curve-based energy-aware routing. | Single GPU (L40S), batch size 1, no parallelism, small models (1B/8B); GPU-level energy only, excluding system overheads. | high | full-text | Directly advocates Energy-per-Token as a standard efficiency metric complementing accuracy benchmarks - central to the tokens-per-watt review framing. |
| wilkins2024hybrid | simulation | Representative LLM dataset of queries with heterogeneous hardware accelerators (energy-efficient processors vs high-performance GPUs) in a hybrid datacenter model | CPU+GPU energy consumption of LLM workloads under a cost-based, workload-aware scheduling framework | The hybrid strategy, which allocates tasks to energy-efficient processors or high-performance GPUs based on input/output token counts per query, reduces CPU+GPU energy consumption by 7.5% compared to a workload-unaware baseline; no absolute energy figures are given in the abstract. | Abstract-only; representative (unspecified) LLM dataset; datacenter-model-level analysis rather than hardware measurements. | moderate | abstract-only | Token-count-aware scheduling across heterogeneous hardware as a datacenter-level energy-efficiency lever for LLM serving. |
| wilkins2024offline | simulation | Several state-of-the-art LLMs on heterogeneous GPU-CPU systems, characterized across different magnitudes of input prompts and output text | Workload-dependent energy consumption and runtime of LLM inference; fit quality (R^2) of energy/runtime models and energy-optimality of offline scheduling | Energy and runtime models fit each LLM with R^2 > 0.96 across prompt/output magnitudes, and a case study of the offline energy-optimal scheduling framework demonstrates advantages of energy- and accuracy-aware scheduling over existing best practices; no absolute numbers are given in the abstract. | Abstract-only; specific models and systems not enumerated; offline scheduling only, no online/adaptive evaluation. | moderate | abstract-only | Shows that workload-dependent energy models of LLM inference can be accurate enough (R^2 > 0.96) to drive energy-optimal scheduling - methodology relevant to energy-per-token modeling. |
| wu2022sustainable | case-study | Meta/Facebook production ML workloads (large language model LM, ranking models RM1-RM5), large-scale models including GPT-3 (750B params) and Switch Transformer (1.5T params), datacenter fleet with PUE ~1.10, and LCA-based hardware embodied-carbon analysis | End-to-end carbon footprint (operational + embodied/manufacturing) and operational power consumption of AI computing across the model development cycle and system hardware life cycle | Manufacturing (embodied) carbon cost is roughly 50% of the location-based operational carbon footprint of large-scale ML tasks at Meta, aggregate cross-stack optimizations reduced operational power by on average 20% every six months, the LM model's carbon footprint is dominated by inference while training-vs-inference footprints are roughly equal for ranking models, and raising GPU utilization to 80% decreases overall carbon footprint by 3x. | Single-company (Meta) experience; location-based carbon intensities assumed and PUE assumed at 1.1 rather than measured per workload; results may not generalize to other providers or workloads | high | full-text | Industry-scale characterization of training vs inference energy split and a call to add efficiency/environmental measures to MLPerf-style leaderboards, directly motivating tokens-per-watt-style benchmarking. |
| xu2021surveygreen | survey | n/a - survey of green deep learning literature: compact networks (e.g., MobileNet family, efficient attention/softmax variants), energy-efficient training (initialization, normalization, progressive training, HPO), energy-efficient inference (pruning, low-rank factorization, quantization incl. 8-bit BERT, I-BERT, TernaryBERT, distillation), and efficient data usage | n/a - organizes methods into 4 categories; discusses candidate efficiency measures: running time, carbon emission (CO2eq), model size, FLOPs, 'fair measure', and intuitive understanding | No new empirical numbers are reported; the survey classifies green deep learning techniques into compact networks, energy-efficient training, energy-efficient inference, and efficient data usage, and recommends reporting FLOPs plus running time for fair comparison while noting FLOPs are theoretical values that diverge from actual runtime. | Descriptive survey without measurements; authors flag that FLOPs are theoretical and parallelism/degree of utilization create a gap between FLOPs and running time; efficiency claims of surveyed methods not independently verified | high | full-text | Provides the measurement-metrics discussion (FLOPs vs runtime vs carbon emission vs fair measure) that informs how a tokens-per-watt benchmark should be defined and reported. |
| yang2026pelm | benchmark | Mobile/edge hardware platforms and datasets (not enumerated in abstract) running on-device LLM inference; compared against state-of-the-art DVFS-based power governing methods | Energy consumption and inference speedup under thermal/power constraints, with task performance maintained | PELM, which augments DVFS frequency tuning with speculative decoding and variable verification depth, achieves up to 23.1% speedup and 52.4% reduction in energy consumption compared to state-of-the-art power governing methods while maintaining comparable task performance (numbers from abstract). | Abstract-only; hardware platforms, datasets, and LLMs not named in the abstract; results specific to thermally constrained mobile scenarios and dependent on speculative decoding quality | moderate | abstract-only | Adds workload-level power knobs (speculative decoding, verification depth) beyond DVFS for energy-efficient on-device LLM inference - relevant to edge tokens-per-watt optimization. |
| yilk2018qualifying | case-study | Four LANL supercomputing platforms: Trinity (separate Haswell and Knights Landing CPU partitions), Grizzly, Fire, Ice; enhanced power monitoring infrastructure; Green500 benchmark | Green500 qualification level (performance per watt) and experience meeting Green500 reporting requirements | All four new LANL platforms were qualified at the highest level of the Green500 benchmark using the enhanced power-monitoring infrastructure (no numeric efficiency values in the abstract). | Experience report focused on the qualification process; no quantitative performance-per-watt or power numbers available in the abstract. | moderate | abstract-only | Documents institutional practice of qualifying HPC systems on the Green500 performance-per-watt benchmark - a methodological precedent for tokens-per-watt-style efficiency qualification. |
| yu2023know | simulation | Cloud-scale ML inference cluster simulation with prototype implementation: 105 servers with three different kinds of GPUs serving five ML models, evaluated with real-world traces | Cloud-scale energy consumption of ML inference serving; energy efficiency of GPU architectures across active-GPU counts and clock frequencies | A hierarchical GPU resource management approach (energy-aware cluster allocation, intra-cluster node scaling, intra-node GPU scaling, and GPU clock scaling) saves up to 28.3% of cloud-scale energy consumption when serving five ML models on 105 servers with three GPU types (from abstract). | Abstract-only; evaluation combines prototype with trace-driven cloud-scale simulation rather than full production deployment; SLO-blind DVFS finding is specific to commercial GPU drivers; generalizability to other workloads/fleets not shown | moderate | abstract-only | Shows GPU energy efficiency varies with architecture, active-GPU count, and clock frequency at constant throughput - cluster-level evidence for facility/cloud-scale tokens-per-watt management. |
| zhao2023sustainable | case-study | GPUs at a research supercomputing center (HPC/datacenter production environment) subjected to power-capping for AI/ML workloads | GPU temperature and power draw under power-capping, plus impact on job performance and hardware lifespan | Power-capping GPUs at a research supercomputing center significantly decreased both GPU temperature and power draw, reducing power consumption and potentially improving hardware lifespan with minimal impact on job performance; the abstract reports no specific magnitudes. | Abstract-only; single-center field study; magnitude of temperature/power reductions and the precise job-performance impact are not quantified in the abstract | moderate | abstract-only | Facility-scale evidence that power-capping AI accelerators cuts power draw and temperature with minimal performance cost - relevant to datacenter-level energy accounting for inference fleets. |

## References

1. Tschand, Arya et al. (2025). *MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from μWatts to MWatts for Sustainable AI*. 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). Abstract only. MLPerf Power defines the industry-standard benchmarking methodology for ML energy efficiency from edge to datacenter - the direct benchmark lineage for tokens-per-watt. [doi:10.1109/hpca61900.2025.00092](https://doi.org/10.1109/hpca61900.2025.00092)
2. Verma, Snehil et al. (2020). *Demystifying the MLPerf Training Benchmark Suite*. IEEE ISPASS 2020. Abstract only. Background on what MLPerf workloads measure (compute/memory behavior), showing why energy efficiency needs a separate measurement dimension (later supplied by MLPerf Power). [doi:10.1109/ispass48437.2020.00013](https://doi.org/10.1109/ispass48437.2020.00013)
3. Reddi, Vijay Janapa et al. (2019). *MLPerf Inference Benchmark*. arXiv preprint. Full text read. Foundational MLPerf Inference methodology (scenarios, LoadGen, quality targets, 99% tail-latency confidence bounds) that later gained a Power/energy division. [doi:10.48550/arxiv.1911.02549](https://doi.org/10.48550/arxiv.1911.02549)
4. Banbury, Colby et al. (2021). *MLPerf Tiny Benchmark*. arXiv (Cornell University). Full text read. Foundational energy-inclusive ML benchmark: makes energy a first-class metric alongside latency and accuracy for ultra-low-power inference - a key reference for energy-in-benchmark methodology. [doi:10.48550/arxiv.2106.07597](https://doi.org/10.48550/arxiv.2106.07597)
5. Reddi, Vijay Janapa et al. (2020). *MLPerf Mobile Inference Benchmark*. arXiv (Cornell University). Full text read. Industry-standard mobile ML benchmark methodology (LoadGen-based) showing rapid stack-level improvements; its lack of an energy metric motivates power-aware mobile benchmarking. [doi:10.48550/arxiv.2012.02328](https://doi.org/10.48550/arxiv.2012.02328)
6. Lange, Klaus-Dieter (2009). *Identifying Shades of Green: The SPECpower Benchmarks*. Computer. Abstract only. Lineage source: the first industry-standard performance-per-watt benchmark (SPECpower_ssj2008), a conceptual ancestor of tokens-per-watt. [doi:10.1109/mc.2009.84](https://doi.org/10.1109/mc.2009.84)
7. Jahanshahi, Ali et al. (2020). *GPU-NEST: Characterizing Energy Efficiency of Multi-GPU Inference Servers*. IEEE Computer Architecture Letters. Abstract only. GPU-NEST characterization methodology plus scheduling insight for multi-GPU inference; system-level tokens-per-watt optimization evidence [doi:10.1109/lca.2020.3023723](https://doi.org/10.1109/lca.2020.3023723)
8. Niu, Chenxu et al. (2026). *TokenPowerBench: Benchmarking the Power Consumption of LLM Inference*. Proceedings of the AAAI Conference on Artificial Intelligence. Abstract only. Directly delivers a tokens-per-watt-style benchmark (joules per token, phase-aligned prefill/decode) motivated by inference being >90% of total LLM power per industry reports. [doi:10.1609/aaai.v40i38.40535](https://doi.org/10.1609/aaai.v40i38.40535)
9. Niu, Chenxu (2025). *Energy Efficient or Exhaustive? Benchmarking Power Consumption of LLM Inference Engines*. ACM SIGEnergy Energy Informatics Review. Abstract only. Stage- and component-level power breakdown of mainstream inference engines - evidence for where inference energy actually goes. [doi:10.1145/3757892.3757900](https://doi.org/10.1145/3757892.3757900)
10. Beitsayadeh, Carl & Darbyshire, Pamayla E. (2026). *Benchmarking AI Inference Efficiency in Public and Private Clouds: An MLPerf-Based Comparative Study*. IEEE Transactions on Cloud Computing. Abstract only. Demonstrates a reproducible performance-per-watt benchmarking framework built entirely on open MLPerf data - directly relevant to tokens-per-watt standardization. [doi:10.1109/tcc.2026.3674888](https://doi.org/10.1109/tcc.2026.3674888)
11. Ferdaus, Farah et al. (2025). *Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf Benchmarks*. 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Abstract only. Head-to-head accelerator energy-efficiency comparison using MLPerf workloads; directly relevant to energy-per-benchmark measurement methodology [doi:10.1109/ccgrid64434.2025.00035](https://doi.org/10.1109/ccgrid64434.2025.00035)
12. Maliakel, Paul Joe et al. (2026). *Characterizing LLM Inference Energy-Performance Tradeoffs Across Workloads and GPU Scaling*. 2026 IEEE 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Full text read. Phase-aware DVFS plus workload-aware routing evidence in joules-per-query terms, with a quantified energy-quality frontier. [doi:10.1109/ccgrid68966.2026.00013](https://doi.org/10.1109/ccgrid68966.2026.00013)
13. Csikai, Dávid et al. (2026). *Scaling Behavior of Energy per Token in LLM Inference on Apple Silicon*. 2026 IEEE 8th International Conference and Workshop Óbuda on Electrical and Power Engineering (CANDO-EPE). Abstract only. Contributes an Apple Silicon-appropriate measurement workflow and empirical scaling laws for energy per token vs context length - relevant to on-device tokens-per-watt benchmarking. [doi:10.1109/cando-epe71091.2026.11569440](https://doi.org/10.1109/cando-epe71091.2026.11569440)
14. Jay, Mathilde (2023). *An experimental comparison of software-based power meters: focus on CPU and GPU*. 2023 IEEE/ACM 23rd International Symposium on. Abstract only. Validation of software power meters against physical meters; critical evidence for the accuracy limits of measurement tooling used in tokens-per-watt studies [doi:10.1109/ccgrid57682.2023.00020](https://doi.org/10.1109/ccgrid57682.2023.00020)
15. Jacquet, Pierre et al. (2026). *Untangling GPU Power Consumption: Job-Level Inference in Cloud Shared Settings*. Proceedings of the 21st European Conference on Computer Systems. Abstract only. Presumably addresses attributing GPU energy to tenants/clients in hyperscale data centers; relevant to per-workload power attribution if full text becomes available [doi:10.1145/3767295.3769333](https://doi.org/10.1145/3767295.3769333)
16. Guan, Xiaoding et al. (2024). *WattScope: Non-intrusive Application-level Power Disaggregation in Datacenters*. ACM SIGMETRICS Performance Evaluation Review. Abstract only. Non-intrusive per-application power measurement from aggregate server power; key methodology for attributing datacenter power to specific AI workloads [doi:10.1145/3649477.3649491](https://doi.org/10.1145/3649477.3649491)
17. Tripp, Charles et al. (2024). *Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations*. arXiv (Cornell University). Full text read. Provides a rigorous watt-meter-based measurement methodology and public dataset (BUTTER-E) relevant to energy benchmarking practice, and challenges FLOP/parameter-count assumptions about efficiency. [doi:10.48550/arxiv.2403.08151](https://doi.org/10.48550/arxiv.2403.08151)
18. Samsi, Siddharth (2023). *From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference*. arXiv (Cornell University). Full text read. Landmark 'from words to watts' study establishing per-output-token Joule figures and power-capping trade-offs for LLM inference benchmarking. [doi:10.48550/arxiv.2310.03003](https://doi.org/10.48550/arxiv.2310.03003)
19. Argerich, Mauricio Fadel & Patiño-Martı́nez, Marta (2024). *Measuring and Improving the Energy Efficiency of Large Language Models Inference*. IEEE Access. Abstract only. Inference-focused LLM energy profiling methodology with quantization and batching as efficiency levers - complements training-centric carbon studies. [doi:10.1109/access.2024.3409745](https://doi.org/10.1109/access.2024.3409745)
20. Hisaharo, Soka (2024). *Optimizing LLM Inference Clusters for Enhanced Performance and Energy Efficiency*. arXiv preprint. Abstract only. Cluster- and model-level redesign for energy-efficient inference; illustrates hardware/software co-optimization levers for tokens-per-watt [doi:10.36227/techrxiv.172348951.12175366/v1](https://doi.org/10.36227/techrxiv.172348951.12175366/v1)
21. Yu, Junyeol (2023). *Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving*. 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Abstract only. Shows GPU energy efficiency varies with architecture, active-GPU count, and clock frequency at constant throughput - cluster-level evidence for facility/cloud-scale tokens-per-watt management. [doi:10.1109/hpca56546.2023.10070943](https://doi.org/10.1109/hpca56546.2023.10070943)
22. Wilkins, Grant (2024). *Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems*. ACM SIGEnergy Energy Informatics Review. Abstract only. Shows that workload-dependent energy models of LLM inference can be accurate enough (R^2 > 0.96) to drive energy-optimal scheduling - methodology relevant to energy-per-token modeling. [doi:10.1145/3727200.3727217](https://doi.org/10.1145/3727200.3727217)
23. Geens, Robin et al. (2024). *Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators*. 2024 IEEE 37th International System-on-Chip Conference (SOCC). Abstract only. Design-space exploration framework for early identification of LLM inference energy bottlenecks; supports energy-per-token modeling before hardware exists [doi:10.1109/socc62300.2024.10737844](https://doi.org/10.1109/socc62300.2024.10737844)
24. Yang, Weisi & Xia, Stephen (2026). *PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling*. Proceedings of the 2026 ACM/IEEE International Conference on Embedded Artificial Intelligence and Sensing Systems. Abstract only. Adds workload-level power knobs (speculative decoding, verification depth) beyond DVFS for energy-efficient on-device LLM inference - relevant to edge tokens-per-watt optimization. [doi:10.1145/3774906.3802783](https://doi.org/10.1145/3774906.3802783)
25. Jahanshahi, Ali et al. (2023). *WattWiser: Power &amp; Resource-Efficient Scheduling for Multi-Model Multi-GPU Inference Servers*. Proceedings of the 14th International Green and Sustainable Computing Conference. Abstract only. GPU load consolidation for power reduction in SLO-bound inference serving; relevant to power-per-request optimization under latency constraints [doi:10.1145/3634769.3634807](https://doi.org/10.1145/3634769.3634807)
26. Wilhelm, Patrick (2025). *Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference*. Proceedings of the 5th Workshop on Machine Le. Full text read. Directly advocates Energy-per-Token as a standard efficiency metric complementing accuracy benchmarks - central to the tokens-per-watt review framing. [doi:10.1145/3721146.3721953](https://doi.org/10.1145/3721146.3721953)
27. Desislavov, Radosvet (2023). *Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning*. Sustainable Computing Informatics and Systems. Abstract only. Challenges the exponential energy-growth narrative for inference by accounting for hardware efficiency and implementation maturity. [doi:10.1016/j.suscom.2023.100857](https://doi.org/10.1016/j.suscom.2023.100857)
28. Stojkovic, Jovan (2024). *Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference*. arXiv (Cornell University). Full text read. Characterizes frequency, batching, and parallelism levers for energy-efficient LLM serving under performance SLOs - core input for tokens-per-watt operating-point selection. [doi:10.48550/arxiv.2403.20306](https://doi.org/10.48550/arxiv.2403.20306)
29. Stojkovic, Jovan (2025). *DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency*. 2025 IEEE International Symposium on High Per. Full text read. First energy-management framework for LLM inference clusters, quantifying how GPU frequency scaling, tensor parallelism and instance scaling trade off against energy-per-request under SLOs - direct evidence for cluster-level tokens-per-watt optimization (energy savings of 23-56%). [doi:10.1109/hpca61900.2025.00102](https://doi.org/10.1109/hpca61900.2025.00102)
30. Poddar, Soham et al. (2025). *Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models*. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics. Abstract only. Peer-reviewed NAACL 2025 study benchmarking LLM inference energy - exactly on topic, but full content must be fetched before use in the review. [doi:10.18653/v1/2025.naacl-long.632](https://doi.org/10.18653/v1/2025.naacl-long.632)
31. Jegham, Nidhal (2025). *How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference*. arXiv (Cornell University). Full text read. Prompt-level Wh/query benchmark across 30 models with facility-level multipliers and DEA efficiency ranking; the closest source to a 'tokens-per-watt'-style comparative benchmark [doi:10.48550/arxiv.2505.09598](https://doi.org/10.48550/arxiv.2505.09598)
32. Husom, Erik Johannes et al. (2025). *Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency*. ACM Transactions on Internet of Things. Abstract only. Hardware-level energy profiling of quantized LLMs on edge hardware; direct example of energy-per-task measurement methodology at the edge [doi:10.1145/3767742](https://doi.org/10.1145/3767742)
33. Ali, Sabiya Banu Masthan et al. (2026). *Assessing the Sustainability of LLM Inference through Energy–Accuracy Analysis*. Proceedings of the 17th ACM International Conference on Future and Sustainable Energy Systems. Abstract only. Empirical energy-accuracy trade-off mapping for open LLM inference at GPU level - a method template for energy-aware model selection. [doi:10.1145/3744255.3811741](https://doi.org/10.1145/3744255.3811741)
34. Chien, Andrew A. (2023). *Reducing the Carbon Impact of Generative AI Inference (today and in 2035)*. Proceedings of the 2nd Workshop on Sustainable Computer Systems. Abstract only. Early carbon-per-query request-routing comparison (Local vs Balance vs CarbonMin) for generative AI inference - carbon-aware direction angle. [doi:10.1145/3604930.3605705](https://doi.org/10.1145/3604930.3605705)
35. Rrapaj, Ermal (2024). *Power Consumption Trends in Supercomputers: A Study of NERSC's Cori and Perlmutter Machines*. ISC High Performance 2024 Research Paper Proceedings (39th International Conference). Abstract only. Facility-level evidence that production power sits far below TDP/peak provisioning - supports power-capped, over-provisioned designs relevant to inference datacenters. [doi:10.23919/isc.2024.10528943](https://doi.org/10.23919/isc.2024.10528943)
36. Zhao, Dan (2023). *Sustainable Supercomputing for AI*. ACM Symposium on Cloud Computing. Abstract only. Facility-scale evidence that power-capping AI accelerators cuts power draw and temperature with minimal performance cost - relevant to datacenter-level energy accounting for inference fleets. [doi:10.1145/3620678.3624793](https://doi.org/10.1145/3620678.3624793)
37. Billah, Waseq et al. (2026). *Is Power Usage Effectiveness (PUE) Still Fit for Purpose? An Integrative Review of Data Center Metrics for Assessing the Environmental Impact of Modern AI Infrastructure*. Journal of Science Policy &amp; Governance. Abstract only. Facility-level metric critique (PUE) relevant to the facility-level efficiency measurement debate for AI infrastructure. [doi:10.38126/jspg280101](https://doi.org/10.38126/jspg280101)
38. Lei, Nuoa & Masanet, Eric (2020). *Statistical analysis for predicting location-specific data center PUE and its improvement potential*. Energy. Abstract only. Statistical PUE-prediction framework relevant to converting inference energy into facility-level efficiency metrics and to PUE target-setting for AI data centers. [doi:10.1016/j.energy.2020.117556](https://doi.org/10.1016/j.energy.2020.117556)
39. Flores-Martin, Daniel et al. (2025). *Improving Energy Efficiency in a Data Center: PUE Analyzing and Tuning*. 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Abstract only. Facility-level PUE optimization via sensors and ML; provides context for facility-level efficiency beyond per-token metrics [doi:10.1109/ccgrid64434.2025.00028](https://doi.org/10.1109/ccgrid64434.2025.00028)
40. Aneli, Stefano et al. (2025). *Modelling and experimental surveys on the energy consumption of a small-scale data center*. Energy Efficiency. Abstract only. Facility-level angle: data-center cooling demand response modeled in energy terms, linking facility efficiency to grid flexibility rather than tokens/watt. [doi:10.1007/s12053-025-10357-7](https://doi.org/10.1007/s12053-025-10357-7)
41. Luccioni, Sasha et al. (2024). *Power Hungry Processing: Watts Driving the Cost of AI Deployment?*. The 2024 ACM Conference on Fairness Accountability and Transparency. Full text read. Foundational inference-phase energy benchmark establishing kWh (and g CO2eq) per 1,000 inferences as a comparison unit and quantifying the energy penalty of multi-purpose generative models - a key precursor to tokens-per-watt benchmarking. [doi:10.1145/3630106.3658542](https://doi.org/10.1145/3630106.3658542)
42. Luccioni, Alexandra Sasha et al. (2022). *Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model*. arXiv (Cornell University). Full text read. Canonical lifecycle carbon accounting of a 176B LLM, quantifying memory-bound, idle-dominated inference energy and advocating granular energy/carbon-intensity/PUE reporting. [doi:10.48550/arxiv.2211.02001](https://doi.org/10.48550/arxiv.2211.02001)
43. Patterson, David A. et al. (2021). *Carbon Emissions and Large Neural Network Training*. arXiv (Cornell University). Full text read. Foundational argument that energy use and CO2e should be first-class ML evaluation metrics; authors collaborated with MLPerf to include energy during training and inference. [doi:10.48550/arxiv.2104.10350](https://doi.org/10.48550/arxiv.2104.10350)
44. Strubell, Emma et al. (2019). *Energy and Policy Considerations for Deep Learning in NLP*. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Abstract only. Seminal training-phase energy/carbon quantification for NLP models that motivated subsequent inference-phase tokens-per-watt benchmarking. [doi:10.18653/v1/p19-1355](https://doi.org/10.18653/v1/p19-1355)
45. Wu, Carole-Jean et al. (2021). *Sustainable AI: Environmental Implications, Challenges and Opportunities*. arXiv (Cornell University). Full text read. Industry-scale characterization of training vs inference energy split and a call to add efficiency/environmental measures to MLPerf-style leaderboards, directly motivating tokens-per-watt-style benchmarking. [doi:10.48550/arxiv.2111.00364](https://doi.org/10.48550/arxiv.2111.00364)
46. Faiz, Ahmad et al. (2023). *LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models*. arXiv (Cornell University). Full text read. Pre-training carbon projection model covering training/inference/storage/embodied phases for dense and MoE LLMs; relevant as an estimation methodology for carbon-per-token style accounting [doi:10.48550/arxiv.2309.14393](https://doi.org/10.48550/arxiv.2309.14393)
47. Dodge, Jesse et al. (2022). *Measuring the Carbon Intensity of AI in Cloud Instances*. 2022 ACM Conference on Fairness, Accountability, and Transparency. Abstract only. Measurement-availability argument: carbon-intensity reporting by cloud providers as the prerequisite for emissions reduction - a measurement-infrastructure contribution. [doi:10.1145/3531146.3533234](https://doi.org/10.1145/3531146.3533234)
48. Henderson, Peter et al. (2020). *Towards the Systematic Reporting of the Energy and Carbon Footprints of\n Machine Learning*. arXiv (Cornell University). Full text read. Foundational standardized energy/carbon reporting framework (experiment-impact-tracker) and RL energy leaderboard; defines the reporting conventions that tokens-per-watt benchmarks should adopt [doi:10.48550/arxiv.2002.05651](https://doi.org/10.48550/arxiv.2002.05651)
49. Xu, Jingjing et al. (2021). *A Survey on Green Deep Learning*. arXiv preprint. Full text read. Provides the measurement-metrics discussion (FLOPs vs runtime vs carbon emission vs fair measure) that informs how a tokens-per-watt benchmark should be defined and reported. [doi:10.48550/arxiv.2111.05193](https://doi.org/10.48550/arxiv.2111.05193)
50. Jiang, Peng et al. (2024). *Preventing the Immense Increase in the Life-Cycle Energy and Carbon Footprints of LLM-Powered Intelligent Chatbots*. Engineering. Abstract only. Life-cycle (training + inference + embodied) energy/carbon framing for LLM chatbots; broadens scope beyond inference-only energy benchmarks [doi:10.1016/j.eng.2024.04.002](https://doi.org/10.1016/j.eng.2024.04.002)
51. Kaneko, Yusuke (2025). *Comparing Electricity Consumption Per Use of Blockchain and Generative AI*. IEEE Access. Abstract only. Provides system-level and per-use electricity numbers showing LLM inference energy (GPT-4 daily >500 MWh) can exceed training, motivating inference energy measurement. [doi:10.1109/access.2025.3573722](https://doi.org/10.1109/access.2025.3573722)
52. García-Martín, Eva et al. (2019). *Estimation of energy consumption in machine learning*. Journal of Parallel and Distributed Computing. Abstract only. Foundational review of energy-estimation approaches and tools for ML; grounds the measurement-methodology side of the review [doi:10.1016/j.jpdc.2019.07.007](https://doi.org/10.1016/j.jpdc.2019.07.007)
53. Sánchez-Mompó, Adrián et al. (2025). *Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations*. Information. Abstract only. Positions itself as a benchmark for estimating total AI energy use across model types in MLOps, emphasizing utilization-dependent LLM efficiency. [doi:10.3390/info16040281](https://doi.org/10.3390/info16040281)
54. Aquino-Brítez, Sergio et al. (2025). *Towards an Energy Consumption Index for Deep Learning Models: A Comparative Analysis of Architectures, GPUs, and Measurement Tools*. Sensors. Abstract only. Proposes a standardized, sensor-validated energy-efficiency index covering both training and inference across architectures - relevant as a measurement-methodology template. [doi:10.3390/s25030846](https://doi.org/10.3390/s25030846)
55. Rubei, Riccardo et al. (2025). *Prompt engineering and its implications on the energy consumption of Large Language Models*. 2025 IEEE/ACM 9th International Workshop on Green And Sustainable Software (GREENS). Abstract only. Shows prompt engineering itself as a lever on inference energy (relevant to energy-per-token optimization), but lacks quantitative measurements. [doi:10.1109/greens66463.2025.00014](https://doi.org/10.1109/greens66463.2025.00014)
56. Li, Yueying et al. (2025). *EcoServe: Designing Carbon-Aware AI Inference Systems*. arXiv preprint. Full text read. Shows carbon-optimal serving differs from energy-optimal; 4R (Reduce/Reuse/Rightsize/Recycle) framework with cross-stack ILP co-design. [doi:10.48550/arxiv.2502.05043](https://doi.org/10.48550/arxiv.2502.05043)
57. Li, Baolin et al. (2023). *Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference Service*. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Abstract only. Early carbon-SLA co-optimization for inference serving (quality mixing + GPU partitioning) rather than a measurement benchmark. [doi:10.1145/3581784.3607034](https://doi.org/10.1145/3581784.3607034)
58. Ding, Yi & Shi, Tianyao (2024). *Sustainable LLM Serving: Environmental Implications, Challenges, and Opportunities : Invited Paper*. 2024 IEEE 15th International Green and Sustainable Computing Conference (IGSC). Abstract only. Positions sustainable LLM serving (inference-phase carbon) as a research agenda - motivational framing for tokens-per-watt benchmarking. [doi:10.1109/igsc64514.2024.00016](https://doi.org/10.1109/igsc64514.2024.00016)
59. Wilkins, Grant (2024). *Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads*. The 15th ACM International Conference on Futu. Abstract only. Token-count-aware scheduling across heterogeneous hardware as a datacenter-level energy-efficiency lever for LLM serving. [doi:10.1145/3632775.3662830](https://doi.org/10.1145/3632775.3662830)
60. Du, Bojun et al. (2026). *From Tokens to Energy Flexibility: Quantization-Enabled Demand Response for Data Centers with LLM Inference Workloads*. arXiv preprint. Full text read. Bridges tokens and grid energy: a quantization-to-power mapping with per-token energy (J/token) and token throughput as dispatchable flexibility parameters for demand response. [doi:10.48550/arxiv.2606.18851](https://doi.org/10.48550/arxiv.2606.18851)
61. Séjourné, Kevin et al. (2026). *Sovereign LLM Inference in the Era of Digital Sobriety: A Comparative Study of Energy Efficiency Across Heterogeneous Architectures*. 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET). Abstract only. Extends per-request energy benchmarking to include embodied carbon and hardware-heterogeneity ('digital sobriety') decisions in sovereign AI infrastructure. [doi:10.1109/icecet65726.2026.11633138](https://doi.org/10.1109/icecet65726.2026.11633138)
62. Wang, Dan et al. (2025). *StoreLLM: Energy Efficient Large Language Model Inference with Permanently Pre-stored Attention Matrices*. Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems. Abstract only. Radical energy-vs-compute substitution idea (replace attention-matrix computation with storage access) claiming large inference energy reductions. [doi:10.1145/3679240.3734604](https://doi.org/10.1145/3679240.3734604)
63. Yilk, Todd (2018). *Qualifying for the Green500: Experience with the newest generation of supercomputers at LANL*. Sustainable Computing: Informatics and Systems. Abstract only. Documents institutional practice of qualifying HPC systems on the Green500 performance-per-watt benchmark - a methodological precedent for tokens-per-watt-style efficiency qualification. [doi:10.1016/j.suscom.2018.02.004](https://doi.org/10.1016/j.suscom.2018.02.004)
64. Chen, Huamin et al. (2026). *The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency*. arXiv preprint. Full text read. Core theoretical contribution for the review: context window as the dominant tok/W lever (1/W law) plus quantitative decomposition of routing-topology vs hardware-generation energy gains. [doi:10.48550/arxiv.2603.17280](https://doi.org/10.48550/arxiv.2603.17280)
65. Schwartz, Roy et al. (2020). *Green AI*. Communications of the ACM. Abstract only. Provides the normative argument ('Green AI') for treating energy efficiency as a first-class evaluation criterion - the conceptual basis for tokens-per-watt benchmarks. [doi:10.1145/3381831](https://doi.org/10.1145/3381831)
