On this page

Benchmarking tokens per watt: how AI inference energy efficiency is measured

What benchmarks and measurement methodologies exist for quantifying AI inference energy efficiency in tokens per watt (or per joule), from accelerator to data-centre level, and what do they establish?

Updated
17 Aug 2026
Sources
65
Years
2009–2026
Confidence
Download Markdown

tokens per wattenergy per tokenMLPerf PowerLLM inference energypower measurementPUEcarbon per querybenchmarking

How this review was made
Databases
OpenAlex, Crossref, arXiv API, arXiv abs pages, Semantic Scholar (best-effort), Unpaywall, ACL Anthology, Zenodo and GitHub (AI Energy Star probe)
Queries (literal)
tokens per watt
tokens per joule
tokens per kilowatt hour
energy per token large language model
LLM inference energy benchmark
MLPerf power energy efficiency machine learning
AI energy star benchmark
benchmark energy consumption LLM inference
energy efficiency large language model inference benchmark
inference energy measurement transformer
GPU power consumption measurement deep learning inference
power metering neural network inference
energy accounting GPU fine-grained
power disaggregation data center application
wattmeter deep learning training measurement
AI data center energy efficiency PUE
generative AI data center energy consumption
machine learning inference energy data center
data center energy efficiency metric AI workload
carbon footprint generative AI inference energy
ChatGPT energy consumption estimate
energy consumption generative AI models
sustainable machine learning inference energy
energy efficient transformer inference
power measurement large language model
speculative decoding energy efficiency
LLM inference energy scaling
green AI benchmark
MLPerf Inference Benchmark
MLPerf Tiny Benchmark
Power Hungry Processing: Watts Driving the Cost of AI Deployment
Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
The growing energy footprint of artificial intelligence
Carbon Emissions and Large Neural Network Training
Energy and Policy Considerations for Deep Learning in NLP
Green AI
Sustainable AI: Environmental Implications, Challenges and Opportunities
LLMCarbon: Modeling the end-to-end carbon footprint of large language models
A Survey on Green Deep Learning
EcoServe: Designing Carbon-Efficient Inference Serving for Large Language Models
Identifying Shades of Green: The SPECpower Benchmarks
Sustainable AI: Environmental Implications, Challenges and Opportunities (Crossref/arXiv/OpenAlex)
EcoServe (arXiv ti)
Green500 (Crossref/arXiv/OpenAlex)
AI Energy Star (Zenodo/arXiv/OpenAlex/GitHub — no records)
Weight-Only Quantization Does Not Always Save Energy (arXiv ti — none)
Power Hungry Processing Watts Driving (arXiv ti)
DynamoLLM (arXiv ti)
The 1/W Law (pool resolution)
Search last run
2026-08-17
Screening
65 sources used · 2009–2026 · deep review

Summary

The short version

The literature on measuring AI inference efficiency in tokens-per-watt terms is young and fragmented. One consortium standard exists — MLPerf Power, which collected 1,841 measurements across 60 systems from microwatts to megawatts 1 — but its workload set predates the generative-LLM era, and the benchmarks that actually report tokens per watt or joules per token are research tools with incompatible boundaries: per-GPU, per-node, or per-request; with or without power meters; over different workloads 81841. The measured numbers span orders of magnitude — about 3–4 joules per output token for a 65B model on A100s 18, 0.002 to 2.9 kWh per 1,000 inferences depending on task 41, a 65× spread across models in commercial data centres 31, and a proposed “1/W law” under which tokens per watt halves every time the context window doubles 64. No retrieved benchmark measures tokens per watt at the data-centre level; facility efficiency is still expressed as PUE, so wall-level tokens per watt is derived from server numbers, not measured 37. Confidence is moderate: 20 of 65 sources were read in full text, and several prominent items — including the widely cited Joule estimate paper and two SSRN preprints — were unreachable in-session.

Why this question

Tokens per watt has become the efficiency currency of AI infrastructure: accelerator purchases, serving-stack choices, power-capping decisions, and data-centre designs are argued in these units. But a unit is not a measurement. What the metric denotes (one GPU or one node or one facility; decode-only or end-to-end; one request length or another), how it is taken (a calibrated wattmeter, a software power model, or vendor telemetry), and on which workload it is taken all change the number by orders of magnitude. An architect comparing H100, B200, and TPU efficiency figures, or deciding whether power capping pays for a decode-heavy workload, is comparing numbers produced by different metrologies — and the published benchmarks are the only place those metrologies are described well enough to be compared.

What turns on the answer is expensive. The prior review in this series (model-to-grid AI datacenter efficiency) established what the efficiency levers are and how strong the evidence for each is. This review asks the measurement question underneath it: when a number like “tokens per watt” is reported — in a paper, a submission, or a vendor slide — what was actually measured, and can it be compared with any other number?

Scope and methods

Question. What benchmarks, metrics, and measurement methodologies exist for quantifying AI inference energy efficiency in tokens per watt (or per joule), from accelerator to data-centre level, and what do they establish about how inference efficiency should be measured and compared?

Inclusion criteria. Works published or posted 2019–2026 (the modern inference-efficiency literature), plus two metric-antecedent sources from 2009 and 2018 (SPECpower, Green500) because the modern metric descends from them; English; peer-reviewed papers and arXiv preprints (this fast-moving field lives partly in preprints); benchmark descriptions, measurement studies, energy- and carbon-per-inference quantifications, and metric or methodology critiques. Exclusion criteria. Training-only efficiency studies except where they define the metric landscape (Zeus, Patterson, Strubell are included for the measurement lineage, not as tokens/watt evidence); power-grid engineering; efficiency levers (power capping, scheduling, serving configuration) as such — already covered by the model-to-grid review — except where they report tokens/watt measurements; microarchitecture-only work; works unreachable through any open archive (see below).

Databases and dates. OpenAlex, Crossref, the arXiv API, arXiv abs pages, Semantic Scholar (best-effort; persistent 429s), Unpaywall for OA checks, the ACL Anthology for abstract rescue, and Zenodo/GitHub for a vendor-benchmark probe. Forty-eight query strings (listed in frontmatter) plus a 13-title landmark-resolution ladder (Crossref query.bibliographic → arXiv ti:) and citation-graph follow-ups. Searches run 2026-08-17.

Screening counts. 1,948 raw records retrieved (1,142 from OpenAlex, 806 from arXiv/Crossref); ~1,566 after DOI/title deduplication; 210 passed an automated topic-and-venue screen; 13 landmark titles were resolved through the resolution ladder (12 resolved cleanly, one resolved to a different but real paper that was kept on its own merits); 66 records were curated into the final selection. One source was dropped at retrieval — the Joule commentary “The growing energy footprint of artificial intelligence” (de Vries) — because every legitimate channel (cell.com, ScienceDirect, the institutional repository, OpenAlex, Crossref, Semantic Scholar) is bot-walled or abstract-less in-session; per this review’s protocol, unfetchable sources are not cited. Two SSRN preprints were dropped at screening (a weight-only-quantization energy study and “Green My LLM”) because SSRN blocks scripted access. A probe for the vendor benchmark “AI Energy Star” found no paper, DOI, Zenodo record, or GitHub repository — it exists only as marketing material, which is not citable under this review’s rules; the absence is noted in Gaps. 65 sources included, all retrieved this session: 20 read in full text, 45 at abstract level (flagged abstract-only and hedged accordingly). Every DOI was verified against Crossref (title similarity 1.00 on all checks; no retraction notices) or, for arXiv DOIs, by abs-page resolution; every 10.48550/arxiv.* DOI was confirmed to resolve.

What this review deliberately does not cover. Training-energy benchmarking as a topic (covered elsewhere in this series); the efficiency-lever evidence base (power capping, scheduling — see the model-to-grid review); vendor marketing figures with no citable artifact.

The landscape

The literature has three neighbourhoods that barely cite each other. The consortium-benchmark neighbourhood (MLPerf family, SPEC, Green500) defines what “efficiency benchmark” means institutionally: rules, submissions, comparability, and rankings in FLOPS/watt or queries-per-watt. The research-measurement neighbourhood (2023–2026) is where tokens-per-watt-style numbers actually live: wattmeter and telemetry studies of LLM inference on specific GPUs, serving engines, and cloud APIs. The carbon-accounting neighbourhood (FAccT, lifecycle-analysis venues) measures the same physical quantity but reports it in grams of CO₂ per query or per day, folding in grid carbon intensity and embodied manufacturing.

A vocabulary finding shapes everything below: the exact phrase “tokens per watt” appears rarely in the peer-reviewed literature. The field speaks in joules per token 826, watt-hours per request 29, kWh per 1,000 inferences 41, and tokens per watt in the few papers that use the unit directly 64. These are different quantities — a per-GPU decode-phase figure is not comparable with an end-to-end per-request figure — and converting between them requires request lengths, batch sizes, and measurement boundaries that most papers report incompletely. The unit heterogeneity is itself a comparability barrier, and no benchmark in the retrieved set standardises the conversion.

Theme 1: The metric lineage — from FLOPS per watt to energy per token

The efficiency metric did not start with language models. Green500 institutionalised FLOPS/watt rankings for supercomputers, with measurement-qualification rules that labs must satisfy to be listed 63. SPEC established SPECpower_ssj2008 as the first industry-standard server power-and-performance benchmark 6. The “Green AI” agenda then made efficiency a first-class evaluation criterion for machine learning, documenting an estimated 300,000× growth in deep-learning compute from 2012 to 2018 65. Strubell, Ganesh and McCallum supplied the first systematic energy accounting for NLP — including inference, not just training — and the policy recommendations that followed 44.

The MLPerf family carries this lineage into machine learning. The training suite was characterised in a peer-reviewed study of its workloads versus earlier benchmarks such as DAWNBench and DeepBench 2; the inference benchmark spans three orders of magnitude in system power and five in performance, and its closed division enforces latency constraints precisely because throughput-only comparisons mislead 3. The mobile and tiny variants extended the family down the power axis — MLPerf Tiny made energy a first-class reported metric alongside accuracy and latency 45.

Against this background, the token-based unit crystallised in 2024–2026. A workshop paper argued that Energy-per-Token should complement accuracy benchmarks, showing that test-time-compute strategies can trade enormous energy for accuracy (chain-of-thought on MATH raised accuracy by 281% at a reported +15,132% energy cost) and that smaller models with controlled reasoning beat larger models on an energy-accuracy frontier (a 40–60% energy saving at similar or better accuracy) 26. TokenPowerBench formalised joules-per-token as its headline metric 8, and the 1/W law paper defined tokens/watt explicitly, decomposing it into single-GPU and fleet-level quantities 64. A survey of green deep learning catalogued the same landscape from the training side and flagged the recurring gap between FLOPs-based and runtime-based efficiency claims 49, and an analysis of inference-energy trends argued that inference energy follows different laws than the performance-versus-parameter scaling curves 27.

Theme 2: The benchmarks that measure inference efficiency

MLPerf Power is the closest thing to a standard. Developed by a consortium of more than 20 organisations, it establishes rules and best practices for measuring ML-system energy efficiency from microwatts to megawatts, and reports 1,841 reproducible measurements across 60 systems 1. Its significance is institutional: it is the only benchmark in this review whose numbers are produced under agreed rules rather than per-lab conventions. Its limitation is scope: its workloads are the classic MLPerf inference set, not generative-LLM serving, so tokens/watt for chat workloads is outside its remit.

The MLPerf-based evaluations that do exist are instructive about comparability. One study measured four accelerators (NVIDIA A100, Intel Gaudi, Graphcore Bow-Pod64, Groq LPU) on two MLPerf workloads to a common target accuracy: which accelerator is most efficient depends on the workload — Gaudi2 had the lowest energy on ResNet50, the A100 on BERT-Large inference 11. A 2026 study of 739 MLPerf Inference v5.0 submissions found no significant public-versus-private-cloud efficiency difference but large accelerator-model differences (the H200 showing roughly 45% higher median throughput-per-TDP efficiency), concluding efficiency scales multiplicatively across accelerator generations 10. Both results cut against any single “best accelerator” narrative.

The research benchmarks report the tokens-per-watt-style numbers. TokenPowerBench measures GPU-, node-, and system-level power without specialised meters, attributes energy to prefill and decode phases per request, and reports joules per token across four model families from 1B to 405B parameters, with batch size, context length, parallelism, and quantisation as configurable dimensions 8. GPU-NEST established an energy-efficiency characterisation methodology for multi-GPU inference servers and showed power as a first-order constraint in them 7, and WattWiser targeted power- and resource-efficient scheduling for multi-model multi-GPU inference servers 25. From Words to Watts benchmarked LLaMA models on V100 and A100 clusters: a 65B model consumed roughly 3–4 joules per output token at 512-token generation, and power-capping the GPUs from 250 W to 175 W cut total energy by 23.21% at a 6.7% time penalty 18. How Hungry is AI benchmarked 30 models in commercial data centres through their APIs, combining public performance data with environmental multipliers: the most energy-intensive models exceeded 29 Wh per long prompt — over 65× the most efficient — and a 0.42 Wh short query, scaled to 700 million queries per day, aggregates to annual electricity comparable to 35,000 US homes 31. A NAACL 2025 study benchmarked inference energy across NLP tasks and found energy correlates strongly with output-token length and response time, with quantisation, optimal batch sizes, and prompt phrasing as significant levers 30. An IEEE Access 2024 profiling study reached a similar conclusion from the architecture side: model size, layer count, parallelised attention, and vocabulary size drive inference energy, while batch size and quantisation level are the tunable levers 19. A cluster-redesign study for GPT-Neo reported substantial throughput, latency, and energy improvements from a novel interconnect and architecture, though without retrievable magnitudes 20.

A second strand benchmarks the serving engines rather than the models. One study decomposed LLM inference engines (vLLM, TensorRT-LLM, DeepSpeed) on a two-H100 node into setup and token-generation stages to locate energy bottlenecks 9. Energy-cost models extend this: a workload-based model fitted energy and runtime across heterogeneous CPU-GPU systems with R² above 0.96 22, and a token-count-aware allocation strategy across heterogeneous clusters saved 7.5% of energy versus a workload-unaware baseline 59. On the edge, 28 quantised LLMs were measured on a Raspberry Pi 4 32, an energy-aware DVFS-plus-speculative-decoding scheme reported a 52.4% energy reduction on-device 24, and measurements on Apple Silicon found energy per token grows monotonically but weakly sublinearly with prompt length 13 — with Apple Silicon reported at 4–7× less energy per request than NVIDIA-based servers on interactive workloads in one comparison 61. Model-level energy estimates round out the set: aggressive weight-only quantisation cut Llama-2-7B inference energy by 4.6× and moved the bottleneck to attention 23, and pre-storing attention matrices cut energy 1.45–2.83× against baselines 62.

Theme 3: The metrology — how the numbers are taken

Underneath the benchmarks sits a measurement-methodology literature, and it is where the comparability problems live. Physical power meters are accurate but expensive and cannot attribute energy to individual services; software-based power meters (power models plus vendor interfaces) are deployable at scale, and a 2023 study compared the leading tools for CPU and GPU measurement, concluding the choice of tool materially changes what you can claim 14. A review of energy-estimation approaches catalogued the same trade-off for machine learning generally 52, and a 2025 MLOps study applied software-based power measurement across discriminative and generative pipelines, concluding architecture and hardware choices dominate pipeline energy 53. An energy-consumption-index proposal aims to standardise cross-architecture DL comparison, though it has been validated on CNNs rather than LLMs 54. On the attribution side, WattScope estimates application-level power from a server’s aggregate draw without OS access, exploiting the low-variability, periodic power signatures of data-centre workloads 16, and a EuroSys 2026 paper tackles the harder version — attributing GPU power to individual jobs in shared cloud settings 15.

Phase attribution is the emerging standard for LLM-specific measurement. TokenPowerBench aligns metrics to prefill versus decode per request 8; the engine study separates setup from token generation 9; and the CCGrid 2026 characterisation shows why phase matters — decode-phase frequency scaling from 2842 MHz to 180 MHz saves 42% of energy at only 1–6% latency cost, while prefill behaves differently 12. Scale matters for methodology too: the BUTTER-E dataset, 63,527 watt-metered runs, found GPU training used a median 9.47 mJ per datum versus 6.16 mJ on CPU in its setting, and that a fitted energy model predicts GPU energy within ±2.74% but CPU energy only within ±20.2% 17.

Reporting practice lags all of this. A 2020 audit found that only about 1 in 100 sampled NeurIPS papers reported energy at all, and proposed a systematic reporting standard; the audit’s own experiments consumed 24.344 kWh and 8.021 kg CO₂eq 48. The configuration space is the final confound: power capping is simultaneously an optimisation lever and a measurement choice. The same GPU measured at different caps produces different tokens-per-watt — 250 W versus 175 W changed total energy by 23.21% on one workload 18, and running at 1.6 GHz instead of 2.0 GHz held throughput constant at under 80% of the energy 28 — so an efficiency number without its power configuration is almost meaningless. Idle power compounds this: the 1/W law analysis shows the B200’s 430 W idle draw becomes a dominant fraction of the bill at long context, where KV-cache limits cut concurrency 64.

Theme 4: What the measurements show

The verified numbers cluster into four findings.

First, per-token energy at the server level is in the single-digit joules on current data-centre GPUs, with a wide spread. A 65B LLaMA on A100s: roughly 3–4 J per output token 18. Llama-2-70B serving: 0.77 to 13.21 Wh per request depending on request class, tensor parallelism, and frequency 29. Per query, 2.9 J for a 1B model rising to 21.0 J for a 32B model — a 7.2× model-size spread 12.

Second, the task and workload spread dwarfs the hardware spread. Per 1,000 inferences, energy ranges from 0.002 kWh for text classification to 2.9 kWh for image generation — a factor above 1,450 — with multi-purpose generative models orders of magnitude costlier than task-specific ones 41. In commercial data centres, the model spread is 65× on long prompts 31. Anyone comparing “tokens per watt” across studies is comparing numbers that vary more with workload than with hardware generation.

Third, inference, not training, dominates steady-state energy — which is the entire justification for the metric. Industry reports cited by TokenPowerBench put inference above 90% of total power consumption 8; a comparison study estimated GPT-4’s training at ~9,450 MWh versus inference exceeding 500 MWh per day 51; BLOOM’s API deployment was estimated at ~19 kg CO₂eq per day versus 24.7–50.5 tCO₂eq for training 42; and the serving phase has now surpassed training in energy terms per one 2024 assessment 58.

Third-and-a-half, model size and accuracy do not move together on energy. Across 14 open-source LLMs (6B–34B) on an A100, larger models often consumed substantially more energy without proportional accuracy gains, with mid-sized models matching larger ones at much lower energy 33. Prompt-engineering choices also change the number: prompt phrasing measurably altered Llama 3’s inference energy on code generation 55.

Fourth, the measured levers are large but regime-dependent. Power/frequency capping: 23.21% energy saved at 6.7% time cost 18, ~20% power reduction with zero latency impact 28, 42% via decode-phase DVFS 12, and near-zero impact at a supercomputing centre where jobs were not power-bound 36. Routing small queries to small models: 80% energy saving at a 6.8-point quality cost, 88% combined with DVFS 12. Cluster-level design: 53% energy, 38% operational carbon, and 61% customer-cost reduction under latency SLOs in DynamoLLM’s evaluated design 29.

A related but distinct literature reports carbon per query, which is the same measurement problem with a grid-intensity multiplier: LLMCarbon models end-to-end carbon within ≤8.2% of measured footprints 46; EcoServe’s carbon-aware design cut emissions by up to 47% with embodied carbon exceeding 50% of lifetime in its analysis 56; Clover reduced inference-service carbon under SLA constraints 57; and cloud-side estimates argued that provider-reported software carbon intensity is the missing prerequisite for mitigation 47. Patterson et al. quantified the upstream lever space — efficiency and siting together can reduce training carbon by 100–1,000× 43 — and a longitudinal analysis found embodied carbon roughly equal to operational footprint and ~20% operational-power reduction per six months of hardware improvement, at an assumed PUE of 1.1 45. A workload model of ChatGPT-class inference compared request-routing policies — local, balanced, and carbon-minimising — for their power and carbon impacts 34, and a life-cycle analysis of LLM-powered chatbots identified eight energy- and carbon-relevant phases with three strategic mitigation pathways 50. At the serving-system level, a cloud-scale characterisation of ML serving reported up to 28.3% energy savings from hierarchical GPU power management across 105 servers 21.

Theme 5: The facility-level gap — tokens per watt for data centres

The review’s central negative finding: no retrieved benchmark or measurement study reports tokens per watt at the data-centre level. The facility literature measures efficiency in PUE — the ratio of total facility energy to IT energy. Studies tune PUE with sensor-plus-ML methodology 39, model PUE statistically across locations 38, measure a small data centre end-to-end including cooling and demand response 40, and — most tellingly — an integrative review asks whether PUE is still fit for purpose for AI infrastructure at all 37. Production-scale power trends are tracked at the facility level (NERSC’s machines stayed well below provisioned power even after GPU transitions 35), and supercomputing centres manage power draw through capping 36.

The bridge between tokens and facilities exists only in fragments. MLPerf Power spans microwatts to megawatts but measures systems, not facilities 1. One 2026 preprint treats tokens served as the dispatchable unit for data-centre demand response, cutting operating cost 34.3% without curtailing token volume 60. And the facility metric that would complete the chain — tokens per watt at the wall, i.e. server-level tokens/watt divided by facility overheads — appears nowhere as a measured quantity. Every published number in this review stops at the server or node boundary; the data-centre boundary is reached only by arithmetic.

Where the evidence disagrees

What the metric should optimise. The energy-per-token advocates argue efficiency belongs alongside accuracy in benchmark scores 26; the MLPerf tradition treats efficiency as one dimension of a comparability regime built for performance 1; the carbon literature argues neither energy nor performance is the right objective — grid intensity and embodied carbon are 565747. These are disagreements about objectives, not numbers, but they change which benchmark you build.

Does quantisation save energy? Yes, measured 4.6× on Llama-2-7B with weight-only quantisation 23, and quantised edge models are the default for efficiency 32. But the one peer-reviewed-adjacent counter-evidence — an empirical study arguing weight-only quantisation does not always save energy across NVIDIA platforms — exists only as an SSRN preprint that was unreachable in-session and is therefore not cited here. The tension is real in the field and unresolved in this review.

Which engine or accelerator wins? Rankings do not transfer across workloads: Gaudi2 beats A100 on ResNet50 energy but loses on BERT inference 11; engine efficiency depends on whether you measure setup or generation phase 9; and the accelerator-generation effect dominates any cloud-type effect in the MLPerf submission record 10. Any single ranking is an artefact of its workload.

Power capping: free lunch or nothing? Large savings on some workloads 1828, 42% on decode with 1–6% latency cost 12, and minimal impact at a centre whose jobs were not power-bound 36. The disagreement tracks workload memory-boundedness — the same regime dependence the model-to-grid review found — and reinforces that tokens-per-watt numbers are configuration-relative.

Energy versus carbon. Optimising energy is not optimising carbon: the grid’s marginal intensity, not the GPU’s watt draw, decides emissions 56574347. Studies that report both (DynamoLLM’s 53% energy versus 38% carbon savings 29) show the gap explicitly.

Gaps and open questions

  • No facility-level tokens/watt benchmark exists. Nothing in the retrieved literature measures tokens per watt at the data-centre boundary. Settling this requires a consortium-standard benchmark that measures LLM serving at node and facility level with a fixed boundary, workload mix, and phase attribution — the MLPerf Power model extended to generative serving and to the wall.
  • No consortium benchmark covers generative-LLM serving energy. MLPerf Power’s workloads predate it; TokenPowerBench is research-grade; the commercial-DC numbers come from API-side estimation, not independent measurement 31. The field is benchmarking around the edges of the workload that matters.
  • No independent cross-cloud measurement. The largest comparative dataset (739 submissions) is vendor-submitted 10; the cloud-carbon literature argues providers must publish software carbon intensity for anyone to act 47.
  • Measurement boundaries are unstandardised. GPU-only versus node versus system; idle power included or excluded; phase attribution present or absent; power configuration reported or not. Henderson’s proposed reporting standard 48 has not been adopted by any major benchmark.
  • Vendor benchmarks are uncitable. The “AI Energy Star”-style figures circulate in marketing with no paper, DOI, or repository; this review can neither verify nor cite them, and the ecosystem’s reliance on them is itself a gap.
  • Idle power and long-context behaviour are under-measured. The 1/W law 64 suggests the efficiency cliff at long context is physical (KV-cache concurrency), but it is one preprint’s analytical model — H100-calibrated, B200-scaled — awaiting independent measurement.

Confidence and limitations

Confidence is moderate. The 20 full-text sources carry the quantitative claims in Themes 2 and 4, and every load-bearing number was re-verified against the retrieved full text or abstract before citation. The 45 abstract-only sources (mostly paywalled IEEE and ACM items) are cited for qualitative claims only. Specific limitations: (1) the de Vries Joule estimate — probably the most-cited per-request energy figure in public discourse — could not be retrieved through any legitimate channel and is not cited; (2) two SSRN preprints (weight-only-quantisation energy study, “Green My LLM”) are excluded for the same reason, leaving a known counter-voice out of the quantisation disagreement; (3) the “AI Energy Star” probe found no citable artifact, which is evidence about the vendor-benchmark ecosystem but not proof of absence; (4) English-only sources, 2009–2026 with the two metric-antecedent exceptions noted; (5) the facility-level gap is an absence finding — no retrieved source measures it, which is strong evidence of a gap but not proof none exists.

Jump to references ↓

Evidence table

keydesignsamplemeasurefindinglimitationsconfidenceaccessnote
ali2026assessingbenchmark14 open-source LLMs (6B-34B params) on 8 benchmark datasets spanning code generation, summarization, mathematical reasoning, and QA; NVIDIA A100 (80 GB) GPU with fine-grained GPU-level power samplingInference-time energy consumption vs accuracy (energy-accuracy trade-off)Across 14 open-source LLMs (6B-34B) on an A100 80GB GPU, increasing model size often raises inference energy consumption significantly without proportional accuracy gains, mid-sized models achieve comparable or superior accuracy at substantially lower energy, and summarization is the most energy-intensive task domain while QA is the least (no absolute energy numbers reported in the abstract).Single GPU platform (A100 80GB) and controlled lab conditions; trade-offs may not generalize to other hardware, serving stacks, or batch configurations; abstract-only access limits methodological detail.moderateabstract-onlyEmpirical energy-accuracy trade-off mapping for open LLM inference at GPU level - a method template for energy-aware model selection.
aneli2025modellingsimulationOne existing data center (experimental survey) with a numerical energy model built in TRNSYS; demand-response scenario exploiting indoor air temperature/humidity fluctuationsData-center electricity consumption, specifically cooling-load reduction under demand responseThe proposed demand-response flexibility scenario reduces the data center's electricity needs for cooling by about 30% in the calibrated and validated TRNSYS model of an existing data center.Single existing data center; results are model-based (TRNSYS) rather than field-measured DR events; the 30% figure applies to cooling load, not whole-facility electricity.moderateabstract-onlyFacility-level angle: data-center cooling demand response modeled in energy terms, linking facility efficiency to grid flexibility rather than tokens/watt.
aquino2025energybenchmarkAlexNet, ResNet18, VGG16, EfficientNet-B3, ConvNeXt-T, and Swin Transformer trained on Imagenette; TITAN XP and GTX 1080 GPUs; OpenZmeter v2, CarbonTracker v1.2.5, CodeCarbon v2.4.1Newly developed energy consumption index for DL models over training and inference; energy efficiency across architectures and GPUsUsing sensor-based (OpenZmeter) and software-based (CarbonTracker, CodeCarbon) tools on TITAN XP and GTX 1080 GPUs, the study reports significant differences in energy efficiency across architectures and GPUs, but no specific energy numbers are given in the abstract.Older GPU generation (TITAN XP/GTX 1080) and small Imagenette dataset; applicability of the index to LLM-scale workloads unshown; no absolute energy values in abstract.moderateabstract-onlyProposes a standardized, sensor-validated energy-efficiency index covering both training and inference across architectures - relevant as a measurement-methodology template.
argerich2024measuringbenchmarkSeveral state-of-the-art LLMs deployed for inference; varied model size, layer count, parallelized attention, and vocabulary size; input batch size and quantization levels as optimization leversEnergy consumption of LLM inference, inference energy efficiency, and latencyProfiling several state-of-the-art LLMs during inference shows that model size, layer count, parallelized attention, and vocabulary size drive inference energy consumption, while input batch size and quantization levels can be tuned to improve inference energy efficiency and latency (no numeric values reported in the abstract).Abstract gives no numeric results or hardware details; measurement scope (which components are counted) and profiling methodology are not described in the abstract.moderateabstract-onlyInference-focused LLM energy profiling methodology with quantization and batching as efficiency levers - complements training-centric carbon studies.
banbury2021mlperftinybenchmarkFour TinyML benchmarks: keyword spotting (Speech Commands v2, DS-CNN with 38.6K params), visual wake words (MSCOCO-2014, MobileNetV1 325KB), image classification (CIFAR-10, ResNet 96KB), anomaly detection (ToyADMOS/DCASE2020, FC-AutoEncoder 270KB); reference implementations on NUCLEO-L4R5ZI with TFLM; EMON energy monitor with electrical-isolation proxyAccuracy, latency, and energy of ML inference on ultra-low-power (sub-milliwatt) TinyML systemsMLPerf Tiny v0.5 measures accuracy, latency, and energy for four ultra-low-power inference benchmarks with quality targets of 90%/80%/85% top-1 accuracy and 0.85 AUC (reference models reaching e.g. 92.2% KWS accuracy), using an EMON energy monitor on systems drawing under a milliwatt, but no aggregate tokens/watt or absolute energy figures are reported.Authors flag benchmark evolution and long-term stability as limitations; measurement-scope questions (what counts in the power measurement, data-path and pre-processing variability) and exclusion of feature extraction from KWS measurement remain open methodology issues.highfull-textFoundational energy-inclusive ML benchmark: makes energy a first-class metric alongside latency and accuracy for ultra-low-power inference - a key reference for energy-in-benchmark methodology.
beitsayadeh2026benchmarkingcohort739 standardized submissions from MLPerf Inference v5.0 across public and private cloud infrastructures and multiple accelerator models (incl. NVIDIA H200-SXM-141GB); efficiency estimated as throughput normalized by vendor-reported TDPInference throughput per watt (TDP-normalized efficiency) across clouds and acceleratorsIn a cross-sectional secondary analysis of 739 MLPerf Inference v5.0 submissions, throughput-per-TDP efficiency showed no significant cloud-type effect but significant accelerator-model differences - NVIDIA H200-SXM-141GB had approximately 45% higher median efficiency - with a log-linear specification fitting best, implying multiplicative efficiency scaling across accelerator generations.Efficiency is a proxy (throughput / vendor-reported TDP), not measured power; MLPerf submissions are self-selected; no absolute power or energy-per-token values.moderateabstract-onlyDemonstrates a reproducible performance-per-watt benchmarking framework built entirely on open MLPerf data - directly relevant to tokens-per-watt standardization.
billah2026ispuesurveyn/a - integrative review of data-center metrics, centered on Power Usage Effectiveness (PUE), for assessing environmental impact of modern AI infrastructureFitness-for-purpose assessment of PUE and alternative data-center environmental metricsNo finding extractable - the available abstract is only a teaser for an integrative review questioning whether Power Usage Effectiveness (PUE) is still fit for purpose as a metric for modern AI infrastructure; no numbers are reported.Fetched abstract is a truncated teaser with no substantive content; method and conclusions cannot be assessed.lowabstract-onlyFacility-level metric critique (PUE) relevant to the facility-level efficiency measurement debate for AI infrastructure.
chen2026oneoverwtheoreticalLlama-3.1-70B, Llama-3.1-405B, Qwen3-235B-A22B, DeepSeek-V3 on H100-SXM5 and B200-SXM (TP=8, fp16); Azure LLM Inference Trace and LMSYS-Chat-1M workloads; inference-fleet-sim queueing framework, logistic GPU power model, AIConfigurator rooflineTokens per watt (tok/W) at single-GPU and fleet levelThe derived 1/W law states tokens per watt halves each time the serving context window doubles - e.g., Llama-3.1-70B on H100 drops from 35.0 tok/W at 2K context to 1.5 tok/W at 64K (roughly 40x spread across 2K-128K) - while two-pool context-length routing (FleetOpt) delivers about 2.5x fleet tok/W (14.08 vs 5.58 tok/W on the Azure trace), H100-to-B200 upgrade about 1.7x, combined 4.25x, and Qwen3-235B-A22B reaches about 37.8 tok/W at 8K on H100 (5.1x Llama-3.1-70B, an upper bound).Authors flag: no empirical validation on B200/H200 (analytical projections with +/-20% uncertainty), MoE dispatch overhead excluded (upper bounds), steady-state traffic assumption, two-pool topology only, static fleet sizing, single workload CDF, and output-only energy accounting.highfull-textCore theoretical contribution for the review: context window as the dominant tok/W lever (1/W law) plus quantitative decomposition of routing-topology vs hardware-generation energy gains.
chien2023reducingsimulationWorkload model of ChatGPT-style generative AI inference; request-direction policies: Local, Balance, CarbonMinPower use and carbon impacts of generative AI inference under different request-direction approachesThe ChatGPT-exemplar workload model compares request-direction approaches (Local, Balance, CarbonMin) assessing their power use and carbon impacts, but the abstract reports no numeric results.Very short abstract with no results; workload-model validity and assumptions unverifiable from abstract alone.lowabstract-onlyEarly carbon-per-query request-routing comparison (Local vs Balance vs CarbonMin) for generative AI inference - carbon-aware direction angle.
csikai2026scalingbenchmarkApple M4 Pro system with a consistent local serving stack and macOS power telemetry; four open-source model families spanning 3B-70B parameters; six prompt categories from 128 to 4096 input tokensToken-normalized energy per generated token (CPU and GPU contributions), power traces sampled at fixed interval and integrated over the model-reported evaluation windowEnergy per token increases monotonically with input prompt length but with weak, sublinear growth over the evaluated range, and log-log fits yield compact scaling parameters separating context-length effects from model-specific baseline efficiency (no absolute energy values reported in the abstract).Single hardware platform (M4 Pro) and runtime path; absolute energy-per-token numbers not available in the abstract; controlled lab conditions may not reflect production serving.moderateabstract-onlyContributes an Apple Silicon-appropriate measurement workflow and empirical scaling laws for energy per token vs context length - relevant to on-device tokens-per-watt benchmarking.
desislavov2023trendstheoreticalRelevant computer-vision and NLP models, comparing first implementations against consolidated versions one to two years later, on newer higher-FLOPS hardware with energy-efficiency optimizationsGrowth of inference energy consumption relative to performance gainsFor consolidated (1-2 years post-breakthrough) CV and NLP models on newer, more efficient hardware, inference energy grows much more softly than previously anticipated for sustained performance increases - contradicting exponential energy-growth expectations - with the caveat that the multiplicative factor of pervasive AI adoption could still drive large totals (no specific numbers in the abstract).Focus on consolidated implementations may understate the energy of first implementations; hardware-efficiency assumptions may not hold for LLM-scale serving; no quantitative data in abstract.moderateabstract-onlyChallenges the exponential energy-growth narrative for inference by accounting for hardware efficiency and implementation maturity.
ding2024sustainablellmsurveyn/a - LLM development/deployment lifecycle spanning training and serving phasesn/a - identifies challenges and research directions for reducing the carbon footprint of LLM servingThe paper reports that the energy consumption of LLM serving has now surpassed that of training, identifies key challenges and outlines research directions for reducing the carbon footprint of LLM serving, but the abstract gives no numeric values.Position/vision paper without new empirical measurements; the claim that serving energy surpasses training rests on cited prior work.moderateabstract-onlyPositions sustainable LLM serving (inference-phase carbon) as a research agenda - motivational framing for tokens-per-watt benchmarking.
dodge2022measuringframeworkn/a - cloud computing and machine-learning workloads; argument targeting cloud providers and data scientistsAvailability of software carbon-intensity information to usersThe paper argues that cloud providers presenting software carbon intensity information to users is a fundamental stepping stone towards minimizing AI/ML emissions, because data scientists today lack easy and reliable access to such measurements (no numeric values in the abstract).Argument/position paper without measurements; feasibility and accuracy of software carbon-intensity reporting are not evaluated.moderateabstract-onlyMeasurement-availability argument: carbon-intensity reporting by cloud providers as the prerequisite for emissions reduction - a measurement-infrastructure contribution.
du2026tokenssimulationRepresentative multi-campus LLM inference data-center system; model-quantization configurations FP16/INT8/W4A8/W4A4 (model architecture metadata from model reports, NVIDIA-spec hardware); gold/silver/bronze request tiers; MILP co-optimization with on-site turbines, BESS, PV, and grid price/carbon signalsTotal data-center operating cost under demand response; per-token energy and token throughput as dispatchable scheduling parametersIn multi-campus case studies, the quantization-enabled demand-response framework (model-instance switching, request routing, precision selection) reduces total data-center operating cost by 34.3% without curtailing served token volume, while the share of on-site generation rises from 7.1% to 12.6% and BESS charge-discharge throughput increases by 47.1%.Operates at the 15-min grid-dispatch timescale rather than millisecond serving dynamics; throughput, per-token energy, and QoS-degradation parameters are calibrated from public benchmarks and reported measurements rather than measured in-house; no absolute energy-per-token values given in the text.highfull-textBridges tokens and grid energy: a quantization-to-power mapping with per-token energy (J/token) and token throughput as dispatchable flexibility parameters for demand response.
faiz2024llmcarbonframeworkDense and MoE LLMs (e.g., GPT-3 175B, OPT family) trained on V100/A100 GPUs; validation on GPT-3 175B inference on 16 A100 GPUs (batch 32, 128-token input); benchmarked against mlco2Predicted vs actual operational carbon footprint (kg CO2eq) across training, inference, experimentation, and storage phases; hardware efficiency under parallelism configurationsLLMCarbon's operational carbon footprint projections for LLM training show disparities of <=8.2% vs actual data, while mlco2's training estimates suffer disparities of more than 69%; predicted inference latency for GPT-3 was 3.1s vs 3s actual, with inference carbon prediction error not exceeding +3.3% (assuming PUE 1.1, carbon intensity 0.429 kg CO2eq/kWh).Higher margin of error for MoE training footprints due to architectural intricacy; accuracy depends on parallelism-configuration assumptions (e.g., hardware efficiency 39%-19.7% for suboptimal settings); embodied carbon modeling remains approximatehighfull-textPre-training carbon projection model covering training/inference/storage/embodied phases for dense and MoE LLMs; relevant as an estimation methodology for carbon-per-token style accounting
ferdaus2025evaluatingbenchmarkFour AI accelerators (Nvidia A100 GPU, Intel Habana Gaudi2 HPU, Graphcore Bow-Pod64 IPU, GroqRack LPU) evaluated on MLPerf BERT-Large and ResNet50 benchmarksEnergy consumption (Wh) and energy efficiency (throughput per watt) for training and inference to reach common MLPerf-specified target accuracyFor ResNet50, Intel Gaudi2 delivered the lowest energy consumption for both training and inference and the highest inference energy efficiency, while Graphcore showed the highest training energy efficiency; for BERT-Large inference, Nvidia A100 achieved the lowest energy consumption (no absolute Wh figures in abstract).Initial study limited to 4 accelerators and 2 benchmarks; relies on vendor-provided power monitoring tools and vendor-optimized models; no absolute energy numbers in the abstractmoderateabstract-onlyHead-to-head accelerator energy-efficiency comparison using MLPerf workloads; directly relevant to energy-per-benchmark measurement methodology
floresmartin2025improvingframeworkReal data center use case with integrated sensor monitoring of key operational variables; machine learning analysis of sensor dataPower Usage Effectiveness (PUE) and data center energy consumptionNo quantitative results reported in the abstract; the step-by-step sensor-plus-ML methodology was validated on a real use case and demonstrates potential to optimize PUE and reduce energy consumption.No numbers available in the abstract; validation limited to a single use case; methodology outcomes depend on sensor coverage and data qualitylowabstract-onlyFacility-level PUE optimization via sensors and ML; provides context for facility-level efficiency beyond per-token metrics
garcia2019estimationsurveyLiterature on energy estimation approaches in computer architecture and machine learning; survey of latest software tools for energy estimation plus two ML use casesReview of methods and tools for estimating energy consumption of ML algorithmsNo quantitative findings in the abstract; the paper reviews energy-estimation approaches and software tools and presents two use cases to guide ML practitioners in measuring energy consumption.Survey/guidance paper; no experimental numbers in the abstract; tools reviewed are dated (2019)lowabstract-onlyFoundational review of energy-estimation approaches and tools for ML; grounds the measurement-methodology side of the review
geens2024energycostsimulationLlama2-7B inference on a representative hardware architecture using a PyTorch-based generalized LLM workload template and extended ZigZag design-space exploration frameworkEnergy cost (J) of prefill and decode stages; energy bottleneck attribution (memory-bound compute, weight fetching, attention)Aggressive weight-only quantization reduces Llama2-7B inference energy cost by 4.6x and shifts the bottleneck from weight fetching to the attention mechanism; memory-bound compute in the decode stage is detrimental to both latency and energy, and prefill's relative energy share grows in edge scenarios.Simulation-based results on a single representative architecture and one model; simulation speedups come at a 'negligible' but nonzero loss of accuracy; no absolute Joules numbers in the abstractmoderateabstract-onlyDesign-space exploration framework for early identification of LLM inference energy bottlenecks; supports energy-per-token modeling before hardware exists
guan2024wattscopeframeworkProduction datacenter workload; server- and rack-level aggregate power measurements already available in datacenters (no OS or application access)Normalized mean absolute error of per-application power disaggregation from aggregate server powerWattScope disaggregates application-level power from external aggregate measurements with high accuracy, often <~10% normalized mean absolute error on a production workload.Relies on datacenter workload power characteristics (low variability, low magnitude, high periodicity) being amenable to disaggregation; machine-learning-based disaggregation may not generalize to workloads violating these assumptions; abstract reports accuracy on a single production workloadmoderateabstract-onlyNon-intrusive per-application power measurement from aggregate server power; key methodology for attributing datacenter power to specific AI workloads
henderson2020systematicframeworkexperiment-impact-tracker framework applied to RL algorithms (Deep RL Energy Leaderboard), image classification and machine translation inference case studies; survey of 100 randomly sampled NeurIPS 2019 papersReal-time energy consumption (kWh) and carbon emissions (kg CO2eq) per experiment; carbon intensity of energy gridsOnly 1 of 100 sampled NeurIPS 2019 papers measured energy (45 measured runtime); the authors' own experiments contributed 8.021 kg CO2eq and 24.344 kWh of electricity; running jobs in carbon-efficient regions can cut emissions by up to 30x (Quebec vs Estonia, 2017 averages).Carbon intensities are region-averaged estimates (electricitymap.org-based); embodied/manufacturing emissions excluded; adoption depends on voluntary self-reporting by researchershighfull-textFoundational standardized energy/carbon reporting framework (experiment-impact-tracker) and RL energy leaderboard; defines the reporting conventions that tokens-per-watt benchmarks should adopt
hisaharo2024optimizingcase-studyRedesigned inference cluster architecture (advanced interconnects, high-bandwidth memory, energy-efficient power management) running a modified GPT-Neo model vs baselineThroughput, latency, and energy consumption of inferenceNo quantitative results in the abstract; the redesigned cluster and modified GPT-Neo model reportedly achieved substantial improvements in throughput, latency, and energy consumption over baseline.No numbers available in the abstract; TechRxiv preprint (not peer-reviewed); details of cluster configuration and measurement methodology not visiblelowabstract-onlyCluster- and model-level redesign for energy-efficient inference; illustrates hardware/software co-optimization levers for tokens-per-watt
husom2025sustainablebenchmark28 quantized LLMs from the Ollama library (default PTQ and weight-only quantization) deployed on a Raspberry Pi 4 (4 GB RAM), benchmarked on CommonsenseQA, BIG-Bench Hard, TruthfulQA, GSM8K, HumanEval with a high-resolution hardware-based energy measurement toolEnergy efficiency (measured power/energy), inference performance (speed), and output accuracy across quantization levels and task typesNo absolute energy numbers in the abstract; the study reveals trade-offs between energy efficiency, inference speed, and accuracy across quantization settings, identifying configurations that optimize LLM deployment on resource-constrained devices.No quantitative figures in the abstract; single edge device (Raspberry Pi 4); limited to Ollama's default quantization schemesmoderateabstract-onlyHardware-level energy profiling of quantized LLMs on edge hardware; direct example of energy-per-task measurement methodology at the edge
jacquet2026untanglingtheoreticaln/a - not described in abstract (presumably GPU workloads in hyperscale data centers leased to diverse clients)GPU energy consumption in hyperscale data centers (attribution presumed from title/context)No findings available; the abstract only frames the problem of GPU energy consumption under scrutiny in hyperscale data centers where accelerators are centralized and leased to diverse clients.Abstract is a single sentence with no methods, results, or numbers; cannot verify design, sample, or findings from available materiallowabstract-onlyPresumably addresses attributing GPU energy to tenants/clients in hyperscale data centers; relevant to per-workload power attribution if full text becomes available
jahanshahi2020gpunestframeworkMulti-GPU cloud inference systems; case studies on multi-GPU scaling, inference scheduling, and non-GPU bottlenecksEnergy efficiency (e.g., performance per watt) of multi-GPU inference systemsInference scheduling improves the energy efficiency of multi-GPU inference systems by as much as 40%.Case-study based on the systems examined; no absolute energy numbers in the abstract; findings may not generalize across scheduling policies and hardwaremoderateabstract-onlyGPU-NEST characterization methodology plus scheduling insight for multi-GPU inference; system-level tokens-per-watt optimization evidence
jahanshahi2023wattwiserframeworkMulti-GPU ML inference serving systems with per-request latency Service-Level Objectives (SLOs); load consolidation to a subset of GPUs (specific systems not stated in abstract)Power consumption minimization subject to SLO latency boundsNo results in the abstract; the work motivates consolidating inference load onto a subset of GPUs (and potentially sharing GPUs) to minimize power consumption without violating SLO.Abstract is truncated and contains no methods or results; no numbers available; relationship to the 2020 GPU-NEST work suggests shared lineage but is unverifiable from this textlowabstract-onlyGPU load consolidation for power reduction in SLO-bound inference serving; relevant to power-per-request optimization under latency constraints
jay2023softwarebenchmarkSeveral software-based power meters for CPU- and GPU-based infrastructures evaluated against high-precision physical power meters under various intensive workloadsAccuracy of software-based power measurement (power models and vendor internal interfaces) vs physical meters at node, application, and service levelsNo quantitative results in the abstract; the empirical comparison highlights the strengths and limitations of each software-based power meter and shows that choosing the right tool for a given need is difficult.No numbers in the abstract; scope limited to the meters and CPU/GPU platforms tested; software meters trade accuracy for deployability vs physical metersmoderateabstract-onlyValidation of software power meters against physical meters; critical evidence for the accuracy limits of measurement tooling used in tokens-per-watt studies
jegham2025howhungrybenchmark30 state-of-the-art LLMs in commercial datacenters; public API performance data, company-specific multipliers (PUE, WUE, CIF), statistical inference of hardware configurations; case studies on GPT-4o scaled usage and GPT-5 adaptive routingWh per prompt (short/long), water and carbon per query, annualized footprints, cross-efficiency DEA ranking of performance vs environmental costThe most energy-intensive models exceed 29 Wh per long prompt, over 65x the most efficient systems; a 0.42 Wh short GPT-4o query scaled to 700M queries/day equals annual electricity of ~35,000 US homes, and GPT-5 ranges from 0.67 Wh (short, minimal reasoning) to 33.8 Wh (long, high reasoning), with the framework's GPT-4o estimate within 19% of OpenAI's reported 0.34 Wh/query.Hardware configurations statistically inferred rather than observed; depends on company-reported PUE/WUE/CIF; excludes idle power of unutilized GPUs in partially loaded nodes; proprietary model sizes classified from API performance; per-prompt figures aggregate wide variancehighfull-textPrompt-level Wh/query benchmark across 30 models with facility-level multipliers and DEA efficiency ranking; the closest source to a 'tokens-per-watt'-style comparative benchmark
jiang2024preventingframeworkLLM-powered intelligent chatbots (e.g., ChatGPT-class systems); life-cycle and interaction analysis across eight development/deployment phases (training, fine-tuning, updating, hardware manufacturing, operations, data management, recycling)Life-cycle energy consumption and carbon emissions of LLM chatbot services (conceptual framework)No quantitative results in the abstract; the paper identifies eight life-cycle phases with energy and carbon implications and proposes a system-level solution with three strategic mitigation pathways.No numbers in the abstract; conceptual life-cycle framing without empirical measurement; mitigation pathways not quantitatively evaluated herelowabstract-onlyLife-cycle (training + inference + embodied) energy/carbon framing for LLM chatbots; broadens scope beyond inference-only energy benchmarks
kaneko2025comparingcase-studyBitcoin and Ethereum blockchains, GPT-4, Visa payment network, and web search (secondary/estimated consumption data for cloud-based services)Electricity consumption (TWh, MWh) at system-wide and per-use level; energy per transaction and per inferenceGPT-4 training required ~9,450 MWh while daily inference exceeded 500 MWh (inference often exceeding training energy), and Bitcoin consumes ~121 TWh (~0.43% of global electricity) with per-transaction energy 720,000x that of Visa, while Ethereum's move to PoS cut energy by 99.988%.Estimates for closed cloud services rely on secondary/external data (author-flagged Scope 3 measurement barrier); abstract-only access; blockchain vs GenAI figures are not directly comparable.moderateabstract-onlyProvides system-level and per-use electricity numbers showing LLM inference energy (GPT-4 daily >500 MWh) can exceed training, motivating inference energy measurement.
lange2009specpowerbenchmarkSPECpower_ssj2008: industry-standard server systems under a Java server-side workload (n/a for AI)Power and performance characteristics of computer systems (performance-per-watt)SPEC established SPECpower_ssj2008 as the first industry-standard benchmark for measuring power and performance of computer systems; the one-sentence abstract reports no quantitative results.Abstract-only (single sentence); benchmark targets general servers, not AI/LLM inference, so transferability to tokens-per-watt is indirect.moderateabstract-onlyLineage source: the first industry-standard performance-per-watt benchmark (SPECpower_ssj2008), a conceptual ancestor of tokens-per-watt.
lei2020statisticalframework17 hyperscale data centers (HDCs) operated by Google and Facebook, modeled with thermodynamics-based PUE models, representative economizer choices, climate variables and energy-system parametersPredicted PUE vs reported PUE; Sobol' total-order sensitivity indices of modeling parameters; minimum achievable PUE via differential evolutionClimate variables and uninterruptible power supply (UPS) efficiencies are the most important PUE model parameters, and predictions verified against reported PUE values of 17 HDCs capture regional and seasonal PUE variations and support point estimates for macro-level data center energy models (no specific PUE values in the abstract).Macro-level point estimations with uncertainty in energy-system parameters and economizer choices; no numeric PUE results available in the abstract; facility-level scope rather than per-inference measurement.moderateabstract-onlyStatistical PUE-prediction framework relevant to converting inference energy into facility-level efficiency metrics and to PUE target-setting for AI data centers.
li2023cloverframeworkML inference services using mixed-quality model pools with GPU resource partitioning (models and hardware not specified in the abstract)Carbon emissions of inference vs accuracy and service-level agreement (SLA) complianceClover, a carbon-friendly ML inference runtime using mixed-quality models and GPU partitioning, substantially reduces carbon emissions while maintaining high accuracy and meeting SLA targets; the abstract reports no numeric results.Abstract reports no quantitative results, and evaluation scope (models, GPU types, workloads) is not stated in the abstract.moderateabstract-onlyEarly carbon-SLA co-optimization for inference serving (quality mixing + GPU partitioning) rather than a measurement benchmark.
li2025ecoserveframeworkTwo Generative AI services at a major cloud provider (traces); Gemma 27B on NVIDIA A100 vs H100; Intel SPR CPU with llama.cpp baseline; Watttime/GreenSKU carbon-intensity traces (261 gCO2/kWh mid-level)Total carbon (operational + embodied) of LLM serving under performance targets and SLOs; decode throughputEcoServe lowers total carbon emissions by up to 47% versus performance-, energy-, and cost-optimized design points while meeting SLOs, based on findings that offline batch inference accounts for up to 55% of serving capacity and embodied carbon can exceed 50% of lifetime emissions.Findings derive from modeling plus traces of two services at one cloud provider; embodied-carbon and carbon-intensity estimates carry assumption error; CPU decode gains (up to 4.03x, avg 1.34x vs llama.cpp) are specific to Intel SPR.highfull-textShows carbon-optimal serving differs from energy-optimal; 4R (Reduce/Reuse/Rightsize/Recycle) framework with cross-stack ILP co-design.
luccioni2022bloomcase-studyBLOOM 176B parameter LM; NVIDIA A100 SXM4 80GB (TDP 400 W); Jean Zay cluster; GCP us-central1 API deployment tracked over ~18 daystCO2eq, kWh, and gCO2eq/kWh across the full training lifecycle and real-time API inference deploymentBLOOM's final training emitted ~24.7 tCO2eq (dynamic power only) or 50.5 tCO2eq full lifecycle (433,195 kWh at ~57 gCO2eq/kWh, plus 256,646 kWh idle), and API inference emitted ~19 kg CO2eq/day (340 kg total) with ~75% of deployment energy spent just keeping the model in memory.Authors could not track real-time power (TDP-based estimates), excluded CPU power (~40x less than GPUs), and grid carbon intensity varies by time and location; deployment tracked on a single GCP instance.highfull-textCanonical lifecycle carbon accounting of a 176B LLM, quantifying memory-bound, idle-dominated inference energy and advocating granular energy/carbon-intensity/PUE reporting.
luccioni2024powerhungrybenchmark88 models (80 task-specific finetuned + 8 multi-purpose zero-shot: Flan-T5 base/large/xl/xxl, BLOOMz-560M/1B/3B/7B) across 10 tasks and 30 datasets (text/image classification, QA, MLM, token classification, text generation, summarization, captioning, object detection, image generation) run sequentially (no batching) 10x on 8x NVIDIA A100-SXM4-80GB GPUs (AWS us-west-2, 297.6 g CO2eq/kWh), energy/carbon measured with CodeCarbonEnergy (kWh) and carbon (g CO2eq) per 1,000 inferences; training-vs-inference cost parity (number of inferences)Per-task mean energy per 1,000 inferences ranges from 0.002 kWh (text classification) to 2.9 kWh (image generation, median 1.35 kWh) - a >1,450x spread - with multi-purpose generative models orders of magnitude more energy-intensive than task-specific ones (e.g., 0.3-0.7 g CO2eq per 1,000 inferences for task-specific QA/sentiment vs 2.34-10 g for zero-shot models), total study consumption 754.66 kWh / 178.97 kg CO2eq, and BLOOMz training-inference cost parity reached at ~205M-593M inferences per model.Single hardware platform and region (A100, us-west-2); sequential non-batched inference (reflects in-situ deployment but not batching gains); idle power of other GPUs included in measurements; open-source models only; authors note study is not representative of all deployment contexts and that proprietary-model transparency is lacking.highfull-textFoundational inference-phase energy benchmark establishing kWh (and g CO2eq) per 1,000 inferences as a comparison unit and quantifying the energy penalty of multi-purpose generative models - a key precursor to tokens-per-watt benchmarking.
maliakel2026characterizingbenchmarkLlama-1B/3B/8B and Qwen-14B/32B (five decoder-only LLMs) on a single NVIDIA RTX PRO 6000 (Blackwell) GPU; BoolQ, HellaSwag, TruthfulQA, NarrativeQA; NVML power sampling at 10 msEnergy per query (J), end-to-end latency, output quality, and energy-performance tradeoffs under GPU SM DVFS (7 frequency levels, 180-2842 MHz)Reducing SM frequency from 2842 to 180 MHz achieves ~42% average energy savings with only 1-6% latency increase (decode dominates 77-91% of time and is frequency-insensitive), and combining workload-aware model routing (32B->3B) with DVFS cuts per-query energy from 20.97 J to 2.52 J (88% saving) in the upper-bound use case.Single-GPU offline setup (no multi-GPU or live serving); routing-based savings trade quality (83.8% -> 77.0% in the use case) and assume accurate difficulty prediction (input length is a weak predictor, r=0.002).highfull-textPhase-aware DVFS plus workload-aware routing evidence in joules-per-query terms, with a quantified energy-quality frontier.
niu2025energybenchmarkvLLM, TensorRT-LLM, and DeepSpeed inference engines on one GPU node with 2x H100 GPUsPower (W) and energy decomposed by inference stage (setup: initialization + model loading vs token generation) and by component (GPU, CPU, DRAM)Provides a fine-grained power benchmark of LLM inference engines on a 2x H100 node, decomposing the inference lifecycle into setup vs token-generation stages and GPU/CPU/DRAM components to identify energy bottlenecks; the abstract reports no numeric results.Abstract-only; single node with 2 H100 GPUs, no multi-node or production serving deployment, and engine/version details not in the abstract.moderateabstract-onlyStage- and component-level power breakdown of mainstream inference engines - evidence for where inference energy actually goes.
niu2026tokenpowerbenchbenchmarkLlama, Falcon, Qwen, and Mistral model series from 1B up to Llama3-405B; declarative config over model/prompt set/inference engine; GPU-, node-, and system-level measurement without specialized power metersJoules per token and other energy-efficiency metrics, with phase-aligned attribution of energy to prefill and decode stages per requestTokenPowerBench, the first lightweight benchmark for LLM-inference power studies, captures GPU/node/system-level power without specialized meters and attributes energy to prefill vs decode per request (joules per token) across models from 1B to Llama3-405B; the abstract reports no numeric results.Abstract-only; measurement fidelity without specialized meters is an inherent design tradeoff, and coverage is limited to four model families and their engines.moderateabstract-onlyDirectly delivers a tokens-per-watt-style benchmark (joules per token, phase-aligned prefill/decode) motivated by inference being >90% of total LLM power per industry reports.
patterson2021carboncase-studyT5, Meena, GShard, Switch Transformer, GPT-3, and Evolved Transformer (NAS); TPU v2/v3, P100, V100; Google datacenters at multiple locationsEnergy use (kWh/MWh) and CO2e (tCO2e) of training and inference, plus average system power (W) per processorSparsely activated DNNs consume <1/10th the energy of dense DNNs at equal accuracy, cloud datacenters are ~1.4-2x and ML accelerators ~2-5x more efficient, and combined DNN/datacenter/processor choices cut carbon footprint up to ~100-1000x, with training CO2e of the studied models ranging 4-552 tCO2e.Estimates rely on measured average system power (TPU v2 221 W, TPU v3 283 W, P100 271 W, V100 325 W) and grid carbon-intensity assumptions; retrospective estimates can be off by 18.7x (average org) to 88x (efficient org) as shown for the NAS case.highfull-textFoundational argument that energy use and CO2e should be first-class ML evaluation metrics; authors collaborated with MLPerf to include energy during training and inference.
poddar2025towardsbenchmarkLLM inference across a wide range of NLP tasks; multiple models, tasks, prompts, and system-related factorsInference energy across tasks, models, and system configurationsFirst broad benchmark of LLM inference energy across NLP tasks: inference energy correlates strongly with output token length and response time, and quantization, optimal batch sizes, and targeted prompt phrasing significantly reduce energy use.Abstract-only; exact magnitudes not reported in the abstract.moderateabstract-onlyAcademic benchmark of LLM inference energy across tasks, models, and prompts.
reddi2020mlperfinferencebenchmarkMLPerf Inference (v0.5): ResNet-50 v1.5, MobileNet-v1, GNMT (NMT) among other workloads; 30+ systems spanning embedded to datacenter; 600+ measurements from 14 organizationsInference latency, throughput (latency-bounded), and accuracy across four scenarios (server, single-stream, multistream, offline); system power/performance rangeML inference systems span three orders of magnitude in power consumption and five in performance, and latency constraints cut server-scenario throughput by 39-55% for NMT versus 3-35% (avg ~20%) for ResNet-50 v1.5 and under 10% for MobileNet-v1.The first-round benchmark did not include energy or power as a submission metric (latency/accuracy only), and results are self-reported by submitters subject to audit.highfull-textFoundational MLPerf Inference methodology (scenarios, LoadGen, quality targets, 99% tail-latency confidence bounds) that later gained a Power/energy division.
reddi2020mlpermobilebenchmarkMobile devices with diverse SoCs (e.g., MediaTek Dimensity 1100, Snapdragon) and software stacks (NNAPI, TFLite, vendor SDKs, OpenVINO); CV and NLP tasksLatency, throughput, and accuracy of on-device ML inference (no energy metric in this version)Across the first two benchmark rounds within six months, offline throughput improved 3x and latency reduced by up to 12x on mobile devices, while an optimized framework delegate (Neuron vs generic NNAPI) delivered over 10% performance difference.Benchmark measures latency/throughput/accuracy only - no energy or power metric in this version; results reflect early rounds (v0.7/v1.0 era) and specific devices.highfull-textIndustry-standard mobile ML benchmark methodology (LoadGen-based) showing rapid stack-level improvements; its lack of an energy metric motivates power-aware mobile benchmarking.
rrapaj2024powercase-studyCori and Perlmutter supercomputers at NERSC; six months of production HPC workload power measurements across the CPU-to-GPU transitionPower draw (W) vs peak provisioned power and TDP over time and across applications/usersPower usage varied considerably but stayed consistently well below peak provisioned power on both machines, and after the GPU transition production power demands did not grow as fast as peak capabilities (further lowering the fraction of TDP used), suggesting machines could be power-capped well below TDP; no numeric values in the abstract.Abstract-only; findings cover HPC workloads at one site (NERSC), not AI/LLM inference, and no numeric values are available in the abstract.moderateabstract-onlyFacility-level evidence that production power sits far below TDP/peak provisioning - supports power-capped, over-provisioned designs relevant to inference datacenters.
rubei2025promptbenchmarkLlama 3 on the CodeXGLUE code-generation benchmark, evaluated in an isolated testing environmentEnergy consumption and accuracy of generated code during LLM inference (carbon/energy impact of prompt engineering techniques)Initial results show that using specific tags to distinguish prompt parts can reduce Llama 3's energy consumption during inference without compromising code-generation performance, though no quantitative energy figures are reported in the abstract.Authors flag the results as initial and requiring more in-depth evaluation; single model and task; no numbers disclosed in the abstract.moderateabstract-onlyShows prompt engineering itself as a lever on inference energy (relevant to energy-per-token optimization), but lacks quantitative measurements.
samsi2023wordsbenchmarkLLaMA 7B/13B/65B on NVIDIA V100 and A100 GPUs with model sharding across up to 32 GPUs, batch sizes 64-512, on Alpaca and GSM8K datasets (4,096 sampled inputs per dataset)Inference energy costs: Joules per output token, energy per response, energy per second (Watts), token rate, and effect of GPU power cappingLLaMA 65B inference consumes roughly 3-4 Joules per output token at ~300 W to ~1 kW across 8-32 shards, and power capping A100s from 250 W to 175 W reduces total energy by ~23.21% at only +6.7% average inference time (a 150 W cap costs +19.49% time).Energy figures are estimates from power draw assumptions on only two GPU generations and one model family (LLaMA); limited datasets; authors note broader power-capping recommendations need additional experimentation.highfull-textLandmark 'from words to watts' study establishing per-output-token Joule figures and power-capping trade-offs for LLM inference benchmarking.
sanchez2025greenmlopsbenchmarkDiscriminative models (various architectures and hyperparameters) and generative LLMs of different sizes across multiple hardware setups in real-world MLOps pipelinesEnergy consumption during training and inference via software-based power measurements; correlations with model size, reasoning complexity, and request-handling capacityFor discriminative models, optimizing architecture, hyperparameters, and hardware significantly reduces energy without sacrificing performance, and for LLMs larger models do not necessarily consume more energy when utilization is low; no numeric figures are given in the abstract.Software-based power measurement rather than watt-meters; abstract provides no quantitative results or model/hardware enumeration.moderateabstract-onlyPositions itself as a benchmark for estimating total AI energy use across model types in MLOps, emphasizing utilization-dependent LLM efficiency.
schwartz2020greenaitheoreticaln/a (position paper; cites aggregate trend: deep learning computation doubling every few months, ~300,000x increase from 2012 to 2018)Proposes efficiency as an evaluation criterion alongside accuracy and reporting of financial cost ('price tag') for developing, training and running modelsArgues deep learning's computations have an estimated 300,000x increase from 2012 to 2018 with a surprisingly large carbon footprint, and advocates making efficiency a standard evaluation criterion and reporting cost baselines to enable greener, more inclusive AI research (no new empirical measurements).Position paper without new measurements; efficiency metrics proposed qualitatively rather than operationalized as a benchmark; no numeric per-model efficiency data.highabstract-onlyProvides the normative argument ('Green AI') for treating energy efficiency as a first-class evaluation criterion - the conceptual basis for tokens-per-watt benchmarks.
sejourne2026sovereignbenchmarkHeterogeneous sovereign (SecNumCloud-compliant) infrastructure: NVIDIA GPUs (L40S, A100), legacy hardware, and Apple Silicon architectures; interactive inference workloads; carbon model integrating operational energy and embodied hardware carbon; custom simulator for hardware-renewal vs legacy-extension arbitrageEnergy per request across hardware; total carbon cost (operational + embodied); scale-to-zero capabilityApple Silicon architectures consume 4x-7x less energy per request for interactive workloads than traditional NVIDIA-based servers while high-end GPUs offer superior efficiency at saturation, and a custom simulator optimizes total carbon cost by arbitrating between hardware renewal and extending legacy equipment lifespans (measured values in Fig. 2 and Table II, not in abstract).Constrained industrial setting (SecNumCloud compliance) and specific interactive workloads; quantitative energy-per-request values referenced to figures/tables not available in the abstract.moderateabstract-onlyExtends per-request energy benchmarking to include embodied carbon and hardware-heterogeneity ('digital sobriety') decisions in sovereign AI infrastructure.
stojkovic2024towardsbenchmarkLlama-2 70B served with vLLM on an NVIDIA DGX-H100 (H100 GPUs, frequencies 800-1980 MHz, tensor parallelism degrees 2/4/8, batch sizes up to 64) under latency SLOs of 5x solo TTFT/TBTGPU power draw, total energy, latency (TTFT/TBT), and throughput under frequency capping, batching, and model-parallelism knobsGPU frequency capping achieves ~20% lower power for most workload configurations with no latency or throughput impact, running at 1.6 GHz instead of 2.0 GHz yields about the same throughput at under 80% of the energy, and reducing maximum batch size during low-throughput phases cuts energy by up to 15%.Single model (Llama-2 70B) and single-node DGX-H100; GPUs offer only GPU-wide frequency control (no fine-grained power knobs); results are workload- and SLO-dependent.highfull-textCharacterizes frequency, batching, and parallelism levers for energy-efficient LLM serving under performance SLOs - core input for tokens-per-watt operating-point selection.
stojkovic2025dynamollmframeworkLLM inference clusters of 8xH100 DGX servers (12 servers baseline; 11 in 24h run) running Llama2-70B (primary) plus Llama2-13B, Llama3-70B, Mixtral-8x7B, Mixtral-8x22B, Falcon-180B on vLLM; workloads from Azure production traces (Coding, Conversation; 1h open-source, 1-day and 1-week traces); requests bucketed into 9 input/output length classes (SS..LL) with TTFT/TBT SLOs at 5x isolated latencyEnergy in Watt-hours (Wh) per request/configuration and per cluster under latency SLOs; operational CO2 emissions; customer cost; GPU power (W)At service level DynamoLLM conserves 53% energy, 38% operational carbon emissions (5.0 vs 3.1 t CO2/week on CAISO), and 61% customer cost (40 -> 24.6 GPU servers, $1,362.7/h saved) while meeting latency SLOs; measured per-request energy for Llama2-70B ranges ~0.77-13.21 Wh depending on request length, tensor parallelism and GPU frequency, and cluster energy is reduced 35% (1h trace), 42% (1-day trace), 23-51% across load levels, and 47-56% in week-long simulations.Evaluated on open-source models and Azure traces only; week-long results come from a discrete-time simulator; considers tensor parallelism only (not pipeline parallelism); reports Wh per request and cluster energy rather than a per-token (J/token) metric; SLO assumptions (5x isolated latency) may not generalize.highfull-textFirst energy-management framework for LLM inference clusters, quantifying how GPU frequency scaling, tensor parallelism and instance scaling trade off against energy-per-request under SLOs - direct evidence for cluster-level tokens-per-watt optimization (energy savings of 23-56%).
strubell2019energybenchmarkA variety of recently successful neural network models for NLP (the paper's measurements covered transformer-based architectures of the era, e.g., Transformer base/large, ELMo, BERT base/large, GPT-2, and NAS) trained on GPU hardwareApproximate financial cost (hardware/cloud), energy (kWh), and carbon emissions (CO2) of model trainingQuantifies the approximate financial and environmental costs of training recently successful NLP models, showing accuracy gains depend on substantial energy consumption, and proposes actionable recommendations to reduce costs (no numeric values in the abstract; the paper's well-known figures are training-cost estimates in kWh/lbs CO2).Estimates tied to specific hardware, electricity prices and grid carbon intensity of the period; training-phase focus only (inference not measured); abstract provides no numbers.moderateabstract-onlySeminal training-phase energy/carbon quantification for NLP models that motivated subsequent inference-phase tokens-per-watt benchmarking.
tripp2024measuringbenchmarkBUTTER-E dataset: 63,527 runs / 30,582 configurations of fully connected networks (13 datasets, 20 sizes, 8 shapes, 14 depths) on NREL Eagle HPC CPU nodes (Xeon Gold 6154) and GPU nodes (2x V100), measured with node-level watt-meters (HPE iLO)Real-world energy consumption (Joules per training datum/epoch, power time series) and accuracy of a proposed hardware-informed energy modelGPU training consumed a median ~9.47 mJ per training datum vs ~6.16 mJ on CPU (GPUs less energy-efficient in this setting), energy per datum rises non-linearly with parameter count until ~2^20 params, and the proposed energy model predicts GPU energy within +/-2.74% and CPU energy within +/-20.2%.Fully connected networks only (authors explicitly call for extending to LLMs, CNNs, GNNs); training rather than inference focus; single HPC datacenter; 163 runs filtered as system artifacts.highfull-textProvides a rigorous watt-meter-based measurement methodology and public dataset (BUTTER-E) relevant to energy benchmarking practice, and challenges FLOP/parameter-count assumptions about efficiency.
tschand2025mlperfbenchmark60 systems spanning edge devices to cloud datacenters running representative workloads from the MLPerf benchmark suite; 1,841 reproducible measurementsEnergy efficiency of ML systems across power levels from microwatts to megawatts, via the MLPerf Power methodology (rules and best practices)1,841 reproducible measurements from 60 systems reveal trade-offs between performance, complexity, and energy efficiency across ML deployment scales; no per-measurement energy figures are given in the abstract.Abstract-only; as a methodology/standards paper, comparability across heterogeneous platforms is itself a stated challenge the rules must address.moderateabstract-onlyMLPerf Power defines the industry-standard benchmarking methodology for ML energy efficiency from edge to datacenter - the direct benchmark lineage for tokens-per-watt.
verma2020demystifyingbenchmarkMLPerf benchmark suite workloads compared against DAWNBench and DeepBench on multi-GPU training systemsCompute rate, memory transactions per second, scaling efficiency, host CPU utilization, and mixed-precision training effectsMLPerf benchmarks exhibit moderately high memory transactions per second and compute rates (vs DAWNBench's high-compute/low-memory and DeepBench's low-compute profiles), with scaling-efficiency variation across models and quantified gains from mixed-precision Tensor Core training; no energy numbers are reported in the abstract.Pre-LLM-era training workloads; characterization targets performance and bottlenecks, not energy efficiency.moderateabstract-onlyBackground on what MLPerf workloads measure (compute/memory behavior), showing why energy efficiency needs a separate measurement dimension (later supplied by MLPerf Power).
wang2025storellmframeworkLLM inference with permanently pre-stored attention matrices, evaluated against state-of-the-art LazyLLM, plus StoreLLM-MoE and StoreLLM-PTQ variantsEnergy consumption of LLM inference (primary outcome) and inference delayStoreLLM outperforms LazyLLM by 1.45x in energy consumption with a sacrifice of only 5.05% in delays, and the StoreLLM-MoE and StoreLLM-PTQ variants achieve 2.64x and 2.83x energy reductions respectively versus state-of-the-art LLM systems.Abstract-only; relies on the observation that attention matrices stay largely unchanged across inferences; storage costs of pre-stored attention matrices are not quantified in the abstract.moderateabstract-onlyRadical energy-vs-compute substitution idea (replace attention-matrix computation with storage access) claiming large inference energy reductions.
wilhelm2025beyondbenchmarkLlama 3.2 1B and 8B (plus 7B-class LLMs for MT-Bench single-token tests) on a single NVIDIA L40S GPU, batch size 1, no parallelism; MMLU (57 categories clustered into 8 domains) and MT-Bench; NVML-based power measurementEnergy-per-Token (Joule) = W_consumed x time / tokens processed, alongside accuracy on MMLU and MT-BenchRouting to Llama 8B instead of CoT on Llama 1B offers similar or better accuracy at 40-60% lower energy, while CoT boosts Math accuracy by 281% at a ~15,132% energy increase (category baselines of 76-84 MJ), motivating their proposed Energy-per-Token metric and operating-curve-based energy-aware routing.Single GPU (L40S), batch size 1, no parallelism, small models (1B/8B); GPU-level energy only, excluding system overheads.highfull-textDirectly advocates Energy-per-Token as a standard efficiency metric complementing accuracy benchmarks - central to the tokens-per-watt review framing.
wilkins2024hybridsimulationRepresentative LLM dataset of queries with heterogeneous hardware accelerators (energy-efficient processors vs high-performance GPUs) in a hybrid datacenter modelCPU+GPU energy consumption of LLM workloads under a cost-based, workload-aware scheduling frameworkThe hybrid strategy, which allocates tasks to energy-efficient processors or high-performance GPUs based on input/output token counts per query, reduces CPU+GPU energy consumption by 7.5% compared to a workload-unaware baseline; no absolute energy figures are given in the abstract.Abstract-only; representative (unspecified) LLM dataset; datacenter-model-level analysis rather than hardware measurements.moderateabstract-onlyToken-count-aware scheduling across heterogeneous hardware as a datacenter-level energy-efficiency lever for LLM serving.
wilkins2024offlinesimulationSeveral state-of-the-art LLMs on heterogeneous GPU-CPU systems, characterized across different magnitudes of input prompts and output textWorkload-dependent energy consumption and runtime of LLM inference; fit quality (R^2) of energy/runtime models and energy-optimality of offline schedulingEnergy and runtime models fit each LLM with R^2 > 0.96 across prompt/output magnitudes, and a case study of the offline energy-optimal scheduling framework demonstrates advantages of energy- and accuracy-aware scheduling over existing best practices; no absolute numbers are given in the abstract.Abstract-only; specific models and systems not enumerated; offline scheduling only, no online/adaptive evaluation.moderateabstract-onlyShows that workload-dependent energy models of LLM inference can be accurate enough (R^2 > 0.96) to drive energy-optimal scheduling - methodology relevant to energy-per-token modeling.
wu2022sustainablecase-studyMeta/Facebook production ML workloads (large language model LM, ranking models RM1-RM5), large-scale models including GPT-3 (750B params) and Switch Transformer (1.5T params), datacenter fleet with PUE ~1.10, and LCA-based hardware embodied-carbon analysisEnd-to-end carbon footprint (operational + embodied/manufacturing) and operational power consumption of AI computing across the model development cycle and system hardware life cycleManufacturing (embodied) carbon cost is roughly 50% of the location-based operational carbon footprint of large-scale ML tasks at Meta, aggregate cross-stack optimizations reduced operational power by on average 20% every six months, the LM model's carbon footprint is dominated by inference while training-vs-inference footprints are roughly equal for ranking models, and raising GPU utilization to 80% decreases overall carbon footprint by 3x.Single-company (Meta) experience; location-based carbon intensities assumed and PUE assumed at 1.1 rather than measured per workload; results may not generalize to other providers or workloadshighfull-textIndustry-scale characterization of training vs inference energy split and a call to add efficiency/environmental measures to MLPerf-style leaderboards, directly motivating tokens-per-watt-style benchmarking.
xu2021surveygreensurveyn/a - survey of green deep learning literature: compact networks (e.g., MobileNet family, efficient attention/softmax variants), energy-efficient training (initialization, normalization, progressive training, HPO), energy-efficient inference (pruning, low-rank factorization, quantization incl. 8-bit BERT, I-BERT, TernaryBERT, distillation), and efficient data usagen/a - organizes methods into 4 categories; discusses candidate efficiency measures: running time, carbon emission (CO2eq), model size, FLOPs, 'fair measure', and intuitive understandingNo new empirical numbers are reported; the survey classifies green deep learning techniques into compact networks, energy-efficient training, energy-efficient inference, and efficient data usage, and recommends reporting FLOPs plus running time for fair comparison while noting FLOPs are theoretical values that diverge from actual runtime.Descriptive survey without measurements; authors flag that FLOPs are theoretical and parallelism/degree of utilization create a gap between FLOPs and running time; efficiency claims of surveyed methods not independently verifiedhighfull-textProvides the measurement-metrics discussion (FLOPs vs runtime vs carbon emission vs fair measure) that informs how a tokens-per-watt benchmark should be defined and reported.
yang2026pelmbenchmarkMobile/edge hardware platforms and datasets (not enumerated in abstract) running on-device LLM inference; compared against state-of-the-art DVFS-based power governing methodsEnergy consumption and inference speedup under thermal/power constraints, with task performance maintainedPELM, which augments DVFS frequency tuning with speculative decoding and variable verification depth, achieves up to 23.1% speedup and 52.4% reduction in energy consumption compared to state-of-the-art power governing methods while maintaining comparable task performance (numbers from abstract).Abstract-only; hardware platforms, datasets, and LLMs not named in the abstract; results specific to thermally constrained mobile scenarios and dependent on speculative decoding qualitymoderateabstract-onlyAdds workload-level power knobs (speculative decoding, verification depth) beyond DVFS for energy-efficient on-device LLM inference - relevant to edge tokens-per-watt optimization.
yilk2018qualifyingcase-studyFour LANL supercomputing platforms: Trinity (separate Haswell and Knights Landing CPU partitions), Grizzly, Fire, Ice; enhanced power monitoring infrastructure; Green500 benchmarkGreen500 qualification level (performance per watt) and experience meeting Green500 reporting requirementsAll four new LANL platforms were qualified at the highest level of the Green500 benchmark using the enhanced power-monitoring infrastructure (no numeric efficiency values in the abstract).Experience report focused on the qualification process; no quantitative performance-per-watt or power numbers available in the abstract.moderateabstract-onlyDocuments institutional practice of qualifying HPC systems on the Green500 performance-per-watt benchmark - a methodological precedent for tokens-per-watt-style efficiency qualification.
yu2023knowsimulationCloud-scale ML inference cluster simulation with prototype implementation: 105 servers with three different kinds of GPUs serving five ML models, evaluated with real-world tracesCloud-scale energy consumption of ML inference serving; energy efficiency of GPU architectures across active-GPU counts and clock frequenciesA hierarchical GPU resource management approach (energy-aware cluster allocation, intra-cluster node scaling, intra-node GPU scaling, and GPU clock scaling) saves up to 28.3% of cloud-scale energy consumption when serving five ML models on 105 servers with three GPU types (from abstract).Abstract-only; evaluation combines prototype with trace-driven cloud-scale simulation rather than full production deployment; SLO-blind DVFS finding is specific to commercial GPU drivers; generalizability to other workloads/fleets not shownmoderateabstract-onlyShows GPU energy efficiency varies with architecture, active-GPU count, and clock frequency at constant throughput - cluster-level evidence for facility/cloud-scale tokens-per-watt management.
zhao2023sustainablecase-studyGPUs at a research supercomputing center (HPC/datacenter production environment) subjected to power-capping for AI/ML workloadsGPU temperature and power draw under power-capping, plus impact on job performance and hardware lifespanPower-capping GPUs at a research supercomputing center significantly decreased both GPU temperature and power draw, reducing power consumption and potentially improving hardware lifespan with minimal impact on job performance; the abstract reports no specific magnitudes.Abstract-only; single-center field study; magnitude of temperature/power reductions and the precise job-performance impact are not quantified in the abstractmoderateabstract-onlyFacility-scale evidence that power-capping AI accelerators cuts power draw and temperature with minimal performance cost - relevant to datacenter-level energy accounting for inference fleets.

Swipe sideways to see all columns.

References

  1. Tschand, Arya et al. (2025). MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from μWatts to MWatts for Sustainable AI — 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). Abstract only. MLPerf Power defines the industry-standard benchmarking methodology for ML energy efficiency from edge to datacenter - the direct benchmark lineage for tokens-per-watt.doi:10.1109/hpca61900.2025.00092
  2. Verma, Snehil et al. (2020). Demystifying the MLPerf Training Benchmark Suite — IEEE ISPASS 2020. Abstract only. Background on what MLPerf workloads measure (compute/memory behavior), showing why energy efficiency needs a separate measurement dimension (later supplied by MLPerf Power).doi:10.1109/ispass48437.2020.00013
  3. Reddi, Vijay Janapa et al. (2019). MLPerf Inference Benchmark — arXiv preprint. Full text read. Foundational MLPerf Inference methodology (scenarios, LoadGen, quality targets, 99% tail-latency confidence bounds) that later gained a Power/energy division.doi:10.48550/arxiv.1911.02549
  4. Banbury, Colby et al. (2021). MLPerf Tiny Benchmark — arXiv (Cornell University). Full text read. Foundational energy-inclusive ML benchmark: makes energy a first-class metric alongside latency and accuracy for ultra-low-power inference - a key reference for energy-in-benchmark methodology.doi:10.48550/arxiv.2106.07597
  5. Reddi, Vijay Janapa et al. (2020). MLPerf Mobile Inference Benchmark — arXiv (Cornell University). Full text read. Industry-standard mobile ML benchmark methodology (LoadGen-based) showing rapid stack-level improvements; its lack of an energy metric motivates power-aware mobile benchmarking.doi:10.48550/arxiv.2012.02328
  6. Lange, Klaus-Dieter (2009). Identifying Shades of Green: The SPECpower Benchmarks — Computer. Abstract only. Lineage source: the first industry-standard performance-per-watt benchmark (SPECpower_ssj2008), a conceptual ancestor of tokens-per-watt.doi:10.1109/mc.2009.84
  7. Jahanshahi, Ali et al. (2020). GPU-NEST: Characterizing Energy Efficiency of Multi-GPU Inference Servers — IEEE Computer Architecture Letters. Abstract only. GPU-NEST characterization methodology plus scheduling insight for multi-GPU inference; system-level tokens-per-watt optimization evidencedoi:10.1109/lca.2020.3023723
  8. Niu, Chenxu et al. (2026). TokenPowerBench: Benchmarking the Power Consumption of LLM Inference — Proceedings of the AAAI Conference on Artificial Intelligence. Abstract only. Directly delivers a tokens-per-watt-style benchmark (joules per token, phase-aligned prefill/decode) motivated by inference being >90% of total LLM power per industry reports.doi:10.1609/aaai.v40i38.40535
  9. Niu, Chenxu (2025). Energy Efficient or Exhaustive? Benchmarking Power Consumption of LLM Inference Engines — ACM SIGEnergy Energy Informatics Review. Abstract only. Stage- and component-level power breakdown of mainstream inference engines - evidence for where inference energy actually goes.doi:10.1145/3757892.3757900
  10. Beitsayadeh, Carl & Darbyshire, Pamayla E. (2026). Benchmarking AI Inference Efficiency in Public and Private Clouds: An MLPerf-Based Comparative Study — IEEE Transactions on Cloud Computing. Abstract only. Demonstrates a reproducible performance-per-watt benchmarking framework built entirely on open MLPerf data - directly relevant to tokens-per-watt standardization.doi:10.1109/tcc.2026.3674888
  11. Ferdaus, Farah et al. (2025). Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf Benchmarks — 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Abstract only. Head-to-head accelerator energy-efficiency comparison using MLPerf workloads; directly relevant to energy-per-benchmark measurement methodologydoi:10.1109/ccgrid64434.2025.00035
  12. Maliakel, Paul Joe et al. (2026). Characterizing LLM Inference Energy-Performance Tradeoffs Across Workloads and GPU Scaling — 2026 IEEE 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Full text read. Phase-aware DVFS plus workload-aware routing evidence in joules-per-query terms, with a quantified energy-quality frontier.doi:10.1109/ccgrid68966.2026.00013
  13. Csikai, Dávid et al. (2026). Scaling Behavior of Energy per Token in LLM Inference on Apple Silicon — 2026 IEEE 8th International Conference and Workshop Óbuda on Electrical and Power Engineering (CANDO-EPE). Abstract only. Contributes an Apple Silicon-appropriate measurement workflow and empirical scaling laws for energy per token vs context length - relevant to on-device tokens-per-watt benchmarking.doi:10.1109/cando-epe71091.2026.11569440
  14. Jay, Mathilde (2023). An experimental comparison of software-based power meters: focus on CPU and GPU — 2023 IEEE/ACM 23rd International Symposium on. Abstract only. Validation of software power meters against physical meters; critical evidence for the accuracy limits of measurement tooling used in tokens-per-watt studiesdoi:10.1109/ccgrid57682.2023.00020
  15. Jacquet, Pierre et al. (2026). Untangling GPU Power Consumption: Job-Level Inference in Cloud Shared Settings — Proceedings of the 21st European Conference on Computer Systems. Abstract only. Presumably addresses attributing GPU energy to tenants/clients in hyperscale data centers; relevant to per-workload power attribution if full text becomes availabledoi:10.1145/3767295.3769333
  16. Guan, Xiaoding et al. (2024). WattScope: Non-intrusive Application-level Power Disaggregation in Datacenters — ACM SIGMETRICS Performance Evaluation Review. Abstract only. Non-intrusive per-application power measurement from aggregate server power; key methodology for attributing datacenter power to specific AI workloadsdoi:10.1145/3649477.3649491
  17. Tripp, Charles et al. (2024). Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations — arXiv (Cornell University). Full text read. Provides a rigorous watt-meter-based measurement methodology and public dataset (BUTTER-E) relevant to energy benchmarking practice, and challenges FLOP/parameter-count assumptions about efficiency.doi:10.48550/arxiv.2403.08151
  18. Samsi, Siddharth (2023). From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference — arXiv (Cornell University). Full text read. Landmark 'from words to watts' study establishing per-output-token Joule figures and power-capping trade-offs for LLM inference benchmarking.doi:10.48550/arxiv.2310.03003
  19. Argerich, Mauricio Fadel & Patiño-Martı́nez, Marta (2024). Measuring and Improving the Energy Efficiency of Large Language Models Inference — IEEE Access. Abstract only. Inference-focused LLM energy profiling methodology with quantization and batching as efficiency levers - complements training-centric carbon studies.doi:10.1109/access.2024.3409745
  20. Hisaharo, Soka (2024). Optimizing LLM Inference Clusters for Enhanced Performance and Energy Efficiency — arXiv preprint. Abstract only. Cluster- and model-level redesign for energy-efficient inference; illustrates hardware/software co-optimization levers for tokens-per-wattdoi:10.36227/techrxiv.172348951.12175366/v1
  21. Yu, Junyeol (2023). Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving — 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Abstract only. Shows GPU energy efficiency varies with architecture, active-GPU count, and clock frequency at constant throughput - cluster-level evidence for facility/cloud-scale tokens-per-watt management.doi:10.1109/hpca56546.2023.10070943
  22. Wilkins, Grant (2024). Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems — ACM SIGEnergy Energy Informatics Review. Abstract only. Shows that workload-dependent energy models of LLM inference can be accurate enough (R^2 > 0.96) to drive energy-optimal scheduling - methodology relevant to energy-per-token modeling.doi:10.1145/3727200.3727217
  23. Geens, Robin et al. (2024). Energy Cost Modelling for Optimizing Large Language Model Inference on Hardware Accelerators — 2024 IEEE 37th International System-on-Chip Conference (SOCC). Abstract only. Design-space exploration framework for early identification of LLM inference energy bottlenecks; supports energy-per-token modeling before hardware existsdoi:10.1109/socc62300.2024.10737844
  24. Yang, Weisi & Xia, Stephen (2026). PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling — Proceedings of the 2026 ACM/IEEE International Conference on Embedded Artificial Intelligence and Sensing Systems. Abstract only. Adds workload-level power knobs (speculative decoding, verification depth) beyond DVFS for energy-efficient on-device LLM inference - relevant to edge tokens-per-watt optimization.doi:10.1145/3774906.3802783
  25. Jahanshahi, Ali et al. (2023). WattWiser: Power &amp; Resource-Efficient Scheduling for Multi-Model Multi-GPU Inference Servers — Proceedings of the 14th International Green and Sustainable Computing Conference. Abstract only. GPU load consolidation for power reduction in SLO-bound inference serving; relevant to power-per-request optimization under latency constraintsdoi:10.1145/3634769.3634807
  26. Wilhelm, Patrick (2025). Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference — Proceedings of the 5th Workshop on Machine Le. Full text read. Directly advocates Energy-per-Token as a standard efficiency metric complementing accuracy benchmarks - central to the tokens-per-watt review framing.doi:10.1145/3721146.3721953
  27. Desislavov, Radosvet (2023). Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning — Sustainable Computing Informatics and Systems. Abstract only. Challenges the exponential energy-growth narrative for inference by accounting for hardware efficiency and implementation maturity.doi:10.1016/j.suscom.2023.100857
  28. Stojkovic, Jovan (2024). Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference — arXiv (Cornell University). Full text read. Characterizes frequency, batching, and parallelism levers for energy-efficient LLM serving under performance SLOs - core input for tokens-per-watt operating-point selection.doi:10.48550/arxiv.2403.20306
  29. Stojkovic, Jovan (2025). DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency — 2025 IEEE International Symposium on High Per. Full text read. First energy-management framework for LLM inference clusters, quantifying how GPU frequency scaling, tensor parallelism and instance scaling trade off against energy-per-request under SLOs - direct evidence for cluster-level tokens-per-watt optimization (energy savings of 23-56%).doi:10.1109/hpca61900.2025.00102
  30. Poddar, Soham et al. (2025). Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models — Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics. Abstract only. Peer-reviewed NAACL 2025 study benchmarking LLM inference energy - exactly on topic, but full content must be fetched before use in the review.doi:10.18653/v1/2025.naacl-long.632
  31. Jegham, Nidhal (2025). How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference — arXiv (Cornell University). Full text read. Prompt-level Wh/query benchmark across 30 models with facility-level multipliers and DEA efficiency ranking; the closest source to a 'tokens-per-watt'-style comparative benchmarkdoi:10.48550/arxiv.2505.09598
  32. Husom, Erik Johannes et al. (2025). Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency — ACM Transactions on Internet of Things. Abstract only. Hardware-level energy profiling of quantized LLMs on edge hardware; direct example of energy-per-task measurement methodology at the edgedoi:10.1145/3767742
  33. Ali, Sabiya Banu Masthan et al. (2026). Assessing the Sustainability of LLM Inference through Energy–Accuracy Analysis — Proceedings of the 17th ACM International Conference on Future and Sustainable Energy Systems. Abstract only. Empirical energy-accuracy trade-off mapping for open LLM inference at GPU level - a method template for energy-aware model selection.doi:10.1145/3744255.3811741
  34. Chien, Andrew A. (2023). Reducing the Carbon Impact of Generative AI Inference (today and in 2035) — Proceedings of the 2nd Workshop on Sustainable Computer Systems. Abstract only. Early carbon-per-query request-routing comparison (Local vs Balance vs CarbonMin) for generative AI inference - carbon-aware direction angle.doi:10.1145/3604930.3605705
  35. Rrapaj, Ermal (2024). Power Consumption Trends in Supercomputers: A Study of NERSC's Cori and Perlmutter Machines — ISC High Performance 2024 Research Paper Proceedings (39th International Conference). Abstract only. Facility-level evidence that production power sits far below TDP/peak provisioning - supports power-capped, over-provisioned designs relevant to inference datacenters.doi:10.23919/isc.2024.10528943
  36. Zhao, Dan (2023). Sustainable Supercomputing for AI — ACM Symposium on Cloud Computing. Abstract only. Facility-scale evidence that power-capping AI accelerators cuts power draw and temperature with minimal performance cost - relevant to datacenter-level energy accounting for inference fleets.doi:10.1145/3620678.3624793
  37. Billah, Waseq et al. (2026). Is Power Usage Effectiveness (PUE) Still Fit for Purpose? An Integrative Review of Data Center Metrics for Assessing the Environmental Impact of Modern AI Infrastructure — Journal of Science Policy &amp; Governance. Abstract only. Facility-level metric critique (PUE) relevant to the facility-level efficiency measurement debate for AI infrastructure.doi:10.38126/jspg280101
  38. Lei, Nuoa & Masanet, Eric (2020). Statistical analysis for predicting location-specific data center PUE and its improvement potential — Energy. Abstract only. Statistical PUE-prediction framework relevant to converting inference energy into facility-level efficiency metrics and to PUE target-setting for AI data centers.doi:10.1016/j.energy.2020.117556
  39. Flores-Martin, Daniel et al. (2025). Improving Energy Efficiency in a Data Center: PUE Analyzing and Tuning — 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Abstract only. Facility-level PUE optimization via sensors and ML; provides context for facility-level efficiency beyond per-token metricsdoi:10.1109/ccgrid64434.2025.00028
  40. Aneli, Stefano et al. (2025). Modelling and experimental surveys on the energy consumption of a small-scale data center — Energy Efficiency. Abstract only. Facility-level angle: data-center cooling demand response modeled in energy terms, linking facility efficiency to grid flexibility rather than tokens/watt.doi:10.1007/s12053-025-10357-7
  41. Luccioni, Sasha et al. (2024). Power Hungry Processing: Watts Driving the Cost of AI Deployment? — The 2024 ACM Conference on Fairness Accountability and Transparency. Full text read. Foundational inference-phase energy benchmark establishing kWh (and g CO2eq) per 1,000 inferences as a comparison unit and quantifying the energy penalty of multi-purpose generative models - a key precursor to tokens-per-watt benchmarking.doi:10.1145/3630106.3658542
  42. Luccioni, Alexandra Sasha et al. (2022). Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model — arXiv (Cornell University). Full text read. Canonical lifecycle carbon accounting of a 176B LLM, quantifying memory-bound, idle-dominated inference energy and advocating granular energy/carbon-intensity/PUE reporting.doi:10.48550/arxiv.2211.02001
  43. Patterson, David A. et al. (2021). Carbon Emissions and Large Neural Network Training — arXiv (Cornell University). Full text read. Foundational argument that energy use and CO2e should be first-class ML evaluation metrics; authors collaborated with MLPerf to include energy during training and inference.doi:10.48550/arxiv.2104.10350
  44. Strubell, Emma et al. (2019). Energy and Policy Considerations for Deep Learning in NLP — Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Abstract only. Seminal training-phase energy/carbon quantification for NLP models that motivated subsequent inference-phase tokens-per-watt benchmarking.doi:10.18653/v1/p19-1355
  45. Wu, Carole-Jean et al. (2021). Sustainable AI: Environmental Implications, Challenges and Opportunities — arXiv (Cornell University). Full text read. Industry-scale characterization of training vs inference energy split and a call to add efficiency/environmental measures to MLPerf-style leaderboards, directly motivating tokens-per-watt-style benchmarking.doi:10.48550/arxiv.2111.00364
  46. Faiz, Ahmad et al. (2023). LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models — arXiv (Cornell University). Full text read. Pre-training carbon projection model covering training/inference/storage/embodied phases for dense and MoE LLMs; relevant as an estimation methodology for carbon-per-token style accountingdoi:10.48550/arxiv.2309.14393
  47. Dodge, Jesse et al. (2022). Measuring the Carbon Intensity of AI in Cloud Instances — 2022 ACM Conference on Fairness, Accountability, and Transparency. Abstract only. Measurement-availability argument: carbon-intensity reporting by cloud providers as the prerequisite for emissions reduction - a measurement-infrastructure contribution.doi:10.1145/3531146.3533234
  48. Henderson, Peter et al. (2020). Towards the Systematic Reporting of the Energy and Carbon Footprints of\n Machine Learning — arXiv (Cornell University). Full text read. Foundational standardized energy/carbon reporting framework (experiment-impact-tracker) and RL energy leaderboard; defines the reporting conventions that tokens-per-watt benchmarks should adoptdoi:10.48550/arxiv.2002.05651
  49. Xu, Jingjing et al. (2021). A Survey on Green Deep Learning — arXiv preprint. Full text read. Provides the measurement-metrics discussion (FLOPs vs runtime vs carbon emission vs fair measure) that informs how a tokens-per-watt benchmark should be defined and reported.doi:10.48550/arxiv.2111.05193
  50. Jiang, Peng et al. (2024). Preventing the Immense Increase in the Life-Cycle Energy and Carbon Footprints of LLM-Powered Intelligent Chatbots — Engineering. Abstract only. Life-cycle (training + inference + embodied) energy/carbon framing for LLM chatbots; broadens scope beyond inference-only energy benchmarksdoi:10.1016/j.eng.2024.04.002
  51. Kaneko, Yusuke (2025). Comparing Electricity Consumption Per Use of Blockchain and Generative AI — IEEE Access. Abstract only. Provides system-level and per-use electricity numbers showing LLM inference energy (GPT-4 daily >500 MWh) can exceed training, motivating inference energy measurement.doi:10.1109/access.2025.3573722
  52. García-Martín, Eva et al. (2019). Estimation of energy consumption in machine learning — Journal of Parallel and Distributed Computing. Abstract only. Foundational review of energy-estimation approaches and tools for ML; grounds the measurement-methodology side of the reviewdoi:10.1016/j.jpdc.2019.07.007
  53. Sánchez-Mompó, Adrián et al. (2025). Green MLOps to Green GenOps: An Empirical Study of Energy Consumption in Discriminative and Generative AI Operations — Information. Abstract only. Positions itself as a benchmark for estimating total AI energy use across model types in MLOps, emphasizing utilization-dependent LLM efficiency.doi:10.3390/info16040281
  54. Aquino-Brítez, Sergio et al. (2025). Towards an Energy Consumption Index for Deep Learning Models: A Comparative Analysis of Architectures, GPUs, and Measurement Tools — Sensors. Abstract only. Proposes a standardized, sensor-validated energy-efficiency index covering both training and inference across architectures - relevant as a measurement-methodology template.doi:10.3390/s25030846
  55. Rubei, Riccardo et al. (2025). Prompt engineering and its implications on the energy consumption of Large Language Models — 2025 IEEE/ACM 9th International Workshop on Green And Sustainable Software (GREENS). Abstract only. Shows prompt engineering itself as a lever on inference energy (relevant to energy-per-token optimization), but lacks quantitative measurements.doi:10.1109/greens66463.2025.00014
  56. Li, Yueying et al. (2025). EcoServe: Designing Carbon-Aware AI Inference Systems — arXiv preprint. Full text read. Shows carbon-optimal serving differs from energy-optimal; 4R (Reduce/Reuse/Rightsize/Recycle) framework with cross-stack ILP co-design.doi:10.48550/arxiv.2502.05043
  57. Li, Baolin et al. (2023). Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference Service — Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Abstract only. Early carbon-SLA co-optimization for inference serving (quality mixing + GPU partitioning) rather than a measurement benchmark.doi:10.1145/3581784.3607034
  58. Ding, Yi & Shi, Tianyao (2024). Sustainable LLM Serving: Environmental Implications, Challenges, and Opportunities : Invited Paper — 2024 IEEE 15th International Green and Sustainable Computing Conference (IGSC). Abstract only. Positions sustainable LLM serving (inference-phase carbon) as a research agenda - motivational framing for tokens-per-watt benchmarking.doi:10.1109/igsc64514.2024.00016
  59. Wilkins, Grant (2024). Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads — The 15th ACM International Conference on Futu. Abstract only. Token-count-aware scheduling across heterogeneous hardware as a datacenter-level energy-efficiency lever for LLM serving.doi:10.1145/3632775.3662830
  60. Du, Bojun et al. (2026). From Tokens to Energy Flexibility: Quantization-Enabled Demand Response for Data Centers with LLM Inference Workloads — arXiv preprint. Full text read. Bridges tokens and grid energy: a quantization-to-power mapping with per-token energy (J/token) and token throughput as dispatchable flexibility parameters for demand response.doi:10.48550/arxiv.2606.18851
  61. Séjourné, Kevin et al. (2026). Sovereign LLM Inference in the Era of Digital Sobriety: A Comparative Study of Energy Efficiency Across Heterogeneous Architectures — 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET). Abstract only. Extends per-request energy benchmarking to include embodied carbon and hardware-heterogeneity ('digital sobriety') decisions in sovereign AI infrastructure.doi:10.1109/icecet65726.2026.11633138
  62. Wang, Dan et al. (2025). StoreLLM: Energy Efficient Large Language Model Inference with Permanently Pre-stored Attention Matrices — Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems. Abstract only. Radical energy-vs-compute substitution idea (replace attention-matrix computation with storage access) claiming large inference energy reductions.doi:10.1145/3679240.3734604
  63. Yilk, Todd (2018). Qualifying for the Green500: Experience with the newest generation of supercomputers at LANL — Sustainable Computing: Informatics and Systems. Abstract only. Documents institutional practice of qualifying HPC systems on the Green500 performance-per-watt benchmark - a methodological precedent for tokens-per-watt-style efficiency qualification.doi:10.1016/j.suscom.2018.02.004
  64. Chen, Huamin et al. (2026). The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency — arXiv preprint. Full text read. Core theoretical contribution for the review: context window as the dominant tok/W lever (1/W law) plus quantitative decomposition of routing-topology vs hardware-generation energy gains.doi:10.48550/arxiv.2603.17280
  65. Schwartz, Roy et al. (2020). Green AI — Communications of the ACM. Abstract only. Provides the normative argument ('Green AI') for treating energy efficiency as a first-class evaluation criterion - the conceptual basis for tokens-per-watt benchmarks.doi:10.1145/3381831