On this page

Energy efficiency across the AI datacenter stack

What does the 2023-2026 literature establish about improving energy efficiency across the AI datacenter stack - from GPU power management and workload scheduling to grid-level coordination - and how strong is the evidence for each lever?

Updated
7 Aug 2026
Sources
93
Years
2023–2026
Confidence
Download Markdown

GPU power cappingtokens per wattcarbon-aware schedulingdata center demand responseLLM inference energydigital twinsMaxQpower smoothing

How this review was made
Databases
OpenAlex, Crossref, arXiv API, arXiv web search, Semantic Scholar (best-effort)
Queries (literal)
GPU power capping
power capping data center
GPU energy efficiency power management
tokens per watt GPU
LLM inference energy efficiency
energy efficient large language model serving
power-aware job scheduling data center
energy-aware scheduling HPC
carbon-aware scheduling data center
carbon-aware computing
data center demand response
grid-aware data center
data center power grid interaction
power smoothing data center
data center power flexibility
data center digital twin
digital twin data center thermal management
GPU cluster power management
data center telemetry anomaly detection machine learning
liquid cooling data center energy efficiency
data center UPS battery grid services
data center power shaving
inference energy optimization
GPU DVFS energy efficiency
energy consumption machine learning training
power management GPU inference
power-aware deep learning training
carbon intensity workload shifting
renewable energy data center scheduling
data center flexibility ancillary services
frequency scaling LLM inference
energy proportional computing
GPU cluster energy optimization
inference serving energy efficiency
quantization energy efficiency inference
embodied carbon computing
green AI sustainable machine learning
data center thermal management reinforcement learning
energy efficiency neural network inference
power consumption prediction GPU
ML training energy optimization
data center liquid cooling machine learning
Zeus Efficiently Optimizing Energy Consumption of ML Training
Perseus Reducing Energy Bloat in Large Model Training
WattScope Non-Intrusive Power Measurement Data Center
Energon Trading Computation for Energy
FAAST GPU Energy-Saving Framework
MLPerf Power Benchmarking Energy Efficiency
SchedInspector Energy-Efficient GPU Scheduling
Towards Carbon-Aware Data Centers
Data Center Power Oversubscription
Green Datacenter Energy Efficiency Survey
Zeus optimizing energy consumption machine learning training (arXiv)
Energon trading computation for energy (arXiv)
FAAST GPU energy saving framework (arXiv)
power oversubscription data center GPU (arXiv)
carbon-aware GPU scheduling cluster (arXiv)
data center digital twin machine learning operations (arXiv)
energy proportional computing data center (arXiv)
demand response data center workload shifting (arXiv)
GPU power capping inference energy (arXiv)
energy efficient transformer training optimizer (arXiv)
power smoothing data center (arXiv)
LLM inference energy efficiency power capping (arXiv)
Energon trading computation for energy ICML (arXiv)
FAAST energy saving GPU server (arXiv)
tokens per watt inference (arXiv)
checkpoint energy large model training (arXiv)
prefill decode energy split inference (arXiv)
data center ancillary services frequency regulation (arXiv)
GPU energy prediction machine learning model (arXiv)
carbon aware inference serving (arXiv)
Perseus reducing energy bloat (arXiv)
SchedInspector GPU scheduling energy (arXiv)
Search last run
2026-08-07
Screening
93 sources used · 2023–2026 · deep review

Summary

The short version

AI datacenters are among the fastest-growing electricity consumers of the 2020s, and the 2023–2026 literature describes a stack of levers for improving their efficiency, from the GPU die up to the power grid. This review synthesises 93 retrieved sources — 64 peer-reviewed papers (HPCA, SIGCOMM, SOSP, SIGMETRICS, SoCC, IPDPS, e-Energy, IEEE TPDS/ToC/TSG, Applied Energy, and others) and 29 preprints. Three findings carry the review. First, the device-level lever — power capping and frequency control — is real but sharply regime-dependent: it recovers ~20–30% energy on compute-bound or older-GPU workloads, yet on memory-bound LLM decode the power caps often never even trigger, and clock locking outperforms capping. Second, the workload-level lever — how LLM inference clusters are configured and served — is where the largest measured savings live (a flagship 2025 study reports 53% energy and 38% carbon reduction under latency SLOs), but most evidence is simulation-based. Third, the grid-level lever — datacenters acting as demand-response assets, “power smoothing”, and grid-interactive operation — is the most novel and the thinnest: one production deployment (a 130 kW Blackwell cluster responding to UK grid events) carries much of the evidence, and most of the rest is preprints and power-engineering journals. Confidence is moderate: 18 sources were read in full text, 75 at abstract level, and several important Model-to-Grid claims exist only as SSRN preprints that could not be retrieved from open archives.

Why this question

The AI infrastructure industry is converging on a claim that the right objective for an AI datacenter is not maximum power draw but maximum tokens per watt — and, further, that a datacenter should be an asset to the electricity grid rather than a liability to it. Vendors now market “Model-to-Grid” orchestration that coordinates GPU clusters with grid conditions in real time. Whether any of this is engineering or marketing depends on what the peer-reviewed record actually shows about each claimed lever, and how strongly.

What turns on the answer is practical and expensive: a 100 MW AI factory deciding between power-capped GPUs, different serving configurations, carbon-aware schedulers, and demand-response contracts is choosing between levers whose evidence quality differs by orders of magnitude. Getting the evidence wrong means buying equipment or signing contracts on the basis of a simulation from one lab. This review asks, for each layer of the stack, what the 2023–2026 literature establishes — and where it stops.

Scope and methods

Question. What does the 2023–2026 literature establish about improving energy efficiency across the AI datacenter stack — GPU power management (capping, DVFS, “MaxQ”), power- and carbon-aware workload scheduling, LLM inference serving configuration, cooling and digital-twin operations, and grid-level coordination (demand response, power smoothing) — and how strong is the evidence for each lever?

Inclusion criteria. Works published or posted 2023-01-01 to 2026-08-07; peer-reviewed systems venues (USENIX NSDI/OSDI/ATC/FAST/HotCarbon, ACM SIGCOMM/SOSP/ASPLOS/EuroSys/SoCC/SIGMETRICS/IMC/e-Energy/CoNEXT/HPDC/ICS/PPoPP, IEEE SC/IPDPS/INFOCOM/HPCA/ISCA/CLUSTER/CCGrid/TPDS/ToC/TSG, ACM TOCS/ToN, Elsevier Energy/Applied Energy/FGCS, and related) plus arXiv and open preprints, which is where much of this fast-moving field lives. Exclusion criteria. Chip-level microarchitecture without datacenter-system implications; power-grid engineering without a datacenter-computing component (unit commitment, power-flow solvers, converters, wind/PV smoothing without DC load); residential/smart-home demand response; wireless “full duplex”; works unreachable through any open archive (see below). Depth: deep (80+ sources targeted).

Databases and dates. OpenAlex, Crossref, the arXiv API, and arXiv web search; Semantic Scholar attempted but effectively unusable from this IP (persistent 429s). Seventy-four query strings (listed in frontmatter) plus citation-graph follow-ups on anchor papers (Zeus, Perseus, WattScope, MLPerf Power, DynamoLLM, Splitwise, Carbon Explorer, CASPER, Power Hungry Processing). Searches run 2026-08-07.

Screening counts. 2,957 raw records retrieved across all queries; ~1,958 unique after DOI/title deduplication; 246 passed a strict venue + topic screen; 113 were curated into the working set; 21 were dropped at retrieval because no open copy exists — 14 paywalled IEEE/Elsevier items with no OA version, and 7 SSRN preprints behind Cloudflare (SSRN blocks both scripted and headless-browser access from this network). 93 sources included, all retrieved this session: 18 read in full text (arXiv HTML / ar5iv), 75 at abstract level (paywalled or abstract-only availability; flagged abstract-only in the evidence table and hedged accordingly). Every DOI was verified against Crossref (title match, no retractions via update-to) or, for arXiv DOIs, by resolution check at doi.org. Note that 10.48550 DOIs are DataCite-minted and return 404 at Crossref, which is expected.

What this review deliberately does not cover. Pre-2023 foundational work (Zeus’s 2022 arXiv version, early demand-response literature) is cited only through its published versions where retrieved; hardware microarchitecture; grid engineering per se; and the economics/policy literature on datacenter electricity markets.

The landscape

The literature has three distinct neighbourhoods that barely cite each other. The systems neighbourhood (NSDI, SIGCOMM, SOSP, SoCC, HPCA, e-Energy) treats energy as a schedulable resource: papers measure GPU power, cap it, schedule around it, and price carbon. The power-engineering neighbourhood (IEEE Transactions on Smart Grid/Power Systems, Applied Energy, Energy) treats the datacenter as a flexible load: papers model demand response, UPS/battery participation, and grid stability with stylised IT workloads. The preprint neighbourhood (arXiv, SSRN) is where “Model-to-Grid” claims live — grid-interactive GPU clusters, curtailment-window training, quantization-enabled demand response — often with no peer-reviewed version at all. The shape of the field explains the review’s shape: the systems literature is rigorous but mostly inside the datacenter; the grid literature is rigorous but mostly outside it; the stack-integration claim from the article that motivated this review (treating model architecture, scheduling, and grid constraints as one optimisation problem) is currently supported mainly by preprints and vendor material, not peer-reviewed systems work.

By era: 2023 work concentrates on GPU power capping and carbon-aware scheduling (largely in response to the 2022–2023 generative-AI wave); 2024 adds LLM-serving energy optimisation and the first production-scale power datasets; 2025–2026 shifts toward grid interaction, with 2026 preprints on grid-responsive inference and curtailment-aware pretraining. One lab or group does not dominate, but a few institutions recur (Meta’s carbon-aware cluster scheduling program; the LLM-serving energy cluster around DynamoLLM/GreenLLM; USENIX-adjacent power-measurement work). No single benchmark dominates the corpus — which is itself a finding: MLPerf Power is the only cross-vendor benchmark, and it measures systems, not tokens.

Theme one — the device lever: power capping, DVFS, and the MaxQ question

The oldest and best-measured lever is controlling the GPU itself — capping power draw (nvidia-smi -pl-style), scaling clocks, or both. The 2023 measurement study From Words to Watts remains the anchor: on LLaMA 65B across V100/A100 shards, inference draws roughly 300 W–1 kW and ~3–4 J per decoded token; capping from 250 W to 175 W costs only ~6.7% more time while cutting energy ~23% 69. The same trade-off shape appears in HPC workloads: capping A100s at ~150–250 W and running GEMMs at 55–70% of TDP recovers ~30% efficiency at small throughput cost 3, and unbalanced capping across nodes in a power-constrained HPC system improves overall energy efficiency 62. Kernel-level DVFS for LLM inference reports ~15% energy saving at <1% slowdown 73; DNN-based frequency selection reports ~27% energy at ~2% performance loss 4; adaptive capping reports ~13% energy at ~3% loss 20; and iteration-level frequency control for LLM inference reports co-optimised energy–latency operating points 24. The same levers appear at other scales: fine-grain DVFS prediction for GPUs 9, DVFS effects for real-time GPU workloads 66, a power-capping evaluation for AI-on-5G edge platforms (FROST) 56, power capping of GPU servers specifically for ML inference optimisation 53. A 2026 study extends this to the datacenter: prediction-based dynamic power capping for latency-sensitive cloud applications 34, and a learning-enabled adaptive capping scheme for cloud datacenters 78. At the server level, energy-aware GPU sharing across heterogeneous GPUs (HeShare) reports energy gains from packing tasks onto efficiency-optimal devices 41. The common thread: capping works when the workload is compute-bound — power falls faster than throughput.

The 2026 result that complicates the whole theme is The Illusion of Power Capping in LLM Decode: on an H200 (700 W), autoregressive decode across four attention architectures draws only 137–300 W, so no power cap in the 280–700 W range ever triggers — decode is memory-bound and simply never approaches its power limit. Static clock locking recovers up to 32% of decode energy at <1% throughput loss and Pareto-dominates capping at every matched operating point 54. DVFS effects also differ between the prefill and decode phases of LLM inference, with phase-aware settings beating uniform ones 84. The same memory-bound logic explains why Splitwise found the token-generation phase tolerates >50% power capping with almost no latency impact while the prefill phase is cap-sensitive 61. Batch size is the regime switch: the “illusion” paper shows that as batch size grows, decode moves toward the compute-bound regime where capping starts to bind 54. Finally, a caution against assuming the adjacent lever of quantization always saves energy: weight-only quantization increases energy ~25–45% on small models (1.1–1.5B) while saving ~23% on 6–9B models, depending on backend 91.

Bottom line: the MaxQ lever is real, measurable, and worth 15–30% on the right workloads — but its applicability flips with hardware generation and batch size, and on modern memory-bound decode the effective control is clock frequency, not the power cap.

Theme two — the workload lever: tokens per watt in LLM serving

The strongest quantitative results in this corpus come from optimising how inference clusters are configured and operated, not the chips. DynamoLLM, which dynamically reconfigures cluster size, tensor parallelism, and GPU frequency under latency SLOs, reports 53% energy and 38% operational-carbon reduction with 61% lower customer cost on production-scale cloud traces 76. Its companion position paper argues energy-efficiency must be a first-class objective in LLM inference design 75. Phase-splitting (running prefill and decode on different machines) yields 1.4× throughput at 20% lower cost, or 2.35× throughput at equal cost and power 61; prefill–decode splitting across edge/cloud also reports large energy reductions for edge inference (ELLIE, ~1.8×) 27. Predictive GPU throttling under SLOs reports up to ~44% lower energy 42; energy–latency co-optimisation with dynamic batching (ELTO) 22; and workload-based energy models that pick energy-optimal serving configurations offline 88, with heterogeneous clusters (mixing GPU generations) lowering LLM inference energy 87.

Measurement studies set the scale of the prize: per-1,000-inference energy ranges from 0.002 kWh (text classification) to 2.9 kWh (image generation, up to 11.5 kWh for the least efficient model), with multi-purpose generative models orders of magnitude more expensive than task-specific ones 51. A 2025 benchmark of inference engines finds large, engine-dependent spreads (the authors’ headline: “energy efficient or exhaustive?”) 60, and a multi-model footprint study reports ~29 Wh per prompt with a 65× spread across models and hardware 40. Simulation-based studies of cluster design for energy-optimal serving 33 and GPU resource management under latency-and-power constraints 46 fill out the design space, while a position paper argues “energy-per-token” should be a first-class metric 86. Notably for billing and carbon accounting, attributing energy to individual requests in batched serving is itself an open measurement problem: naive per-token attribution deviates from exact Shapley values, and a calibrated method reduces error to ~0.12–0.18 normalized L1 at negligible overhead 52.

Bottom line: tokens-per-watt is decided mostly at the serving layer — batch composition, phase placement, frequency, and cluster configuration — and the measured headroom (tens of percent) exceeds what device-level capping alone delivers. The caveat: the flagship numbers are simulation/trace-driven; only the measurement studies are laboratory-hard.

Theme three — the scheduler lever: power- and carbon-aware scheduling

Scheduling work splits into power-constrained scheduling (HPC) and carbon-aware scheduling (cloud/HPC), with different evidence quality.

Power-constrained HPC. Scheduling with lightweight power predictions in power-constrained HPC platforms 14, energy-hardware-workload-aware scheduling for interconnected HPC environments 18, and the uncertainty-aware multi-objective HARMONIC 2 all report energy savings in simulation on real workload traces. Production-adjacent work includes an LLM-based power prediction for energy-aware HPC scheduling (SC ’25 workshop) 57, which its follow-up uses to quantify the economic potential of energy-aware scheduling — a ~8% electricity-bill reduction on one UK HPC operator’s tariffs 58 — plus a SLURM plugin for automatic energy-efficient scheduling 74 and an account of energy-aware operation of German HPC systems 77. A SoCC position paper argues for sustainable supercomputing for AI as a first-class systems problem 92, and a scheduling study for cloud-native ML workloads (EELAS) reports energy savings under latency constraints 79. The methodological caution comes from job-level power disaggregation work showing that GPU power in shared cloud settings is hard to attribute per job — a precondition for any of these schedulers 36, and from Linux-level process scheduling experiments showing OS energy accounting is a work in progress 65.

Carbon-aware computing. The influential ISCA-published Carbon Explorer frames datacenter design around carbon, not just energy 1; CASPER schedules distributed web services to carbon-optimal times and locations 72. The strongest systems-venue result is SIGCOMM 2025’s carbon- and precedence-aware scheduling for data-processing clusters, reporting ~33% carbon reduction in simulation 44. Around these: carbon-aware scheduling of AI datacenter workloads using grid forecasts 10, a public-data decision-support framework for carbon-aware workload allocation 49, a reinforcement-guided carbon-aware GPU orchestration framework reporting ~85% CO₂e reduction with no SLO violations in simulation 63, geographic load-shifting scenario modelling 82, and an estimate of generative-AI inference’s carbon trajectory to 2035 15. HPC-specific carbon accounting 45 and a proposal to locate datacenters at renewable curtailment hotspots (“follow the curtailment”) 19 round out the design space.

Two results discipline this theme. First, the sunk carbon fallacy (SoCC 2024): including embodied carbon — a sunk cost — in operational scheduling decisions can produce decisions that do not reduce total carbon at all; operational decisions should be driven by marginal carbon 7. Second, a scenario study finds realistic geographic load-shifting saves only ~5% carbon in current grids — far below simulation claims — because renewable surplus is rarely both large and simultaneously available across regions 82. Carbon-intensity attribution across providers is itself contested 55.

Bottom line: scheduling levers are plausible and simulation-supported, but the gap between headline simulation savings (~30–85%) and conservative scenario estimates (~5%) is the widest in this review, and production deployments are essentially absent.

Theme four — the grid lever: demand response, power smoothing, grid-interactive AI factories

The article that motivated this review claims datacenters can be grid assets. The peer-reviewed record is thin but the preprint record is suddenly rich. The standout is a production report of a 130 kW Blackwell GPU cluster operated as a grid-interactive asset: 100% compliance across 200+ UK National Grid events, ~30% power reduction within 40 seconds (reserve-comparable response), sustained 10–40% curtailment for 2–10 hours, and geographic load shifting (a 375 W GPU power cap in Ashburn shifting load to Chicago) at ~30 ms added time-to-first-token 89. On the demand-response side, a learning-based data-center model enables efficient DR bidding 16; surveys of datacenter flexibility potential 17 and of datacentres as grid flexibility sources 80 argue the theoretical headroom is large; UPS flexibility is explicitly modelled as an economic resource 93; and quantification work frames datacenter flexibility as a resource-adequacy asset 26, with seasonal-hourly analysis showing 80–160 MW of spatial flexibility in one regional case 35.

The preprints push furthest toward “Model-to-Grid”: a framework making LLM quantization level a dispatchable demand-response lever reports 34% cost reduction without curtailing tokens 25; distributed LLM pretraining scheduled into renewable-curtailment windows preserves training quality while cutting operational emissions to 5–12% of single-site baselines, with 97% of energy consumed in curtailment windows 85; grid-interactive thermal management of AI datacenters uses distributionally robust optimisation to shift cooling load 70; strategic load shifting is analysed for market-efficiency and transmission value 12; and “night-window batching” for clinical AI workloads closes ~78% of the carbon gap of full carbon-aware scheduling at a fraction of the complexity 23. A privacy-preserving framework for coordinating hierarchical datacenters with power networks 47 and, at the facility edge, off-grid solar/battery/hydrogen powering of hyperscale AI datacenters 11 round out the space.

Bottom line: grid-level claims have the weakest peer-reviewed base of any theme — one production deployment, several surveys, and a cluster of 2025–2026 preprints — but the production deployment is unusually concrete (200+ grid events, sub-minute response), and the preprint direction (tokens as dispatchable load, curtailment-window training) is exactly the “Model-to-Grid” claim of the motivating article, now with numbers attached.

Theme five — the facility lever: cooling, thermal control, digital twins, telemetry

Cooling is typically 20–40% of datacenter energy, and the 2023–2026 record has two strands: control and simulation.

Cooling control. RL-based cooling management with safety constraints reports ~13% cooling-power reduction (SafeCool) 83; a thermal-aware LLM-inference scheduler for cooling-regulated datacenters reports ~41% throughput gain at elevated supply temperatures (TAWS) 50; a mixture-of-experts multi-critic RL for sustainable datacenter management 48 and safety-constrained multi-agent RL for hybrid air-liquid cooling (reporting better COP and lower pump power) 64 continue the pattern; and an operational measurement study of a direct liquid-cooled facility quantifies the real lever — raising supply temperature from ~22 °C to 25 °C cut cooling power ~63% in that facility 29. At the urban-infrastructure edge, underground thermal-energy storage for datacenter cooling (e-Energy 2026) proposes seasonal shifting of heat rejection 30.

Digital twins and telemetry. The digital-twin strand is almost entirely simulation: a digital-twin-driven energy-management framework for heat-pipe cooling 71, a scalable digital-twin framework reporting PUE improvement from 1.85 to 1.70 in simulation 31, physics-informed ML thermal surrogates for datacenter “Physical AI” operations 13, RL benchmarking for liquid-cooling optimisation 59, and ML-based anomaly detection for cooling equipment 67. The measurement infrastructure strand is stronger: MLPerf Power (HPCA 2025) is the field’s first systematic cross-vendor energy benchmark 81, with independent evaluations of AI accelerators 28 and of public/private clouds 8; WattScope (SIGMETRICS 2024) achieves non-intrusive application-level power disaggregation in production datacenters 32; the PM100 dataset provides job-level power traces from a production HPC system 6; Kepler estimates container energy 5; an experimental comparison of software power meters quantifies their error 38; NERSC’s Cori/Perlmutter power trends document real HPC power behaviour 68; a systematic classification of GPU workload power characteristics (Minos) cuts profiling time ~89% 37; node-level energy models for AI inference servers map workload- and hardware-dependent efficiency 43; a 2023 HPCA study characterises the energy-performance behaviour of machine-learning services in cloud settings 90; an early survey documents trends in AI-inference energy consumption beyond parameter-count scaling laws 21; and a reliability-modeling study quantifies availability risks of liquid-cooling systems, the other side of the cooling-efficiency coin 39.

Bottom line: cooling control has real measured headroom (tens of percent, with the supply-temperature result verified operationally), while digital-twin claims are uniformly simulation-grade — a gap worth naming, since vendors describe digital twins as “real-time engines” (the motivating article’s own claim) rather than offline planning tools.

Where the evidence disagrees

1. Does power capping help LLM inference? The 2023–2024 record says yes: ~23% energy for ~7% time (LLaMA 65B) 69, >50% cap tolerance on decode 61. The 2026 record says the cap is an illusion on modern hardware: H200 decode draws 137–300 W, caps never trigger, and clock locking beats capping 54. The apparent conflict dissolves along a hardware-generation × batch-size axis: capping binds when the workload is compute-bound (older GPUs, small batches, prefill); decode on H-class hardware with large batches is memory-bound and never approaches the cap. The two sides study different operating points, not different physics — and the “illusion” paper says so explicitly. This is a resolved disagreement, and it is the single most important nuance for anyone buying “MaxQ” infrastructure.

2. How much does carbon-aware scheduling actually save? Simulation papers report 33% 44 to 85% 63 carbon reductions; a scenario model of real grids reports ~5% for geographic shifting 82, and the sunk-carbon argument says naive metrics can push decisions that don’t reduce carbon at all 7. The disagreement is metric- and assumption-driven: simulations assume accurate marginal-carbon forecasts and grid regions with genuinely different intensities; the scenario work and the fallacy critique attack exactly those assumptions. Both can be right — the field has not yet produced a deployment that settles which regime dominates.

3. Is quantization an energy lever? The working assumption in serving papers is that lower precision saves energy; a 2026 empirical study finds the sign of the effect flips with model size and backend (energy up 25–45% on small models) 91. Unresolved; the practical reading is that quantization energy effects must be measured per deployment.

4. Digital twins: real-time control or offline simulation? Vendors and the motivating article claim real-time digital-twin control; every retrieved digital-twin paper evaluates offline or in simulation 31 71 13. No peer-reviewed deployment of a closed-loop real-time digital twin for datacenter energy was located. This is an absence, not a contradiction — but an expensive one to bet on.

Gaps and open questions

  • No standard tokens-per-watt benchmark. MLPerf Power measures system-level workload energy 81; nothing standardised measures energy per served token across serving stacks, which is precisely the metric the industry says it optimises. The attribution problem is only beginning to be solved 52.
  • The grid-level lever lacks systems-venue evidence. Only one production grid-interactive GPU-cluster deployment was retrieved 89, and it is a preprint. Demand-response work lives in power-engineering venues with stylised workloads 17, while systems venues have not yet published the orchestration layer. What would settle it: a peer-reviewed deployment study with SLO and grid-event data.
  • Unretrievable but load-bearing preprints. Several key Model-to-Grid claims exist only as SSRN preprints (quantization-enabled demand response, algorithmic virtual inertia for frequency response, power-profile reshaping of GPU clusters) that could not be retrieved from open archives — SSRN is Cloudflare-walled to automated access. This is a reproducibility problem for the field: the most on-topic claims are the least accessible.
  • The GB200/GB300 era is unstudied. The highest-power measurements retrieved are H100/H200-class nodes and one 130 kW Blackwell cluster; rack-level grid services for ~100–150 kW liquid-cooled racks are extrapolation.
  • Checkpoint- and all-reduce-aware power smoothing — the synchronized power spikes from collective communication and checkpointing that motivate “power smoothing” — have no dedicated peer-reviewed study in this corpus (only adjacent scheduling work 36).
  • Digital-twin control loops need a validated deployment before the “real-time engine” claim is taken seriously.

Confidence and limitations

Confidence: moderate. The device- and workload-level findings rest on multiple independent, partially full-text-verified measurements and agree on direction and rough magnitude (15–30% energy levers). The scheduling theme is simulation-dominated; its conclusions are directionally consistent but quantitatively fragile. The grid theme rests on one production deployment plus preprints; treat its headline numbers as promising rather than established.

Limitations of this review. (1) Access: 75 of 93 sources were read at abstract level — findings from those are hedged to what abstracts assert; 18 were read in full text, and all load-bearing numbers quoted in the body were spot-verified against the full text where available. (2) Twenty-one candidate sources were unreachable (14 paywalled with no OA copy, 7 SSRN-only) and are absent rather than cited; this biases the grid theme toward what is openly accessible. (3) English-language literature only. (4) Date cutoff 2026-08-07; the field is moving quickly and 2026 preprints dominate the grid theme. (5) Semantic Scholar was unusable from this network, so citation-graph snowballing leaned on OpenAlex/Crossref and arXiv; some citing-side work may have been missed.

Jump to references ↓

Evidence table

keydesignsamplemeasurefindinglimitationsconfidenceaccessnote
acun2023carbonanalytical-modelFramework analysis over datacenter design configurations (renewable mix, energy storage, workload scheduling) across geographic locations and workloads; open-sourced (CarbonExplorer)Trade-off between operational and embodied carbon; achievable 24/7 carbon-free operationCarbon Explorer analyzes the solution space for 24/7 carbon-free datacenter operation, optimizing the mix of complementary renewables, storage, and workload scheduling; results reported qualitatively (no quantitative results in the abstract).Framework-level analysis; abstract reports no validated numbers; results depend on assumed location and workload inputsmoderateabstract-onlyContributes a holistic carbon-aware datacenter design methodology (operational vs embodied carbon trade-offs) relevant to carbon-aware scheduling and 24/7 CFE planning.
adimora2025harmonicsystem-design+evaluationHPC resource management; validated in simulated environments and controlled testbeds (exascale-facility scenario); workloads/jobs not further specified in abstractEnergy consumption, throughput, and performance variability of scheduled jobsHARMONIC uncertainty-aware multi-objective scheduler achieves 10-19% energy reduction, 16-25% throughput improvement, and 18-32% performance variability reduction versus state-of-the-art schedulers, with multimillion-dollar annual savings projected per exascale facility.Results are from simulated environments and controlled testbeds, not a production exascale deployment; abstract gives no workload detail or energy model basis.moderateabstract-onlyEvidence that multi-objective (energy+resilience) scheduling can cut HPC energy use, supporting energy-aware HPC scheduling as a datacenter efficiency lever.
afzal2025gromacsbenchmark/measurement4 NVIDIA GPUs (A40, A100, L4, L40); 6 GROMACS biomolecular workloads + Pi Solver (compute-bound) and BabelStream/STREAM Triad (memory-bound); 200,000 MD steps, single precision, ECC on A40/A100Throughput vs GPU graphics clock frequency and power cap; power draw under capsUnder power capping, performance stays stable until workload/architecture-specific thresholds (on the A100 around 150 W for small and 250 W for larger benchmarks), with the A100 reaching near-peak performance at moderate caps and BabelStream stabilizing just below 200 W (Pi Solver below 140 W on all devices), while the low-power L4 degrades sharply under restricted caps.Single-GPU only (no multi-GPU scaling); synthetic benchmarks never approach TDP; results specific to GROMACS-style molecular dynamics workloadshighfull-textDirect measurement evidence that power capping has workload- and architecture-specific thresholds and that high-end GPUs (A100) tolerate caps with near-peak performance, supporting GPU power capping in HPC.
ali2023performancesystem-design+evaluationNVIDIA Ampere (training) and Volta (portability) GPUs; SPEC-ACCEL/DGEMM/STREAM microbenchmarks for training; LAMMPS, NAMD, GROMACS, LSTM, BERT, ResNet50 for evaluationAccuracy of predicted execution time and power across the GPU DVFS design space; energy savings at given performance lossDNN-based DVFS frequency selection models reach 89-98% accuracy on Ampere and >93% on Volta, achieving maximum energy savings of 27% at a 1.8% performance loss.Limited workload set (6 applications); single-vendor (NVIDIA) GPUs; accuracy on broader production workloads unverified.moderateabstract-onlyShows ML-guided GPU DVFS can deliver large energy savings with negligible performance impact, supporting frequency scaling as a GPU power-efficiency lever.
amaral2023keplerbenchmark/measurementContainerized applications at process/container/Kubernetes-pod level across multiple architectures; per-process power measured in a controlled environment with RAPL and hardware countersPower estimation accuracy (mean squared error) of regression power modelsA generic power model trained on measured per-process power (hardware counters + RAPL as regressors) achieves MSE as low as 0.010 versus 0.16 for a simple ratio approach and 0.92 when trained on aggregated workload power.Requires a controlled measurement procedure per architecture; accuracy may degrade on unseen hardware/workload mixeshighabstract-onlyEnables fine-grained per-process/container power accounting (GHG-protocol consistent) that underpins datacenter power provisioning, capping, and tuning.
antici2023pm100dataset~230K jobs with power consumption values from the M100 production supercomputer at CINECA (Italy), derived from workload-manager logs and node power metricsJob-level power consumption values (dataset content); enables power prediction and cappingProvides a methodology and a novel large dataset of around 230K jobs with corresponding power consumption values from a production HPC system; no prediction-accuracy results reported in the abstract.Single-site dataset (M100/CINECA); collection methodology imposes structure on what can be predictedhighabstract-onlyPublishes the job-power data needed for workload-manager-level power forecasting and power capping research in HPC.
bashir2024sunkanalytical-modeln/a (conceptual argument about carbon footprint metrics for datacenter scheduling)total carbon footprint outcome of operational scheduling/placement decisionsThe paper argues that including embodied carbon emissions (a sunk cost) in operational job scheduling and placement decisions can lead to decisions that do not actually reduce total carbon footprint — reported qualitatively, no quantitative results in the abstract.Position/argument paper; no empirical evaluation or numbers presented in the abstract; the claim's validity depends on the specific lifecycle accounting assumed.moderateabstract-onlyConceptual grounding for carbon-aware scheduling: operational decisions should be driven by marginal/operational carbon rather than amortized embodied carbon, avoiding the 'sunk carbon fallacy'.
beitsayadeh2026benchmarkingbenchmark/measurement739 standardized MLPerf Inference v5.0 submissions across public and private cloud infrastructures; efficiency estimated as throughput normalized by vendor-reported TDPinference throughput and energy efficiency (performance-per-watt proxy) by cloud type, accelerator model, and their interactionNo significant cloud-type effect on inference efficiency, but significant accelerator-model differences with NVIDIA H200-SXM-141GB showing approximately 45% higher median efficiency, and a log-linear specification best fitting multiplicative efficiency scaling across accelerator generations.Efficiency is a TDP-normalized proxy, not measured power; cross-sectional secondary analysis of self-reported MLPerf submissions; vendor TDP may not reflect real operating power.moderateabstract-onlyProvides a reproducible, public-data framework for comparing inference performance-per-watt across accelerators and cloud types, useful for hardware-efficiency benchmarking in AI datacenters.
bharadwaj2023predictsystem-design+evaluationn/a (position/design on GPU DVFS circuits; IVR transition times shrinking from microsecond to nanosecond regime)Energy efficiency gains from fine-grain predictive DVFSArgues that nanosecond-scale DVFS transition times unlock fine-grain predictive DVFS mechanisms that adapt to rapid workload fluctuation; results reported qualitatively (no quantitative results in the abstract).No evaluation or numbers in the abstract; motivational/position content onlylowabstract-onlyBackground on why predictive (rather than reactive) fine-grain GPU DVFS is the direction for energy efficiency.
bhavsar2026carbonsimulationSimulated AI job-level workloads (MLPerf-style profiles) combined with ComStock building loads for California datacenters; CAISO hourly carbon intensity, Open-Meteo temperature forecasts, PG&E time-of-use pricing as inputsEstimated CO2 emissions, total energy cost, and cooling demand of carbon-optimized vs baseline schedulesInitial simulations suggest the hybrid rule+ML carbon-aware scheduler could reduce carbon emissions and energy costs by 5-10% for flexible workloads, with greater gains during high renewable output or extreme heat (context: US datacenters consumed 176 TWh in 2023, 4.4% of national use, projected up to 12% by 2028).Simulation-based and explicitly preliminary ('initial simulations suggest'); results depend on author-defined workload flexibility assumptions and California-specific data.lowabstract-onlyIllustrates carbon-aware scheduling with public grid/environmental data (carbon intensity, temperature, TOU pricing) and hybrid rule-based + ML decision making for AI workloads.
bika2025poweringanalytical-modelHypothetical 1-GW AI data center powered off-grid by solar PV, battery, and underground salt-cavern hydrogen storage (Perspective paper with LCOE modeling)Levelized cost of energy and annual cost of firm power supplyA solar + battery + hydrogen configuration achieves firm power for a 1-GW AI data center with energy savings of $1.1 billion yr-1 compared to a battery-only configuration.Perspective/back-of-envelope analysis, not a deployed system; costs of salt-cavern hydrogen infrastructure and electrolyzer/fuel-cell maturity are idealized.lowabstract-onlyArgues off-grid self-generation (with seasonal hydrogen storage) is the key pathway for firm AI-datacenter power, relevant to grid-interconnection-constrained siting.
brenner2025strategicanalytical-modelStylized two-zone electricity market with flexible (shiftable) data center load; bilevel economic-dispatch/consumer-cost model; illustrative numerical example with flexible load at 20% of demandSocial welfare, locational marginal prices, and value of transmission expansion under strategic load shiftingStrategic price-driven load shifting can be socially inefficient (illustrative example: LMP at zone A drops from $25/MWh to $0/MWh while total system generation cost rises) and can fully offset the benefits of transmission expansion, so conventional price signals misrepresent system value.Theoretical model with illustrative numbers only; no empirical validation against real operator behavior or grid data.moderatefull-textImportant caution for carbon/price-aware geographic workload shifting: naive price signals can backfire at grid scale.
cao2025transformingsystem-design+evaluationCase study on a large-scale datacenter hall: 500 CFD/HT scenarios (340k mesh cells, supply temp 18-24 degC, server loads 1000-2000 W), surrogate trained on 4 NVIDIA V100 GPUs, inference on 1 V100Median absolute temperature prediction error and inference latency of a physics-informed neural surrogate vs full CFD/HT simulationThe PIML surrogate predicts thermal and airflow profiles with a median absolute temperature error of 0.18 degC (most errors within 2.5 degC) at ~0.01 s inference latency on one V100 GPU, achieving real-time simulation vs time-consuming CFD/HT (no explicit speedup factor stated).Single case study; error bounds of up to 2.5 degC in parts of the hall; requires extensive domain expertise and high-fidelity simulation data to build; no direct energy-savings measurement.moderatefull-textKey evidence for cooling/digital-twin review: physics-informed ML surrogates plus digital twins (Omniverse/PhysicsNeMo) enable real-time thermal simulation for proactive DC thermal management.
carastansantos2025schedulingsimulationMarconi 100, a 980-node supercomputer; workload submission logs with power monitoring dataSupercomputer power/energy consumption vs scheduling performance (QoS, starvation, deadline behavior)Simulation shows a lightweight history-based power-prediction method feeding a greedy-knapsack power-capping scheduler improves energy management of a large-scale supercomputer compared to energy-unaware scheduling with no significant negative performance impact; magnitude reported qualitatively.Simulation on one platform's logs; abstract reports no exact energy-savings percentagesmoderateabstract-onlyEvidence that lightweight power prediction suffices for power-capped HPC scheduling without QoS loss.
chien2023reducingsimulationWorkload model of ChatGPT-like generative AI inference; request direction policies Local, Balance, CarbonMin compared for power use and carbon impactPower use and carbon emissions of inference request direction approachesReported qualitatively - the abstract describes the workload model and policy comparison but gives no quantitative results.Abstract only; no numbers for the stated comparison; simulation-based workload model rather than measured deployment.lowabstract-onlyEarly framing of request-direction (spatial load shifting) for carbon-aware generative AI inference, precursor to later quantitative work.
clark2024learningsystem-design+evaluationCONDOR ML method mapping datacenter configuration (power, performance, load characteristics) to an objective combining DR savings, energy cost, and workload QoS complianceSpeed and accuracy of average power estimates and flexibility forecasts for demand-response participationCONDOR achieves speed increases of around 15,000x in computing accurate power/flexibility forecasts compared to simulation-based estimation, enabling large real-world datacenters to participate in DR programs without debilitating computational overhead.Abstract gives no accuracy numbers for the learned model relative to simulation ground truth, nor deployment-scale details; surrogate accuracy under unseen configurations is unquantified here.moderateabstract-onlyML-based flexibility/power forecasting for datacenter demand response - addresses the computational bottleneck of simulation-based DR estimation.
crozier2025potentialsurvey/reviewReview of data center flexibility literature (mechanisms incl. protection systems, pre-cooling, CPU/GPU clock scaling, smart job scheduling; cryptocurrency mining) plus authors' analysis of real energy usage dataPotential scale and cost of data center demand flexibility; price sensitivity of data center electricity consumptionMany flexibility mechanisms have been proposed but little is understood about their scale or cost, and real energy usage data suggest data centers are not likely to be sensitive to electricity price except during extreme events.Review-level synthesis; the underlying real energy data (n, centers, period) is not specified in the abstract.moderateabstract-onlyProvides a grounding caution for grid demand-response potential of datacenters: technical mechanisms exist but economic responsiveness may be weak.
damico2024energysimulationTwo cooperating CPU clusters in a heterogeneous multi-cluster HPC machine; workloads modeled from real-world applications; EAMC policy implemented in SlurmWorkload energy consumption, response time, makespan, and cluster utilization vs runtime-minimizing and energy-minimizing baselinesThe EAMC policy reduces response time and makespan by up to 25% and 6% while saving up to 20% of total energy versus runtime-minimizing policies, and by 49%, 26%, and 6% respectively versus energy-minimizing policies.Evaluation is simulation-based; only CPU clusters (no GPUs); results depend on the accuracy of the predictive performance/energy models for job-resource combinations.moderateabstract-onlyPower-aware HPC scheduling: energy-aware job placement across heterogeneous clusters/frequencies, showing energy savings are compatible with response-time improvements.
das2026followanalytical-modelSeven-factor composite scoring model across all 50 U.S. states using renewable curtailment-to-Fortune-500 density ratio; validation against Microsoft's announced investment pipelineState-level suitability ranking for training vs inference data center siting; installed capacity shareStates with the strongest case for training co-location (Nebraska, North Dakota, Iowa, Oklahoma, South Dakota) hold less than 1% of installed U.S. data center capacity, and a 500 MW training cluster in rural Nebraska would cost roughly half the energy of one in Ashburn, Virginia; shifting model weights toward training raises North Dakota 36 places and Nebraska 18 while Georgia falls 24.Composite scoring is an analytic argument rather than an empirical measurement; capacity and energy-cost figures are model-derived estimates.moderateabstract-onlyArgues training (unlike inference) can be sited at high-curtailment, low-cost locations - a spatial carbon/curtailment-aware siting strategy.
desai2025adaptivesystem-design+evaluationGPU systems with tree-based ML model predicting optimal power cap from GPU utilization, memory utilization, temperature, and frequency (workloads/n not specified in abstract)Energy consumption, operating temperature, and execution time under predicted power capsThe ML-based power-capping approach achieves a maximum energy saving of 12.87% and temperature reduction of 11.38% with only a 2.69% increase in execution time.Abstract omits hardware model, workload set, and number of experiments; single-node GPU setting only.moderateabstract-onlyEvidence that learned power caps can trade ~13% energy and ~11% temperature for <3% slowdown, directly supporting GPU power capping in datacenters.
desislavov2023trendsanalytical-modelAnalysis of relevant CV and NLP models and their consolidated (1-2 year later) implementations, focusing on inference rather than trainingGrowth of inference energy consumption relative to performance and parameter-count growthFor sustained performance increases, inference energy consumption grows much more softly than parameter/performance scaling laws suggest, once hardware efficiency gains and consolidated implementations are accounted for — reported qualitatively, no quantitative results in the abstract.No numbers in the abstract; trend conclusions depend on assumptions about hardware efficiency improvements and which implementations are counted.moderateabstract-onlyCounterpoint to parameter-scaling doom narratives: argues inference energy growth is sub-exponential because hardware FLOPS efficiency and algorithmic consolidation dominate raw parameter growth.
dokuchaev2025eltosystem-design+evaluationResNet18 and ResNet50 vision models served via NVIDIA Triton Server on an NVIDIA GeForce RTX 3060Ti under varying batch sizes and request ratesWeighted sum of normalized predicted latency and energy (scaled operational cost) of inference batchingELTO's cost-based batch-size selection reduces average scaled operational cost by 75.5% for ResNet18 and 48.2% for ResNet50 relative to a no-batching strategy.Single consumer GPU; two vision models only; results are relative to no-batching baseline rather than optimal batching.moderateabstract-onlyShows dynamic batching is a major energy-latency lever in ML inference serving, with order-of-magnitude cost reductions on small models.
doshi2026nightsimulationSimulated mixed-GPU cluster (8-GPU baseline); 13 scheduling rules; synthetic patient-style jobs with urgency tiers and hard deadlines; time-of-day carbon traces; 2000 jobs/run; gate and geo grids; up to 48 jobs/hourAverage kg CO2e per policy, critical/overall deadline-miss rates, p95 turnaround, weighted tardinessThe overnight-window rule closes about 78% of the average carbon gap between urgency-only scheduling and carbon-urgency mixing (CUCA_0.45) while missing ~2.4x fewer critical deadlines, whereas the carbon-only CarbonShift rule lets about 46% of the most urgent jobs miss their deadlines.Authors flag results as exploratory (wide run-to-run spread, no statistical adjustment); synthetic arrivals; operational carbon only (no Scope 3, no migration/data-transfer overheads); geo test uses one shared carbon shapemoderatefull-textShows simple night-window batching approximates carbon-aware scheduling when clinical deadlines dominate, and that carbon-only policies violate SLOs - relevant to carbon-aware scheduling under deadline constraints.
dossantosgonalves2026investigatinbenchmark/measurementAI workloads on GPUs under DVFS (abstract truncated; details of GPUs and workloads unavailable)Energy efficiency of AI workloads as a function of DVFS settingsNo quantitative results in abstract - the text is truncated before any findings are reported.Truncated abstract provides only motivation; results, hardware, and methods unavailable.lowabstract-onlyPlaceholder for a DVFS-on-AI-workloads measurement study; cannot be used for quantitative claims.
du2026tokenscase-studyNumerical case study of a representative multi-campus LLM-inference data center (GPU clusters with on-site gas turbines, batteries, PV; grid price and carbon signals; 15-min demand-response timescale; MILP co-optimization)Total data center operating cost and served token volume under quantization-enabled demand responseThe quantization-enabled DR framework reduces total data-center operating cost by 34.3% without curtailing served token volume, validating model quantization as an IT-side flexibility lever.Simulation case study, not a real deployment; quantization-to-power parameters are modeled rather than measured; 15-min timescale abstracts away millisecond-level serving dynamics.moderatefull-textNovel lever for grid demand response: LLM quantization (precision selection) as dispatchable flexibility, alongside price/carbon-aware multi-campus co-optimization.
dunlap2026quantifyinganalytical-modeln/a (eleven-parameter framework; U.S. grid resource adequacy planning context for AI data centers)Quantified flexibility/resource-adequacy contribution of AI data centersPresents an eleven-parameter framework for quantifying AI data center flexibility as a resource adequacy asset, arguing against treating AI data centers as inflexible firm loads; abstract truncated, results reported qualitatively.Abstract is truncated; no results or validation visiblelowabstract-onlyFrames AI datacenter demand flexibility as a grid resource-adequacy asset, relevant to grid demand response and interconnection planning.
fan2025elliesystem-design+evaluationIntel AI PC platform with integrated CPU, GPU, and NPU; diverse LLMs and prompt types; offline latency/power profiling plus a lightweight output-token-length predictorEnergy consumption, Energy-Delay Product (EDP), and latency of dynamically chosen prefill/decode device mappings vs static GPU-only mappingWhen optimizing for EDP, ELLIE reduces energy consumption by 1.8x, improves EDP by 1.5x, and achieves latency comparable to GPU-only inference, on average across diverse LLMs and prompt types.Single edge platform class; benefit is conditional on prompt characteristics, model, and hardware; runtime prediction overhead and worst-case behavior not quantified in the abstract.moderateabstract-onlyPhase-aware (prefill/decode split) energy optimization for LLM inference on heterogeneous edge SoCs - evidence that heterogeneous device mapping plus output-length prediction can cut energy at equal latency.
ferdaus2025evaluatingbenchmark/measurementFour AI accelerators - NVIDIA A100, Intel Habana Gaudi2, Graphcore Bow-Pod64, GroqRack - on MLPerf BERT-Large and ResNet50 to a common target accuracyEnergy consumption, throughput, and energy efficiency (per-benchmark) for training and inferenceReported qualitatively per accelerator: Gaudi2 gave highest throughput and lowest energy for ResNet50 training and inference, Graphcore highest training energy efficiency, A100 lowest time and energy for BERT-Large inference, and Graphcore highest throughput in both BERT pre-training and inference; no numeric energy values are given in the abstract.Initial study on two benchmarks; vendor-provided power tools and optimized models may not be directly comparable; no absolute energy numbers in abstract.moderateabstract-onlyCross-vendor accelerator energy benchmarking shows winner varies by workload, cautioning against single-hardware assumptions in AI-datacenter efficiency analysis.
gheni2026operationalcase-study21 months of operational data from the direct liquid-cooled HAWK supercomputer run at different supply water temperatures (17 to 25 degC), plus datacenter-level power/thermal model simulationsLiquid cooling system power consumption, heat transfer to server room air, IT equipment power, and overall datacenter energy performance vs supply water temperatureRaising supply temperature from 17 to 25 degC reduced liquid cooling system power consumption by 63.3% at an outdoor wet-bulb temperature of 19 degC, but nearly doubled heat transfer to server room air (from ~2.6% to ~5.0%) and increased IT equipment power by about 3.15%.Single site (HAWK) and single climate condition for the headline number; trade-offs are specific to the facility's cooling architecture and environmental conditions.highabstract-onlyQuantitative field evidence for supply-water-temperature optimization in direct liquid-cooled datacenters and for holistic cross-system (liquid + air + IT) modeling of cooling control.
gnibga2026improvingsystem-design+evaluationn/a (design proposal for underground thermal energy storage (UTES) for datacenter cooling)Cooling power share of datacenter power, IT capacity, cooling space use, datacenter lifetimeMotivates UTES with the claims that cooling can be up to 30% of overall datacenter power and rack densities grew from under 20 kW (2010) to over 120 kW (today) with 1 MW possible by 2029; UTES results reported qualitatively.Abstract provides motivation and context numbers only; no evaluation or quantitative UTES resultslowabstract-onlyThermal energy storage as cooling-side flexibility that relieves peak-grid-coincident cooling loads and extends datacenter IT capacity/lifetime.
gonalves2026scalablesystem-design+evaluationControlled small-scale testbed (Dell PowerEdge R440 servers, sensors at 1 Hz, Raspberry Pi 4 gateway, CPU workload varied 10-90% via stress-ng) with cloud-based LSTM forecastingTotal energy consumption, PUE, LSTM vs linear-regression forecast error (MAPE) at 1/6/24 h horizons, and cloud operating costThe digital twin framework reduced total energy consumption by approximately 10-10.4%, improved PUE from about 1.85 to 1.70, achieved LSTM MAPE of 4.0%/7.8%/14.0% at 1/6/24 hours (vs 6.2%/11.3%/19.0% for linear regression), at ~USD 100-120 monthly cloud cost.Small-scale, partially simulated environment; authors explicitly note limited generalization to large-scale datacenters; single testbed and workload generator.moderatefull-textDemonstrates a low-cost, IoT+cloud+LSTM digital twin for DC energy monitoring/forecasting with measured PUE and energy improvements in a constrained environment.
guan2024wattscopesystem-design+evaluationProduction datacenter workload; server- and rack-level aggregate power measurements (non-intrusive, no OS/app access); ML disaggregation of building power adapted to serversNormalized mean absolute error of application-level power estimationWattScope estimates per-application power from aggregate server/rack measurements with high accuracy, often <~10% normalized mean absolute error on a production workload.Evaluated on a single production workload; accuracy may degrade for workloads lacking the low-variability/high-periodicity power characteristics the method exploits.moderateabstract-onlyEnables application-level power attribution without privileged access - foundational metrology for datacenter power monitoring and per-workload efficiency.
hisaharo2024optimizingsystem-design+evaluationRedesigned LLM inference cluster running a modified GPT-Neo model (advanced interconnects, high-bandwidth memory, energy-efficient power management)Throughput, latency, energy consumption, scalabilityReports that a redesigned inference cluster with architectural and algorithmic changes achieved substantial improvements in throughput, latency, and energy consumption versus baseline models; no quantitative results in the abstract.No numbers, hardware details, or methodology in the abstract; claims not verifiable from abstract alonelowabstract-onlyAnecdotal evidence that cluster-level redesign (interconnect, memory, power management) can improve LLM inference energy efficiency.
huang2026dynamicsystem-design+evaluationAlibaba production traces replayed in simulations plus a testbed; heterogeneous servers; latency-sensitive cloud applications under power over-subscriptionRequest latency, power utilization, and capping decision quality under power-adjustment-interval constraintsPPE-DPC (prediction-based, power-efficiency-aware dynamic capping) outperforms existing capping solutions on latency and power utilization, and only at 3x the true prediction error does it become weaker than the existing optimal capping algorithm (exact magnitudes not given in abstract).Abstract gives no numeric latency/power improvements; simulation+testbed evidence rather than full production deployment.moderateabstract-onlyAddresses cluster-level power capping for latency-sensitive workloads, including robustness to prediction error - directly relevant to oversubscribed AI datacenters.
hur2026seasonalsimulationJeju Island power system case study: 16 operating scenarios (4 seasons x 4 time-of-day periods), DC optimal power flow co-optimizing generation dispatch and geographically distributed workload allocationCongestion-constrained datacenter integration limit (MW) and thermal violations under fixed vs spatially flexible workload allocationUnder fixed allocation the congestion-constrained integration limit ranges from 200 MW (summer evening) to 360 MW (spring afternoon), spatial redistribution expands the limit by 80-160 MW across all scenarios without transmission reinforcement, and ~30% workload migration capability suffices to eliminate thermal violations in most conditions.Single-island case study; results hinge on modeled wind availability, workload mobility assumptions, and OPF simplifications; no real workload telemetry.moderateabstract-onlyShows seasonal-hourly granularity (not peak-based analysis) is required to size spatial workload flexibility for renewable-rich grids - directly relevant to grid demand-response planning.
jacquet2026untanglingsystem-design+evaluationCloud shared GPU settings where accelerators are leased to diverse clients; job-level power inference (abstract truncated after motivation)Job-level GPU power consumption attribution/estimation in multi-tenant cloudsNo quantitative results in abstract - the abstract contains only the motivation sentence.Truncated abstract; method, evaluation, and results unavailable.lowabstract-onlyPlaceholder on job-level GPU power attribution for multi-tenant AI clouds; essential for fair energy accounting but not yet evidenced.
jain2026minossystem-design+evaluation18 graph analytics, HPC, HPC+ML, and ML workloads on GPU-accelerated HPC clustersProfiling time reduction; power and performance prediction errorMinos, a low-cost profiling-based classifier for GPU workload power/performance behavior, reduces profiling time by 89% when predicting frequency-capping behavior of unseen applications and achieves mean errors of 4% for power and 3% for performance predictions across 18 workloads, improving on state-of-the-art approaches by 10%.Classification tuned on specific cluster workloads; assumes stable workload classes; single-cluster evaluation contexthighabstract-onlyLow-cost workload classification enables power/performance management (e.g., frequency capping) on power-constrained HPC clusters.
jay2023experimentalbenchmark/measurementSeveral software-based power meters for CPU and GPU infrastructures evaluated against high-precision physical power meters while running various intensive workloadsAccuracy of software power meters vs physical meters; qualitative strengths/limitations per toolReported qualitatively - the abstract describes a comparative experimental evaluation of software-based power meters against physical meters but gives no numeric accuracy results.Abstract gives no tool names or numeric errors; coverage of meters/workloads unspecified.moderateabstract-onlyMethodological reference for choosing power measurement tooling in datacenter/GPU efficiency studies - grounding for measurement validity.
jayatilleka2026reliabilityanalytical-modeln/a (methodology for liquid cooling system (LCS) reliability/availability estimation across New Product Introduction phases)System availability integrating MTBF, MTTR, and Mean Logistics Delay Time (MLDT), with PERT-based MTTR estimatesThe paper advocates availability (integrating MTBF, MTTR, MLDT) over MTBF alone as a more meaningful reliability metric for liquid cooling systems and proposes PERT-based MTTR allocation across reliability-critical components — reported qualitatively, no quantitative results in the abstract.No empirical data or quantitative results presented in the abstract; MTTR/MLDT data are acknowledged as often lacking in early development phases; methodology-level contribution.lowabstract-onlyReliability/availability modeling for liquid cooling infrastructure - relevant to cooling-system dependability for high-density AI datacenters, but contains no energy numbers.
jegham2025howbenchmark/measurement30 state-of-the-art LLMs (OpenAI, Anthropic, Meta, DeepSeek) in commercial datacenters on DGX A100/H100/H200/H800; public API latency/throughput data + company PUE/WUE/CIF multipliers; hardware inferred statisticallyPer-prompt energy (Wh), water, carbon emissions; DEA eco-efficiency rankingThe most energy-intensive models exceed 29 Wh per long prompt (over 65x the most efficient systems), and a 0.42 Wh short query scaled to 700M queries/day aggregates to annual electricity comparable to 35,000 U.S. homes, evaporative freshwater equal to the annual drinking needs of 1.2M people, and carbon emissions requiring a Chicago-sized forest to offset.Energy/water/carbon are estimates from inferred hardware and published multipliers, not direct measurements; depends on public API data and company disclosuresmoderatefull-textInfrastructure-aware per-prompt benchmarking (PUE/WUE/CIF scaling) quantifying LLM inference footprint - key for inference energy benchmarking and policy.
jiang2025hesharesystem-design+evaluationHeterogeneous multi-GPU datacenter systems; energy-aware task scheduling plus adaptive MPS and DVFS configuration per GPU; compared against a state-of-the-art frameworkAverage energy cost and job completion time of multi-task GPU sharingHeShare reduces average energy costs by 26% and improves job completion time by 31% compared to the state-of-the-art framework, balancing energy efficiency and performance.Abstract omits hardware specifics, workload mix, cluster size, and the identity/definition of the baseline framework; single comparison point.moderateabstract-onlyEvidence for combining GPU power capping/DVFS with sharing (MPS) in heterogeneous GPU scheduling to cut energy while improving completion time.
kakolyris2025throttllsystem-design+evaluationLLM inference traces; framework combining instance-size scaling and GPU frequency scaling with KV-cache/batch-size projections; compared against NVIDIA Triton serverEnergy consumption and energy efficiency under service-level objectives (SLOs)throttLL'eM achieves up to 43.8% lower energy consumption and an energy-efficiency improvement of at least 1.71x under SLOs compared to NVIDIA's Triton server, with an ML model scoring R^2 > 0.97 and missing performance by less than 1 iteration per second on average.Evaluation on LLM inference traces (not specified as production); results depend on SLO configuration and workload mix.moderateabstract-onlyDemonstrates SLO-preserving GPU throttling (frequency + instance scaling) for LLM inference - a direct GPU power-efficiency lever for AI datacenters.
krishnaram2026feedingbenchmark/measurement8x NVIDIA H100 GPU AI inference server with AMD EPYC vs Intel Xeon head nodes; representative inference workloads including multi-GPU benchmarks and a CPU-hosted recommenderNode-level power draw (CPU, memory, I/O, storage, cooling, GPU compute/memory/interconnect), throughput, and performance-per-watt by workloadBoth configurations peak at ~8-9 kW with 8 GPUs fully loaded (GPU-dominated), the AMD EPYC head node achieves up to ~10% higher multi-GPU inference throughput at similar power, the head node contributes <15% of system power when GPUs are saturated, and batching/parallelism improved throughput-per-watt by an order of magnitude in one case.Single server configuration and two specific CPU platforms; results are workload-dependent; measurements at one site without stated repetition statistics.moderateabstract-onlyNode-level energy breakdown for multi-GPU inference servers: quantifies when head-node choice matters and when GPU efficiency dominates; motivates batching, DVFS, and GPU power capping.
lechowicz2025carbonsystem-design+evaluation100-node Kubernetes cluster running data processing jobs with precedence constraints; time-varying carbon intensity signalsCarbon footprint and makespan/efficiency trade-off of schedulingA moderate configuration of PCAPS, a carbon- and precedence-aware scheduler, reduces carbon footprint by up to 32.9% without significantly impacting total efficiency.Prototype-scale (100-node) evaluation; carbon savings depend on grid signal quality and user-set carbon/makespan priority.moderateabstract-onlyShows precedence constraints are key for carbon-aware scheduling of data-processing jobs - extends carbon-aware scheduling beyond independent tasks.
li2023towardanalytical-modelHPC systems analyzed via hardware-component carbon footprint modeling, regional carbon intensity analysis, and experimental life-cycle characterizationCarbon footprint of HPC systems across hardware manufacturing and operational stagesThe work presents a comprehensive carbon-footprint analysis covering both manufacturing and operation stages of HPC systems — reported qualitatively, with no quantitative results in the abstract.Abstract contains no results or numbers; contribution is methodological (modeling + characterization framework).lowabstract-onlyLifecycle carbon accounting (embodied + operational) framework for HPC - complements carbon-aware scheduling by clarifying where emissions actually accrue.
liu2023efficientsystem-design+evaluationHardware testbed; DL inference workloads; knobs studied include batch size, GPU frequency, and GPU spatial sharing under latency and power constraintsThroughput under fixed latency and power-cap constraintsMorak, a multi-knob GPU resource management framework (spatial sharing + frequency and batch-size search), achieves up to 67.7% throughput improvement over several state-of-the-art baselines under tight latency and power constraints.Testbed-scale evaluation; workloads/models not enumerated in abstract; power cap levels unspecified.moderateabstract-onlyShows combining spatial sharing with DVFS/batch tuning maximizes inference throughput within power caps - core GPU power-capping technique evidence.
liu2025synergisingsimulationHierarchical data centers (cloud-fog-edge) coupled with power networks; day-ahead co-dispatch problem solved via distributed privacy-preserving optimization; numerical simulationsOperating cost, communication delay, and peak load balancing of the integrated DC-power systemReported qualitatively - the distributed approach is validated as effective, optimal, and scalable in simulations, with computing tasks delayable and migratable across hierarchical DCs to support peak load balancing of the power network; no magnitudes given in the abstract.Simulation-based; no real-system validation; convergence/optimality guarantees are for the proposed reformulation, not measured deployments.moderateabstract-onlyEvidence for DC-grid co-dispatch where IT workload migration serves power-network peak shaving, with privacy preserved across agents.
liu2026mixturesimulationData center microgrid with coupled workload scheduling, IT, cooling, and energy subsystems; multi-reward MDP solved by mixture-of-experts multi-critic deep RL; experiments against SOTA DRL and rule-based baselinesOverall energy consumption and carbon emissions of the DC microgrid under multi-objective controlReported qualitatively - the MoE multi-critic DRL outperforms state-of-the-art DRL and rule-based baselines in reducing overall energy consumption and carbon emissions; no numeric reductions are given in the abstract.Simulation experiments; abstract gives no magnitudes, baselines detail, or workload scenarios.moderateabstract-onlyMulti-objective DRL approach for jointly managing workload, cooling, and energy assets - relevant to DC microgrid energy/carbon management though quantitative evidence is absent.
liu2026publicanalytical-modelChina's AI-datacenter computing hubs as illustrative case; public-source inputs (regional carbon intensities, etc.) with author-defined workload mobility assumptions and a network-flow allocation modelEstimated CO2 emissions baseline (Mt CO2) and percentage emission reductions under PUE-only, flexibility, and workload-aware allocation scenariosThe current baseline is estimated at 20.30 Mt CO2, with PUE-only improvement giving a 3.63% reduction, a verified 10% flexibility case 4.47%, and the main workload-aware allocation scenario 11.78% (an upper-bound renewable/storage sensitivity reaches 29.23% but is not treated as a headline result).Results rest on author-defined workload mobility assumptions and public-data quality; claim-tier rules separate main scenarios from exploratory stress tests; not validated against operator data.moderateabstract-onlyTransparent, tiered-evidence methodology for estimating carbon-aware spatial workload-allocation benefits from public data - useful template for decision-support rather than a deployment result.
lu2025thermalsystem-design+evaluationLLM inference serving in cooling-regulated datacenters (millions of jobs; Ray Serve-like scheduler assigning jobs and batch sizes to GPUs; Singapore 28-32C level-4 and EU 35C reference temperatures)Inference throughput, thermal-throttling probability, GPU frequency degradationIn cooling-regulated datacenters, existing schedulers increase the probability of thermal throttling by 10x with performance degradation up to 34.2%, while the proposed thermal-aware scheduler TAWS improves LLM inference throughput by up to 40.94% at 41C ambient.Evaluation appears simulation-based; results tied to specific ambient-temperature regimes and workload patternsmoderateabstract-onlyShows interaction between cooling regulation (higher setpoints) and GPU thermal throttling, motivating thermal-aware scheduling for AI datacenters.
luccioni2023powerbenchmark/measurement8x NVIDIA A100-SXM4-80GB node on AWS (us-west-2); task-specific models across 10 tasks plus multi-purpose Flan-T5 (222M-11B) and BLOOMz (560M-7B) families; 1,000 inferences per model/dataset, 10 repetitions each, measured with Code CarbonEnergy (kWh) and carbon (g CO2eq) per 1,000 inferencesPer-1,000-inference energy ranges from 0.002 kWh (text classification) to 2.9 kWh (image generation, up to 11.49 kWh for the least efficient model), and multi-purpose generative models are orders of magnitude more expensive than task-specific ones (e.g., ~0.3g vs ~10g CO2eq per 1,000 inferences for extractive QA), with the whole study consuming 754.66 kWh and emitting 178.97 kg CO2eq.Single cloud region and GPU type; idle power of co-located GPUs included in measurements; model/hardware landscape has evolved since 2023.highfull-textFoundational measurement of LLM/generative-AI inference energy and carbon - anchors cost-per-inference estimates for AI datacenter operational efficiency.
luo2026requestbenchmark/measurementvLLM serving of Qwen2.5-7B/14B, Llama-3.1-8B, Mistral-7B on NVIDIA A800 (plus A40 and H100); static and continuous batching; 16 model/workload runs; NVML power sampling at 100 ms with idle-power subtractionNormalized L1 deviation of attribution rules (token-proportional, standalone-measurement, JCalib) from exact Shapley request-level energy; batch energy savings vs solo servingToken-proportional attribution deviates from exact Shapley energy by 0.440 normalized L1 under static batching and 0.458 under continuous batching (reproduced on three GPU types), JCalib reduces this to 0.116/0.177 at ~0.003 ms/request overhead, and batching saves 77% (static: 1989 J vs 8532 J) and 69% (continuous: 2545 J vs 8286 J) of energy versus solo serving.GPU-only power (60-75% of node power per cited work); attribution ground truth built from counterfactual replay of request subsets, not organic production traffic; prefill/decode stage energy not separated.highfull-textShows token-based energy attribution is unreliable under batched LLM serving and that measured Shapley ground truth plus a cheap calibration model enables fair request-level energy accounting for sustainability reporting.
ma2025powersystem-design+evaluationGPU servers with a host CPU and multiple GPUs processing ML inference workloads (abstract truncated before any experiments)ML inference performance under joint host-CPU + multi-GPU power capping (proposed; not yet evidenced)No quantitative results in abstract - it motivates a joint CPU+GPU server-level power capping solution but reports no evaluation.Abstract ends at motivation; no method details, experiments, or results available.lowabstract-onlyIdentifies the gap that server-level (CPU+multi-GPU) capping, not single-GPU capping, is what GPU servers actually need for power oversubscription.
ma2026illusionbenchmark/measurementSingle NVIDIA H200 SXM (700 W TDP, HBM3e); 4 approx. 4B-parameter models (GQA, MLA, Gated DeltaNet, Mamba2) served via vLLM BF16; BS 1-32; seq 1K-64K; 5 SM clock levels 390-1980 MHz; 5 power caps 280-700 W; NVML 50 ms power samplingEnergy per token (mJ/tok) for prefill and decode; throughput; power draw vs configured capAutoregressive decode across all four attention paradigms draws only 137-300 W on the 700 W H200 so no power cap (280-700 W) ever triggers, whereas SM clock locking recovers up to 32% of decode energy at less than 1% throughput loss and Pareto-dominates power capping at every matched operating point.Single GPU and decode-pool scenario; approx. 4B models only; ~44% of prefill configs under 100 ms use snapshot-power x latency energy fallback; firmware clock-clamping confounds require clock locking to controlhighfull-textCentral evidence for the review: power capping is structurally ineffective for memory-bound LLM decode; clock locking is the effective lever - directly challenges GPU power-capping practice.
maji2024untanglinganalytical-modeln/a (analysis of renewable energy attribution and PPA double-counting in grid carbon intensity estimation)Correctness of grid carbon intensity and organization-level carbon accounting when PPAs are (or are not) netted outThe work argues that private PPAs cause double counting of renewable generation in grid carbon reports, letting PPA-holding organizations understate emissions, and that there is no consensus method for PPA-aware carbon intensity — reported qualitatively, no quantitative results in the abstract.No quantitative results in the abstract; analysis is scenario/conceptual and depends on accounting-rule assumptions.lowabstract-onlyCautionary source for carbon-aware computing: carbon intensity signals and renewable credits can be double-counted, so scheduler-visible carbon metrics need consistent PPA attribution.
mavromatis2023frostsystem-design+evaluationML pipelines on O-RAN/5G platforms (RAN consumes 73% of network energy); energy profiling with online hardware reconfigurationML pipeline energy consumption, model accuracy, inference time delayFROST, a flexible reconfiguration method with online system tuning that profiles ML pipeline energy and limits power draw, achieves energy savings of up to 26.4% without compromising model accuracy or introducing significant time delays.O-RAN-specific context; evaluation scale and hardware not detailed in abstractmoderateabstract-onlyShows power limiting via online hardware tuning for AI pipelines in edge/telecom settings, extending power-capping evidence beyond datacenters.
menear2025energysystem-design+evaluationHPC cluster jobs (Slurm-based); per-job power prediction from LLM embeddings of job scripts; scheduler evaluated in simulation for shifting load to on-site solarPer-job power prediction MAE; MWh of load shifted to solar without throughput lossThe LLM-embedding-based power predictor reduces per-job power MAE by 15% versus the current state of the art, and the simulated scheduler shifts 4.0 MWh onto on-site solar without throughput loss.Scheduler results are from simulation; prediction evaluated on a single HPC environment; no absolute MAE values in abstract.moderateabstract-onlyPractical path to energy-aware HPC scheduling (HPC as a grid-responsive load) without modifying Slurm core - relevant to demand response and renewables self-consumption.
menear2026quantifyingcase-studyFive months (Jan-May 2025) of metered operations from a U.S. national laboratory HPC center; CPU job workloads; TOU energy prices plus monthly demand chargesTotal electricity bill and demand-charge reductions of tariff-aware scheduling vs the actual metered baselineThe optimal envelope (Stage A) reduces the total electricity bill by 7.97% (demand charges -17.56%), and the feasible offline schedule (Stage B) still reduces the bill by 5.73% (demand charges -12.62%), showing multi-percent full-bill savings from software-only tariff-aware scheduling.Offline/retrospective analysis on historical data; CPU workloads only (no GPU/AI jobs); single facility and tariff structure; assumes perfect foresight in Stage A.highabstract-onlyQuantifies the economic value of time-of-use + demand-charge-aware HPC scheduling with real metered data and released reproducible methodology - strong evidence for power-aware scheduling economics.
naug2025lcsimulationGymnasium-based RL benchmark environment built on a Modelica digital twin of Oak Ridge National Lab's Frontier supercomputer cooling system (site-level cooling towers to cabinets and server blade groups)Multi-objective trade-off between local thermal regulation and global cooling energy efficiency; interpretability of distilled control policiesPresents LC-Opt, an RL benchmark environment for liquid cooling control (supply temperature, flow rate, valve actuation, cooling-tower setpoints) with centralized/decentralized multi-agent RL baselines, policy distillation to trees, and LLM-based action explanations; no quantitative results reported in the abstract.Environment/baseline paper; no reported controller energy or thermal results; single-site digital twinmoderateabstract-onlyDemocratizes a high-fidelity liquid-cooling simulation for RL-based cooling control research in high-density AI datacenters.
niu2025energybenchmark/measurementSingle GPU node with 2 H100 GPUs; LLM inference engines vLLM, TensorRT-LLM, DeepSpeed; lifecycle decomposed into setup (init + model loading) and token generation stagesPower consumption per lifecycle stage and per component (GPU, CPU, DRAM)Benchmarks inference-engine power consumption on a 2xH100 node with a fine-grained stage- and component-level breakdown to identify energy bottlenecks; results reported qualitatively (no numbers in the abstract).Single node and engine versions; no quantitative results in the abstractmoderateabstract-onlyComponent-level (GPU/CPU/DRAM) energy breakdown of the inference lifecycle helps locate inference energy bottlenecks.
patel2023splitwisesystem-design+evaluationBLOOM-176B and Llama-70B on DGX-A100 (400W) and DGX-H100 (700W); production coding and conversation traces (median prompt 1500/1020 tokens, median output 13/129 tokens); real vLLM implementation on Azure VMs plus event-driven cluster simulatorThroughput, cost, power, and phase-level latency (TTFT/TBT/E2E) of LLM inference clustersSplitting prompt and token-generation phases onto separate machines yields 1.4x higher throughput at 20% lower cost, or 2.35x more throughput at the same cost and power budgets, and the token-generation phase tolerates >50% power capping (700W to 350W) with almost no latency impact while the prompt phase is highly power-cap-sensitive.Cluster-level results rely on an event-driven simulator with a performance model; characterization on two LLMs and two trace types.highfull-textKey evidence that LLM inference token generation is memory-bound and power-cappable without SLO harm - central to GPU power capping and heterogeneous (e.g., older-GPU) deployment strategies.
piolant2025improvingbenchmark/measurementMultiple NVIDIA GPU architectures; compute-intensive GEMM kernel study, then dense linear algebra (matrix multiplication, Cholesky factorization) on a heterogeneous node with 4 GPUs under a task-based runtimeEnergy efficiency (energy per operation) under different power capsCompute-intensive GEMM kernels are up to 30% more energy efficient when the GPU is capped at 55-70% of its TDP, and on a 4-GPU heterogeneous node applying a cap to all GPUs improves matrix-multiplication energy efficiency by up to 24.3% (double precision) and 33.78% (single precision).Linear-algebra kernels only; results platform-dependent; cap levels tuned from the GEMM studyhighabstract-onlyDirect evidence that capping GPUs at 55-70% TDP improves HPC energy efficiency, and that runtime schedulers can exploit heterogeneous capped devices.
purnachandrareddy2026cargosimulation168-hour, 8-region workload of 9,479 jobs derived from public AI-cluster trace characteristics; 5 reference policies + 2 ablations; six seeds; inference, fine-tuning, and pre-training jobsCO2e emissions, inference SLO violations, deadline misses, p99 latency, planner convergenceCARGO achieves an 85.2% reduction in CO2e emissions versus a latency-greedy production baseline (340.9 vs 2306.7 kg, mean over six seeds) while maintaining 0.0% inference SLO violations and 0.0% deadline misses, whereas aggressive carbon-greedy placement reaches lower emissions only at the cost of 14.4% SLO violations.Simulation-based evaluation on one derived workload profile; results depend on trace realism and carbon-intensity assumptionsmoderateabstract-onlyGeo-distributed carbon-aware GPU scheduling with explicit SLO feasibility (NSGA-III planner + PPO controller) - key carbon-aware scheduling evidence.
qi2026safetysystem-design+evaluationShared cold-source hybrid air-liquid cooling system (cooling towers, pumps, CRAH, CDU); simulation with held-out 2024 Hefei summer meteorological data; deployment on a physical China Unicom testbed in GuangdongMean COP, total cooling power, thermal constraint violations, and CDU pump power vs PID baselineThe safety-constrained multi-agent D3QN framework achieves mean COP 8.57 vs 7.97 for the PID baseline (7.6% improvement) with cooling power reduced by 7.0% and zero thermal violations in simulation, and on the physical testbed reduces CDU pump power by 6.2% and 10.4% at 124 kW and 166 kW steady-state loads.Training data from one climate (Hefei summer) with deployment in a different zone (Guangdong) - validated but single testbed and two load levels; simulation-trained policies may not cover all failure modes.moderateabstract-onlyEvidence for multi-agent RL with decoupled safety mechanisms in coordinated cooling control, including rare sim-to-real deployment results for hybrid air-liquid cooling.
qiao2024energysystem-design+evaluationLinux datacenter servers; per-process workloads measured via eBPF at millisecond-scale granularityPer-process energy accounting accuracy/overhead; effect of energy-informed scheduling policiesIntroduces Wattmeter, an eBPF-based per-process energy accounting framework for Linux (no kernel changes, millisecond-scale granularity, low overhead) plus two proof-of-concept scheduling policies (energy equalization and per-process energy capping); results reported qualitatively (no numbers in the abstract).No quantitative results in the abstract; policies are proof-of-conceptmoderateabstract-onlyOS-level per-process energy accounting enables energy-aware scheduling and per-process energy caps in datacenters.
radman2026impactbenchmark/measurementGPU energy consumption under DVFS for real-time systems (abstract truncated; hardware and workloads unspecified)Energy consumption vs performance under DVFS in real-time GPU environmentsNo quantitative results in abstract - the text is truncated at the motivation stage.Truncated abstract; no method, results, or hardware details.lowabstract-onlyPlaceholder on DVFS for real-time GPU workloads; cannot support quantitative claims in the review.
repeva2025anomalycase-studyMultidimensional time series from multiple air-conditioner sensors in a real containerized data center monitoring systemAnomaly-detection effectiveness and lead time for cooling-equipment malfunctions (freon leakage)Among three open-source libraries (Merlion, Darts, Anomaly Detection Toolkit), methods based on LevelShiftAD and VolatilityShiftAD were the most effective at identifying malfunctions such as freon leakage approximately two and a half hours before complete equipment failure.Single facility and fault type; method ranking may not generalize to other cooling equipment or failure modesmoderateabstract-onlyML-based early fault detection for datacenter cooling equipment, supporting cooling reliability and maintenance efficiency.
rrapaj2024powerbenchmark/measurementSix months of power measurements from NERSC's Cori and Perlmutter supercomputers; per-application and per-user power analysis; CPU to GPU-node transitionPower usage versus peak provisioned power and thermal design power (TDP) fractionProduction power usage varies considerably and is consistently significantly below peak provisioned power, and as NERSC transitioned to GPU-accelerated nodes the peak power capability increased while production workload power demand did not rise at the same rate, further decreasing the fraction of TDP used; magnitudes reported qualitatively.Two machines at one site; abstract gives no exact numbers; workload mix is NERSC-specifichighabstract-onlyEvidence that production HPC power stays well below TDP, supporting power-capped, over-provisioned machine designs and power-aware scheduling.
samsi2023wordsbenchmark/measurementLLaMA 7B/13B/65B inference on NVIDIA V100 (32 GB) and A100 (80 GB) GPUs, sharded across up to 32 GPUs, on Alpaca and GSM8K datasets, batch sizes 64-512, max generation lengths 256/512/1024; power-capping subset on 4x A100Energy per second, energy per decoded token, throughput, and inference time under GPU power capsLLaMA 65B inference draws on the order of 300 W to 1 kW across 8-32 shards and about 3-4 J per decoded output token (max length 512), and power capping from 250 W to 175 W increases inference time by an average of 6.7% while reducing total energy by ~21.8-24.0% (at 150 W: +15.3-21.7% time, -32.8-34.7% energy).GPU energy only (rank-0 node extrapolated to multi-node); power-capping experiments are a limited set (one model, one dataset, three cap levels); sharding leaves only 20-25% GPU memory utilized (over-provisioning).highfull-textFoundational LLM-inference energy benchmark: documents per-token energy on V100/A100 and the time-vs-energy trade-off of GPU power capping - a core reference for inference power-efficiency reviews.
shen2026gridsimulationHigh-fidelity EnergyPlus co-simulations of an AI datacenter cooling plant (CRAC fans, chillers, cooling towers, pumps; H100-class GPUs with >700 W TDP and >50 kW racks); context features from queue and day-ahead/real-time grid signalsThermal violation rate, operational cost (energy + carbon + thermal penalty), robustness premium vs Min-Max MPCA Contextual Distributionally Robust Optimization (CDRO) controller with adaptively calibrated Wasserstein radius achieves near-zero thermal violations under extreme AI workload spikes while reducing the operational cost premium of robustness by approximately 13.7 percentage points relative to standard Min-Max Model Predictive Control.Co-simulation rather than field deployment; performance depends on context features, radius calibration, and the modeled plant (cooling cited at 30-40% of facility energy)moderatefull-textGrid-interactive cooling control that safely unlocks demand-side flexibility under AI workload uncertainty - key for grid demand response from cooling.
shuai2024datasystem-design+evaluationUnity3D-based datacenter simulation platform receiving temperature-field information, with a deep neural network for device energy consumption prediction (environment details not specified in abstract)Device energy consumption prediction and datacenter energy efficiency evaluation via 3D temperature visualizationThe digital-twin platform renders 3D temperature cloud maps for energy-efficiency evaluation and uses a DNN to predict device energy consumption — reported qualitatively, no quantitative results in the abstract.No quantitative results or evaluation details in the abstract; prototype-level description.lowabstract-onlyIllustrates digital-twin visualization + DNN prediction for DC energy management; qualitative only, useful as an approach reference rather than evidence.
souza2023caspersystem-design+evaluationGeo-distributed web services hosted across cloud regions with varying carbon intensity; multi-objective optimization over carbon intensity and network latency constraintsCarbon footprint reduction vs baseline methods while respecting Service Level Objectives (latency)CASPER achieves substantial carbon emission reductions of up to 70% compared to baseline methods with no latency performance degradation.Abstract lacks evaluation scale, workload mix, region count, and baseline definitions; 'up to' headline may not reflect typical-case savings.moderateabstract-onlyCarbon-aware load balancing across cloud regions for interactive web services - evidence that carbon migration can be SLO-preserving.
spaan2026reducingsystem-design+evaluationGPT-3 training run as case study; fine-grained kernel-level DVFS on GPU(s) (GPU model not stated in available text); also tests data and tensor parallelism transferabilityEnergy consumption and slowdown (waste reduction) of LLM trainingKernel-level DVFS saves as much as 14.6% of GPT-3 training energy with only a 0.6% slowdown, versus 2% energy saving for a pass-level approach with no performance loss, and discovered frequencies transfer across data and tensor parallelism.Available full text is a partial HTML conversion (abstract, intro, background); GPU model and full experimental details not present; single-model case study.moderatefull-textReframes DVFS goal as performance-neutral 'waste reduction' rather than EDP trade-offs - argues fine-grained DVFS is adoptable without slowing LLM training.
springborg2023automaticsystem-design+evaluationSingle-node HPC system with SLURM; application-specific energy models via a Python plugin; HPCG benchmarkEnergy consumption of scheduled jobs (proof-of-concept)The energy-model-driven SLURM plugin approach demonstrates an 11% energy saving for the HPCG benchmark on a single-node HPC system.Proof-of-concept on one node and one benchmark; no multi-node or production validation.lowabstract-onlyShows energy-aware scheduling can be added to SLURM as a plugin; pathway toward scheduling jobs when energy is cheap/renewable.
stojkovic2024towardsbenchmark/measurementLlama-2-70B served with vLLM on an NVIDIA DGX-H100 (8 GPUs; tensor parallelism 2/4/8); GPU frequency 800-1980 MHz in 200 MHz steps; 9 input/output workload buckets (100/50, 500/128, 1024/256 tokens); latency SLOs at 5x solo-request TTFT/TBTTTFT, TBT, maximum throughput, power draw, and total energy vs frequency, parallelism, and batch sizeFrequency capping yields ~20% lower power for most configurations with no latency or throughput impact (halving frequency cuts throughput by only ~20% for small workloads), 1.6 GHz instead of 2 GHz gives about the same throughput at less than 80% of the energy, and reducing max batch size during low-throughput phases can cut energy by up to 15%.Single model (Llama-2-70B) and single platform (DGX-H100); generality to other models asserted via prior correlation results; GPU-level rather than full-system measurements.highfull-textSystematic quantification of DVFS/frequency-capping, tensor-parallelism, and batching levers for energy-efficient LLM serving under latency SLOs - core evidence for GPU power management in inference.
stojkovic2025dynamollmsystem-design+evaluationDGX H100 servers with vLLM; models Llama2-13B/70B, Llama3-70B, Mixtral-8x7B/8x22B, Falcon-180B; one week of Azure Coding and Conversation invocation traces; loads 650/2K/4K tokens/s; TTFT SLOs 250/400/2000 ms and TBT SLO 100 ms (5x isolated latency)Energy (Wh), operational carbon, customer cost, latency SLO (TTFT/TBT) complianceDynamoLLM, which dynamically reconfigures inference clusters (instance scaling, tensor parallelism, GPU frequency) under latency SLOs, conserves 53% energy and 38% operational carbon emissions and reduces customer cost by 61% in a large GPU cluster evaluation with production-level cloud traces.Traces from one cloud provider's two services; open models only; SLOs defined relative to isolated latency (5x); reconfiguration overheads managed but not detailed in abstracthighfull-textFlagship evidence that dynamic DVFS + parallelism + instance scaling yields large LLM inference energy/carbon savings under SLOs.
suarez2025energysurvey/reviewGerman HPC centers (case studies) where electric power per installation reaches and often exceeds 20 MW; national energy policy contextn/a (review of strategies: heterogeneous hardware, monitoring, high-temperature cooling, energy-aware scheduling, dynamic power management)Reviews state-of-the-art energy-efficiency strategies in German HPC centers, including heterogeneous architectures, advanced monitoring, high-temperature cooling, energy-aware scheduling, and dynamic power management; no quantitative comparisons reported.Descriptive survey; no measured energy or cost numbersmoderateabstract-onlyCatalog of in-production HPC energy-efficiency practices (cooling, scheduling, power management) useful as practice evidence.
sun2025learningsimulationCloud data center operational environment simulator built on real-world production traces from Alibaba; power capping as a partially observable MDP with model-based RLData center power load reshaping toward electricity price/market signals; capping decision optimalityReported qualitatively - numerical experiments validate the uncertainty-aware MBRL power-capping scheme as effective for reshaping power load to market signals; no magnitude results are given in the abstract.Simulator-based evaluation (albeit on Alibaba traces); abstract gives no numeric savings or comparison baselines.moderateabstract-onlyRL-based dynamic power capping aligned to electricity prices - evidence direction for price-responsive datacenter load shaping, pending quantitative detail.
syrigos2023eelassystem-design+evaluationCloud/fog/edge continuum resources; ML inference workloads; Kubernetes-integrated prototype evaluated in real-world settingsOverall energy consumption of cloud-to-things resources vs provisioning/access latencyEELAS, an ILP-based scheduling platform with a lower-complexity heuristic for energy- and latency-aware allocation of ML inference workloads across the continuum, achieves significant energy gains in real-world settings with minimum possible latency from far-edge devices; no numbers in the abstract.No quantitative results in the abstract; heuristic-vs-ILP gap not quantifiedmoderateabstract-onlyEnergy-latency-aware ML scheduling spanning cloud to edge, extending datacenter scheduling evidence to the continuum.
takc2025datasurvey/reviewReview of power-system flexibility requirements, datacenter flexibility assets and operational capabilities, plus illustrative real-world case studies (UK-focused statistics)Qualitative assessment of datacenters' flexibility potential, benefits, and barriers (QoS/SLA, legislation, market structures)The review concludes datacenters have high potential to meet growing power-system flexibility needs (context statistics: datacenters use 1.7% of global electricity, ~2.5% of UK electricity, contribute ~1% of energy-related GHG, with consumption projected to roughly double to ~1000 TWh by 2026 vs 2022) — reported qualitatively.Review-level synthesis with no new measurements; quantitative claims are contextual statistics rather than experimental results.moderateabstract-onlyFraming source for datacenter demand-response/flexibility: catalogs flexibility assets and identifies QoS/SLA compliance and energy-market regulation as key barriers.
tschand2025mlperfbenchmark/measurement1,841 reproducible power measurements from 60 ML systems spanning microwatt to megawatt scales, using MLPerf benchmark workloads; methodology developed by a consortium from 20+ organizationsEnergy efficiency (performance per energy) and its trade-offs with performance and complexity across ML deployment scalesThe analysis reveals trade-offs between performance, complexity, and energy efficiency across the full range of ML system scales and provides rules/best practices for comparability — reported qualitatively in the abstract, without specific numerical results.Abstract reports no specific numbers (results are in the full paper); measurement comparability across heterogeneous systems remains inherently difficult.moderateabstract-onlyStandard-setting methodology for ML energy-efficiency benchmarking from microwatts to megawatts - key reference for defining reproducible power-efficiency metrics.
vanderbauwhede2026modellinganalytical-modelAnalytical scenarios of carbon-aware geographic load shifting of compute workloads (model explicitly ignores grid capacity, demand, and curtailment - i.e., optimistic)Emission reductions from geographic load shiftingEven under optimistic assumptions, realistic emission reductions from carbon-aware geographic load shifting are small, of the order of 5%, which is not enough to compensate for the growth in emissions from global data centre expansion.Deliberately optimistic model (ignores grid capacity/curtailment), so real reductions are expected to be smaller; scenario-level analysis rather than measurement.moderateabstract-onlyBounds expectations for carbon-aware spatial workload shifting - important counterweight to optimistic carbon-scheduling claims in the review.
wan2023safecoolsimulationData center cooling system simulation driven by a real-world workload trace; actor-critic model-based RL with transition and risk models, MPC and risk-guided explorationCooling power consumption; safety (constraint violation avoidance); sample efficiencySafeCool, a model-based RL cooling controller with MPC-based safety and risk-guided exploration, saves up to 13.18% cooling power compared with state-of-the-art MBRL data center cooling solutions in simulations using a real-world workload trace.Simulation-based evaluation; comparison limited to MBRL baselines; single workload tracemoderateabstract-onlyRL-based cooling optimization with safety guarantees, directly relevant to datacenter cooling energy reduction.
wang2025decoupledbenchmark/measurementLLM inference workloads decoupled into prefill and decode stages under varying workload conditionsEnergy efficiency and performance under DVFS, per inference stageCase-study measurements show DVFS improves energy efficiency in both prefill and decode stages with different sensitivities to workload variations, motivating stage-aware DVFS policies; no quantitative results in the abstract.No numbers in the abstract; case-study scope and hardware not specifiedmoderateabstract-onlyStage-aware DVFS for LLM serving - complements phase-aware power/energy studies of inference.
wiesner2026distributedsystem-design+evaluation561M-parameter (d20) transformer pretrained on 12.8B tokens across three geo-distributed GPU clusters (4x NVIDIA A100 each) using Flower; curtailment windows derived from WattTime marginal carbon intensity tracesTraining quality (perplexity), energy (kWh), and operational emissions under curtailment-aligned elastic schedulingCurtailment-aware scheduling preserves training quality (best train perplexity 14.6 vs 14.7 centralized) with operational emissions reduced to 5-12% of single-site baselines and 97% of energy consumed in curtailment windows, at comparable total energy (36.0-37.7 kWh across scenarios; 37.7 kWh for the elastic system).Preliminary feasibility study: small 561M model, prototype scale, and traces replayed rather than live grid signals; per-round overhead ~115 s (~80% compute utilization at 600 s rounds) slightly increases energy.moderatefull-textFeasibility evidence for pretraining LLMs inside renewable curtailment windows - a demand-following strategy that converts otherwise-wasted clean energy into compute.
wilhelm2025beyondanalytical-modelSmall vs large language models on the MMLU benchmark; test-time compute strategies (Chain-of-Thought, majority voting); transformer input-output token dynamicsEnergy-accuracy trade-off; proposed Energy-per-Token metric; reasoning-depth controlAnalyzes how test-time compute strategies let small models approach larger-model accuracy but add energy costs, proposing Energy-per-Token and related metrics plus energy-aware routing and controlled reasoning; results reported qualitatively (no numbers in the abstract).Position/metrics paper; no quantitative results in the abstractlowabstract-onlyProposes energy-per-token metrics for LLM inference, a conceptual foundation for inference energy benchmarking.
wilkins2024hybridanalytical-modelRepresentative LLM query dataset; hybrid cluster of energy-efficient processors and high-performance GPUs with cost-based workload-aware scheduling by input/output token countsCPU+GPU energy consumption of LLM inferenceA workload-aware hybrid-cluster strategy that routes queries by input/output token counts reduces CPU+GPU energy consumption by 7.5% compared to a workload-unaware baseline on a representative LLM dataset.Dataset-based analysis; single baseline; latency/SLO implications not addressed in the abstractmoderateabstract-onlyHardware heterogeneity plus token-aware routing yields modest but concrete LLM inference energy savings.
wilkins2024offlineanalytical-modelSeveral state-of-the-art LLMs characterized on heterogeneous GPU-CPU systems across different input prompt and output text magnitudesAccuracy (R2) of workload-dependent energy and runtime models; energy savings of offline energy-optimal scheduling in a case studyThe authors develop energy and runtime models per LLM with R2 > 0.96 and demonstrate through a case study that energy- and accuracy-aware scheduling beats existing best practices — the scheduling savings magnitude is not quantified in the abstract.Offline modeling only (no online adaptation); savings magnitude and case-study scale not reported in the abstract.moderateabstract-onlyWorkload-dependent energy/runtime models for LLM inference on heterogeneous systems - enables energy-optimal offline scheduling decisions.
williams2026powercase-study130 kW GPU cluster of 96 NVIDIA Blackwell Ultra GPUs at a Nebius London facility run for 5 days under 22 grid dispatch events (EPRI/National Grid); geo-shift between Oracle Ashburn VA and Chicago IL clusters (60 kW, 80 H100 each, Qwen2.5-32B via vLLM/Dynamo)Power-reduction response time and magnitude, dispatch compliance, TTFT, throughput, and cross-cluster load migrationThe cluster achieved 100% compliance across 200+ power targets including ~30% load reduction within 40 seconds and 40% within about a minute, sustained 10-40% curtailment for 2-10 hours with priority-job performance preserved, and a 375 W GPU power cap in Ashburn shifted load (3.1 kW increase in Chicago) with only a ~30 ms average TTFT increase and 10% of live inference traffic migrated VA-to-IL during a Dominion winter peak event.Industry-partner demonstration with controlled, replayed grid events rather than long-term operation; specific cluster sizes and configurations; TTFT degradation and workload-mix effects only partially characterized.highfull-textFlagship production-scale evidence that GPU clusters can act as grid-interactive assets (demand response, emergency curtailment, carbon-aware following, geo-shifting) via power capping and workload orchestration.
yu2023knowsystem-design+evaluationCloud-scale ML inference service with 105 servers containing three different kinds of GPUs serving five ML models, evaluated with real-world traces (prototype plus cloud-scale simulation)Cloud-scale energy consumption under hierarchical GPU management (cluster allocation, node scaling, GPU scaling, GPU clock scaling)The proposed hierarchical energy-aware GPU resource management saves up to 28.3% of cloud-scale energy when serving five ML models on 105 servers with three GPU types, and the paper finds that SLO-blind DVFS drivers on commercial GPUs maintain immoderately high clock frequencies.Evaluation combines prototype with cloud-scale simulation; savings are trace- and hardware-mix-specific; 'up to' headline figure.moderateabstract-onlyCharacterization + hierarchical DVFS/clock-scaling and allocation for GPU inference clusters - direct evidence for GPU frequency control as an energy lever at cloud scale.
zhang2026weightbenchmark/measurement270 measured configurations: 6 decoder-only model families (1.1B-9B), FP16/NF4/INT8 precisions, NVIDIA RTX 4090D/RTX 5090/A800 GPUs, 5 batch sizes, direct NVML power sampling (plus supplementary Tesla T4 runs)Quantized LLM inference energy vs FP16 baseline, as a function of model scale, precision, and GPU platformWeight-only quantization shows a model-size-dependent sign reversal under bitsandbytes: NF4 increases energy by approximately 25-45% for 1.1-1.5B models but saves about 23% for 6B-9B models, while INT8 shows ~33-55% small-model overhead and about 15% savings for larger models.Single quantization backend (bitsandbytes); INT8 results specific to the evaluated mixed-precision implementation; crossover values are operational reference bounds, not exact hardware constants; T4 runs are supplementary only.highabstract-onlyEmpirical caution for inference energy efficiency: quantization is not universally energy-saving; precision selection must be guided by model scale, hardware, and backend - important nuance for LLM serving energy claims.
zhao2023sustainablecase-studyGPUs at a research supercomputing center under power capping; GPU temperature and power draw measured with job performance impact assessed (first such analysis at supercomputing scale)GPU temperature, power draw, and job performance under power cappingPower capping GPUs yields significant decreases in both temperature and power draw with minimal impact on job performance, potentially improving hardware lifespan — reported qualitatively, with no numerical results in the abstract.No quantitative results in the abstract; single research center; performance impact described only as 'minimal' without bounds.moderateabstract-onlyEarly supercomputing-scale evidence that GPU power capping cuts power and temperature at low performance cost - supports power capping as a datacenter energy/cooling lever.
zhao2025internetsimulationInternet data center with idle UPS assets modeled for power backup and energy storage; QPSO-solved optimization; simulation studyTotal operating cost (power purchase cost + cost of calling UPS flexibility)Proposes a UPS flexibility model (power backup + energy storage) and a QPSO-based optimization for internet data center economic operation, verified through simulation including the impact of UPS flexibility capacity; no quantitative results in the abstract.Simulation-based; no numbers in the abstract; single-asset type (UPS) consideredmoderateabstract-onlyUPS/battery assets as grid demand-response flexibility for datacenter economic operation.

Swipe sideways to see all columns.

References

  1. Acun, Bilge (2023). Carbon Explorer: A Holistic Framework for Designing Carbon Aware Datacenters — Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. Abstract only. Selected for theme ?.doi:10.1145/3575693.3575754
  2. Adimora, Kyrian C. (2025). HARMONIC: Uncertainty-Aware Multi-Objective Optimization for Energy-Efficient HPC Resource Management — IEEE Transactions on Parallel and Distributed. Abstract only. Selected for theme ?.doi:10.1109/tpds.2025.3610354
  3. Afzal, Ayesha (2025). GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs — arXiv.org. Full text read. Selected for theme ?.doi:10.48550/arxiv.2510.06902
  4. Ali, Ghazanfar (2023). Performance-Aware Energy-Efficient GPU Frequency Selection using DNN-based Models — Proceedings of the 52nd International Conference on Parallel Processing. Abstract only. Selected for theme ?.doi:10.1145/3605573.3605600
  5. Amaral, Marcelo (2023). Kepler: A Framework to Calculate the Energy Consumption of Containerized Applications — 2023 IEEE 16th International Conference on Cloud Computing (CLOUD). Abstract only. Selected for theme ?.doi:10.1109/cloud60044.2023.00017
  6. Antici, Francesco (2023). PM100: A Job Power Consumption Dataset of a Large-scale Production HPC System — Proceedings of the SC '23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis. Abstract only. Selected for theme ?.doi:10.1145/3624062.3624263
  7. Bashir, Noman (2024). The Sunk Carbon Fallacy: Rethinking Carbon Footprint Metrics for Effective Carbon-Aware Scheduling — Proceedings of the ACM Symposium on Cloud Com. Abstract only. Selected for theme ?.doi:10.1145/3698038.3698542
  8. Beitsayadeh, Carl & Darbyshire, Pamayla E. (2026). Benchmarking AI Inference Efficiency in Public and Private Clouds: An MLPerf-Based Comparative Study — IEEE Transactions on Cloud Computing. Abstract only. Selected for theme ?.doi:10.1109/tcc.2026.3674888
  9. Bharadwaj, Srikant (2023). Predict; Don't React for Enabling Efficient Fine-Grain DVFS in GPUs — Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4. Abstract only. Selected for theme ?.doi:10.1145/3623278.3624756
  10. Bhavsar, Adyant (2026). Carbon-Aware Scheduling of AI Data Center Workloads Using Environmental and Energy Grid Forecasts — arXiv preprint. Abstract only. Selected for theme ?.doi:10.22541/essoar.177248651.16591616/v1
  11. Bika, Anil (2025). Powering Hyperscale AI Data Centers with Off-grid Solar, Battery, and Hydrogen — arXiv preprint. Abstract only. Selected for theme ?.doi:10.26434/chemrxiv-2025-1z3cp
  12. Brenner, Aron et al. (2025). Strategic Data Center Load Shifting : Implications for Market Efficiency and Transmission Value — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2510.20805
  13. Cao, Zhiwei et al. (2025). Transforming Future Data Center Operations and Management via Physical AI — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2504.04982
  14. Carastan‐Santos, Danilo (2025). Scheduling With Lightweight Predictions in Power-Constrained HPC Platforms — IEEE Transactions on Parallel and Distributed. Abstract only. Selected for theme ?.doi:10.1109/tpds.2025.3586723
  15. Chien, Andrew A. (2023). Reducing the Carbon Impact of Generative AI Inference (today and in 2035) — Proceedings of the 2nd Workshop on Sustainable Computer Systems. Abstract only. Selected for theme ?.doi:10.1145/3604930.3605705
  16. Clark, Quentin (2024). Learning a Data Center Model for Efficient Demand Response — ACM SIGEnergy Energy Informatics Review. Abstract only. Selected for theme ?.doi:10.1145/3727200.3727215
  17. Crozier, Constance (2025). The Potential of Data Center Energy Demand To Provide Grid Flexibility — Current Sustainable/Renewable Energy Reports. Abstract only. Selected for theme ?.doi:10.1007/s40518-025-00258-9
  18. D'Amico, Marco (2024). Energy hardware and workload aware job scheduling towards interconnected HPC environments — IEEE Transactions on Parallel and Distributed. Abstract only. Selected for theme ?.doi:10.1109/tpds.2021.3090334
  19. Das, Abhijit (2026). Follow the Curtailment: Data Center Location as a Grid Coordination Strategy — arXiv preprint. Abstract only. Selected for theme ?.doi:10.2139/ssrn.6823640
  20. Desai, T. (2025). Adaptive GPU Power Capping: Balancing Energy Efficiency,Thermal Control and Performance — IEEE International Symposium on High-Performa. Abstract only. Selected for theme ?.doi:10.1145/3731545.3735119
  21. Desislavov, Radosvet (2023). Trends in AI inference energy consumption: Beyond the performance-vs-parameter laws of deep learning — Sustainable Computing Informatics and Systems. Abstract only. Selected for theme ?.doi:10.1016/j.suscom.2023.100857
  22. Dokuchaev, Ivan (2025). ELTO: Energy-Latency Trade-off Optimization for Machine Learning Inference with Dynamic Batching — 2025 23rd International Symposium on Network. Abstract only. Selected for theme ?.doi:10.1109/nca67271.2025.00035
  23. Doshi, Nishi & Shah, Shrey (2026). Night-Window Batching versus Carbon - Aware Scheduling for Clinical AI GPU Workloads — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2606.01766
  24. dos Santos Gonçalves, Thiago (2026). Investigating the Impact of DVFS on the Energy Efficiency of AI Workloads on GPUs — Communications in Computer and Information Sc. Abstract only. Selected for theme ?.doi:10.1007/978-3-032-24923-4_11
  25. Du, Bojun et al. (2026). From Tokens to Energy Flexibility: Quantization-Enabled Demand Response for Data Centers with LLM Inference Workloads — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2606.18851
  26. Dunlap, Chris (2026). Quantifying AI data center flexibility as a resource adequacy asset — arXiv preprint. Abstract only. Selected for theme ?.doi:10.21203/rs.3.rs-9829457/v1
  27. Fan, Haoyang (2025). ELLIE: Energy-Efficient LLM Inference at the Edge Via Prefill-Decode Splitting — 2025 IEEE 36th International Conference on Ap. Abstract only. Selected for theme ?.doi:10.1109/asap65064.2025.00031
  28. Ferdaus, Farah et al. (2025). Evaluating Energy Efficiency of Ai Accelerators Using Two Mlperf Benchmarks — 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Abstract only. Selected for theme ?.doi:10.1109/ccgrid64434.2025.00035
  29. Gheni, Mashhur (2026). Operational analysis of the cooling system in a direct liquid-cooled data center: a measurement and simulation study on the impact of supply water temperature — Applied Energy. Abstract only. Selected for theme ?.doi:10.1016/j.apenergy.2025.127061
  30. Gnibga, Wedan Emmanuel & Chien, Andrew A. (2026). Improving Datacenter IT Capacity, Cooling Efficiency and Lifetime with Underground Thermal Energy Storage Systems — Proceedings of the 17th ACM International Conference on Future and Sustainable Energy Systems. Abstract only. Selected for theme ?.doi:10.1145/3744255.3798111
  31. Gonçalves, Raphael Hendrigo de Souza & Santos, Wendel Marcos dos (2026). A Scalable Digital Twin Framework for Energy Optimization in Data Centers — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2605.05581
  32. Guan, Xiaoding et al. (2024). WattScope: Non-intrusive Application-level Power Disaggregation in Datacenters — ACM SIGMETRICS Performance Evaluation Review. Abstract only. Selected for theme ?.doi:10.1145/3649477.3649491
  33. Hisaharo, Soka (2024). Optimizing LLM Inference Clusters for Enhanced Performance and Energy Efficiency — arXiv preprint. Abstract only. Selected for theme ?.doi:10.36227/techrxiv.172348951.12175366/v1
  34. Huang, Huikang (2026). Dynamic Power Capping for Latency-Sensitive Cloud Applications: A Prediction-Based Power Efficiency Approach — IEEE Transactions on Computers. Abstract only. Selected for theme ?.doi:10.1109/tc.2025.3647020
  35. HUR, JIN (2026). Seasonal–Hourly Assessment of Data Center Spatial Flexibility under Renewable Variability: A Case Study of Jeju Island — arXiv preprint. Abstract only. Selected for theme ?.doi:10.2139/ssrn.6753967
  36. Jacquet, Pierre et al. (2026). Untangling GPU Power Consumption: Job-Level Inference in Cloud Shared Settings — Proceedings of the 21st European Conference on Computer Systems. Abstract only. Selected for theme ?.doi:10.1145/3767295.3769333
  37. Jain, Rutwik (2026). Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters — Proceedings of the ACM on Measurement and Ana. Abstract only. Selected for theme ?.doi:10.1145/3805644
  38. Jay, Mathilde (2023). An experimental comparison of software-based power meters: focus on CPU and GPU — 2023 IEEE/ACM 23rd International Symposium on. Abstract only. Selected for theme ?.doi:10.1109/ccgrid57682.2023.00020
  39. Jayatilleka, Sarath (2026). Reliability Modeling of Liquid Cooling Systems for Data Center Availability — 2026 Annual Reliability and Maintainability S. Abstract only. Selected for theme ?.doi:10.1109/rams50514.2026.11424468
  40. Jegham, Nidhal (2025). How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference — arXiv (Cornell University). Full text read. Selected for theme ?.doi:10.48550/arxiv.2505.09598
  41. Jiang, Zhuolong (2025). HeShare: Energy-Aware and Efficient Multi-Task GPU Sharing in Heterogeneous GPU-Based Computing Systems — IEEE Transactions on Computers. Abstract only. Selected for theme ?.doi:10.1109/tc.2025.3628924
  42. Kakolyris, Andreas Kosmas (2025). throttLL’eM: Predictive GPU Throttling for Energy Efficient LLM Inference Serving — 2025 IEEE International Symposium on High Per. Abstract only. Selected for theme ?.doi:10.1109/hpca61900.2025.00103
  43. Krishnaram, Prasanna Chandran Melnatami (2026). Feeding the GPUs: Node-Level Energy Modeling and Workload-Aware Efficiency in AI Inference Servers — 2026 IEEE Green Technologies Conference (GreenTech). Abstract only. Selected for theme ?.doi:10.1109/greentech68285.2026.11471616
  44. Lechowicz, Adam (2025). Carbon- and Precedence-Aware Scheduling for Data Processing Clusters — Proceedings of the ACM SIGCOMM 2025 Conferenc. Abstract only. Selected for theme ?.doi:10.1145/3718958.3750478
  45. Li, Baolin (2023). Toward Sustainable HPC: Carbon Footprint Estimation and Environmental Implications of HPC Systems — Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. Abstract only. Selected for theme ?.doi:10.1145/3581784.3607035
  46. Liu, Di (2023). Efficient GPU Resource Management under Latency and Power Constraints for Deep Learning Inference — 2023 IEEE 20th International Conference on Mobile Ad Hoc and Smart Systems (MASS). Abstract only. Selected for theme ?.doi:10.1109/mass58611.2023.00074
  47. Liu, Junhong et al. (2025). Synergising Hierarchical Data Centers and Power Networks: A Privacy-Preserving Approach — arXiv preprint. Abstract only. Selected for theme ?.doi:10.48550/arxiv.2505.20575
  48. Liu, Qiong et al. (2026). Mixture-of-experts based multi-critic deep reinforcement learning for sustainable management of data center microgrids — Applied Energy. Abstract only. Selected for theme ?.doi:10.1016/j.apenergy.2026.127561
  49. liu, xingliang (2026). A public-data decision-support framework for carbon-aware workload allocation in AI data-center hubs: evidence from China — arXiv preprint. Abstract only. Selected for theme ?.doi:10.2139/ssrn.7052138
  50. Lu, Rui (2025). A Thermal-Aware Workload Scheduler for High-Performance LLM Inference in Cooling-Regulated Datacenters — ACM SIGEnergy Energy Informatics Review. Abstract only. Selected for theme ?.doi:10.1145/3757892.3757906
  51. Luccioni, Alexandra Sasha (2023). Power Hungry Processing: Watts Driving the Cost of AI Deployment? — arXiv (Cornell University). Full text read. Selected for theme ?.doi:10.48550/arxiv.2311.16863
  52. Luo, Qi et al. (2026). Request-Level Energy Attribution for Batched LLM Serving — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2608.00026
  53. Ma, Yuan (2025). Power Capping of GPU Servers for Machine Learning Inference Optimization — International Conference on Parallel Processi. Abstract only. Selected for theme ?.doi:10.1145/3754598.3754670
  54. Ma, Bole (2026). The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures — arXiv.org. Full text read. Selected for theme ?.doi:10.48550/arxiv.2605.11999
  55. Maji, Diptyaroop (2024). Untangling Carbon-free Energy Attribution and Carbon Intensity Estimation for Carbon-aware Computing — The 15th ACM International Conference on Futu. Abstract only. Selected for theme ?.doi:10.1145/3632775.3662164
  56. Mavromatis, Ioannis (2023). FROST: Towards Energy-efficient AI-on-5G Platforms – A GPU Power Capping Evaluation — 2023 IEEE Conference on Standards for Communications and Networking (CSCN). Abstract only. Selected for theme ?.doi:10.1109/cscn60443.2023.10453214
  57. Menear, Kevin (2025). Energy-Aware HPC Scheduling with LLM-Based Power Prediction — Proceedings of the SC '25 Workshops of the In. Abstract only. Selected for theme ?.doi:10.1145/3731599.3767563
  58. Menear, Kevin (2026). Quantifying the Economic Potential of Energy-Aware Scheduling in HPC — Proceedings of the Supercomputing Asia and In. Abstract only. Selected for theme ?.doi:10.1145/3784828.3785162
  59. Naug, Avisek et al. (2025). LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers — arXiv preprint. Abstract only. Selected for theme ?.doi:10.48550/arxiv.2511.00116
  60. Niu, Chenxu (2025). Energy Efficient or Exhaustive? Benchmarking Power Consumption of LLM Inference Engines — ACM SIGEnergy Energy Informatics Review. Abstract only. Selected for theme ?.doi:10.1145/3757892.3757900
  61. Patel, Pratyush (2023). Splitwise: Efficient generative LLM inference using phase splitting — arXiv (Cornell University). Full text read. Selected for theme ?.doi:10.48550/arxiv.2311.18677
  62. Piolant, Albert d'Aviau de (2025). Improving energy efficiency of HPC applications using unbalanced GPU power capping — IEEE International Symposium on Parallel & Di. Abstract only. Selected for theme ?.doi:10.1109/ipdpsw66978.2025.00132
  63. Purnachandra Reddy, Gokul Chandra (2026). CARGO: A Carbon-Aware, Reinforcement-Guided Orchestration Framework for GPU Workload Scheduling Across Geo-Distributed Cloud Regions — arXiv preprint. Abstract only. Selected for theme ?.doi:10.2139/ssrn.6774820
  64. Qi, Minxuan et al. (2026). Safety Constraint Multi-Agent Deep Reinforcement Learning for Hybrid Air-Liquid Cooling System in Data Centers — arXiv preprint. Abstract only. Selected for theme ?.doi:10.2139/ssrn.6864511
  65. Qiao, Feitong (2024). Energy-Aware Process Scheduling in Linux — ACM SIGEnergy Energy Informatics Review. Abstract only. Selected for theme ?.doi:10.1145/3727200.3727214
  66. Radman, Gamil (2026). Impact of Dynamic Voltage on GPU Energy Consumption for Real-Time Systems — Research Square. Abstract only. Selected for theme ?.doi:10.21203/rs.3.rs-9423012/v1
  67. Repeva, Marina (2025). Anomaly Detection In The Performance Of Data Center Cooling System Devices Based On Machine Learning And Time Series Analysis — 2025 IEEE Ural-Siberian Conference on Biomedi. Abstract only. Selected for theme ?.doi:10.1109/usbereit65494.2025.11054137
  68. Rrapaj, Ermal (2024). Power Consumption Trends in Supercomputers: A Study of NERSC's Cori and Perlmutter Machines — ISC High Performance 2024 Research Paper Proceedings (39th International Conference). Abstract only. Selected for theme ?.doi:10.23919/isc.2024.10528943
  69. Samsi, Siddharth (2023). From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference — arXiv (Cornell University). Full text read. Selected for theme ?.doi:10.48550/arxiv.2310.03003
  70. Shen, Jiachen et al. (2026). Grid-Interactive Thermal Management of AI Data Centers via Contextual Distributionally Robust Optimization — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2607.00099
  71. Shuai, Huang (2024). Data Center Energy Consumption Simulation and Prediction from the Perspective of Digital Twin — 2024 China International Conference on Electr. Abstract only. Selected for theme ?.doi:10.1109/ciced63421.2024.10754439
  72. Souza, Abel (2023). CASPER: Carbon-Aware Scheduling and Provisioning for Distributed Web Services — Proceedings of the 14th International Green and Sustainable Computing Conference. Abstract only. Selected for theme ?.doi:10.1145/3634769.3634812
  73. Spaan, Jeffrey (2026). Reducing Compute Waste in LLMs through Kernel-Level DVFS — arXiv.org. Full text read. Selected for theme ?.doi:10.48550/arxiv.2601.08539
  74. Springborg, Anders Aaen (2023). Automatic Energy-Efficient Job Scheduling in HPC: A Novel SLURM Plugin Approach — Proceedings of the SC '23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis. Abstract only. Selected for theme ?.doi:10.1145/3624062.3624265
  75. Stojkovic, Jovan (2024). Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference — arXiv (Cornell University). Full text read. Selected for theme ?.doi:10.48550/arxiv.2403.20306
  76. Stojkovic, Jovan (2025). DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency — 2025 IEEE International Symposium on High Per. Full text read. Selected for theme ?.doi:10.1109/hpca61900.2025.00102
  77. Suarez, Estela (2025). Energy-aware operation of HPC systems in Germany — Frontiers in High Performance Computing. Abstract only. Selected for theme ?.doi:10.3389/fhpcp.2025.1520207
  78. Sun, Yimeng (2025). Learning-Enabled Adaptive Power Capping Scheme for Cloud Data Centers — IEEE Transactions on Smart Grid. Abstract only. Selected for theme ?.doi:10.1109/tsg.2025.3598070
  79. Syrigos, Ilias (2023). EELAS: Energy Efficient and Latency Aware Scheduling of Cloud-Native ML Workloads — 2023 15th International Conference on COMmunication Systems &amp; NETworkS (COMSNETS). Abstract only. Selected for theme ?.doi:10.1109/comsnets56262.2023.10041344
  80. Takcı, Mehmet Türker (2025). Data centres as a source of flexibility for power systems — Energy Reports. Abstract only. Selected for theme ?.doi:10.1016/j.egyr.2025.03.020
  81. Tschand, Arya et al. (2025). MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from μWatts to MWatts for Sustainable AI — 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). Abstract only. Selected for theme ?.doi:10.1109/hpca61900.2025.00092
  82. Vanderbauwhede, Wim (2026). Modelling Scenarios for Carbon-Aware Geographic Load Shifting of Compute Workloads — IEEE Transactions on Sustainable Computing. Abstract only. Selected for theme ?.doi:10.1109/tsusc.2026.3702249
  83. Wan, Jianxiong (2023). SafeCool: Safe and Energy-Efficient Cooling Management in Data Centers With Model-Based Reinforcement Learning — IEEE Transactions on Emerging Topics in Compu. Abstract only. Selected for theme ?.doi:10.1109/tetci.2023.3234545
  84. Wang, Lei (2025). Decoupled Analysis of DVFS Effects in Prefill and Decode Stages of Large Language Model Inference — 2025 IEEE 9th Conference on Energy Internet a. Abstract only. Selected for theme ?.doi:10.1109/ei268505.2025.11425606
  85. Wiesner, Philipp et al. (2026). Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2602.22760
  86. Wilhelm, Patrick (2025). Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference — Proceedings of the 5th Workshop on Machine Le. Abstract only. Selected for theme ?.doi:10.1145/3721146.3721953
  87. Wilkins, Grant (2024). Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads — The 15th ACM International Conference on Futu. Abstract only. Selected for theme ?.doi:10.1145/3632775.3662830
  88. Wilkins, Grant (2024). Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems — ACM SIGEnergy Energy Informatics Review. Abstract only. Selected for theme ?.doi:10.1145/3727200.3727217
  89. Williams, Chris et al. (2026). Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive Compute — arXiv preprint. Full text read. Selected for theme ?.doi:10.48550/arxiv.2606.25098
  90. Yu, Junyeol (2023). Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning Serving — 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Abstract only. Selected for theme ?.doi:10.1109/hpca56546.2023.10070943
  91. Zhang, Hongping (2026). Weight-Only Quantization Does Not Always Save Energy: An Empirical Study of LLM Inference Across NVIDIA GPU Platforms — arXiv preprint. Abstract only. Selected for theme ?.doi:10.2139/ssrn.6854700
  92. Zhao, Dan (2023). Sustainable Supercomputing for AI — ACM Symposium on Cloud Computing. Abstract only. Selected for theme ?.doi:10.1145/3620678.3624793
  93. Zhao, Yuxin (2025). Internet Data Center Economic Operation Considering UPS Flexibility — 2025 8th International Conference on Energy,. Abstract only. Selected for theme ?.doi:10.1109/ceepe64987.2025.11033969