The environmental impact of AI, from chips to prompts

The Environmental Impact of Training Models Like ChatGPT
The Environmental Impact of Training Models Like ChatGPT

Artificial intelligence (AI) runs on physical infrastructure. Training and serving models require electricity, cooling water, chips, data centers, and grid capacity, while hardware production adds emissions and material demand. The size of that footprint varies sharply by model, workload, location, and measurement boundary. Here is what current evidence can establish, what remains unknown, and how technical teams can reduce impact without relying on misleading per-prompt comparisons.

The short answer

AI has a real and growing environmental footprint. Its direct effects include electricity demand, greenhouse gas emissions, water consumption, hardware production, land use, and electronic waste. Its indirect effects can be positive or negative, depending on what an AI application replaces and whether greater efficiency drives more total use.

There is no single defensible number for “the footprint of AI.” A short text completion from a small model and a long multimodal agent run are different workloads. A kilowatt-hour from a low-carbon grid has a different carbon footprint from the same energy on a fossil-heavy grid. Water demand also changes with cooling design, weather, power generation, and local scarcity.

The most credible assessment therefore follows the full system over its life cycle and reports a functional unit, such as energy per 1,000 completed tasks, alongside model quality.

What counts as AI's environmental impact?

The visible prompt is only the last step in a physical supply chain. The UN Environment Programme recommends considering the full life cycle, from raw-material extraction and chip manufacturing through model development, deployment, and end-of-life equipment.

Life-cycle stageMain environmental effectsCommonly missed costs
Chips and data-center constructionEmbodied carbon, minerals, manufacturing water, landServers, networking gear, concrete, steel, replacement cycles
Data preparation and experimentationElectricity, cooling, storageFailed runs, hyperparameter searches, synthetic-data generation
Pretraining and fine-tuningConcentrated electricity and water demandMultiple candidate models and repeated fine-tunes
Inference and servingRecurring electricity, cooling, network trafficIdle capacity, long context, retries, tool calls, monitoring
End of lifeElectronic waste and lost materialsAccelerators retired before their technical end of life
Effects of the applicationAvoided or added energy, emissions, travel, and consumptionRebound effects and a weak comparison baseline

Four distinctions keep the discussion accurate:

  • Power and energy differ. Power, measured in watts, describes demand at a moment. Energy, measured in watt-hours, describes use over time. Both matter because a data center can strain local grid capacity even when its annual national share looks modest.
  • Energy and carbon differ. Electricity use becomes an emissions estimate only after applying a grid carbon intensity and a defined accounting method.
  • Water withdrawal and consumption differ. Withdrawal is water taken from a source. Consumption is the portion unavailable for immediate return, often because it evaporates. Data centers can use water on-site for cooling and indirectly through electricity generation.
  • Operational and embodied impacts differ. Electricity and cooling occur during use. Mining, manufacturing, construction, transport, maintenance, and disposal sit elsewhere in the life cycle.

How much electricity does AI use?

The strongest top-down figures describe data centers as a whole. They include AI, cloud services, storage, streaming, and other workloads. Treating the total as “AI electricity use” overstates what the evidence can show.

The International Energy Agency's 2026 data-center outlook estimates 485 terawatt-hours (TWh) of global data-center electricity consumption in 2025 and projects about 950 TWh in 2030, close to 3% of global electricity demand. Within that total, electricity use at AI-focused data centers is projected to triple between 2025 and 2030. The projection is a scenario, and the IEA identifies financing, chip supply, grid connections, and efficiency gains as major uncertainties.

The US picture shows why location matters. A Berkeley Lab report estimates that US data centers used 176 TWh in 2023, or 4.4% of US electricity. Its 2028 scenarios range from 325 to 580 TWh, equal to 6.7% to 12% of projected US electricity use. These figures cover all data-center workloads and reflect wide uncertainty about graphics processing unit (GPU) shipments, utilization, and cooling.

Aggregate percentages also hide local effects. Data centers cluster around available fiber, land, power, and water. A new cluster can create a large, continuous load on one utility territory even when the global percentage remains small.

Training gets attention, but inference can dominate over time

Pretraining a foundation model concentrates a large amount of computation into a bounded project. Inference is each subsequent use of the trained model. Fine-tuning, evaluation, safety testing, failed experiments, and model updates sit between those phases.

A US Government Accountability Office review illustrates the spread in published training estimates. It lists about 1,287 megawatt-hours (MWh) and 552 metric tons of carbon dioxide equivalent (tCO2e) for GPT-3, compared with 21,588 MWh and 8,930 tCO2e for Llama 3.1 405B. These are different models trained years apart, and their reporting boundaries are not identical. They show scale and variability rather than a fair efficiency ranking.

Training is only part of a deployed model's lifetime footprint. A model used rarely may never accumulate inference energy comparable with training. A model serving millions of long responses can. The crossover depends on request volume, input and output length, batching, quantization, hardware, utilization, and how long the model stays in service.

This is also why GPT-3 training estimates do not describe ChatGPT today. ChatGPT is a service that can route work across models and tools. OpenAI does not publish the complete energy, water, hardware, and location data needed to calculate a universal ChatGPT footprint.

Why “energy per AI prompt” is a fragile number

Per-prompt figures can help when the workload and boundary are explicit. They become misleading when repeated as a property of all AI.

Google measured a median text prompt in its Gemini Apps at 0.24 watt-hours (Wh), 0.03 grams of CO2e, and 0.26 milliliters of water in May 2025. Its technical measurement paper includes active accelerators, host systems, idle capacity, and facility overhead. The carbon and water results use fleet-wide averages. The finding applies to Google's median Gemini text prompt under that methodology. It does not establish the footprint of ChatGPT, an image, a video, a large reasoning task, or an agent that performs many model and tool calls.

A useful per-request disclosure should state:

  • the model and hardware version;
  • input, output, and cached-token counts;
  • modality, batch size, and latency target;
  • accelerator, host, memory, network, and idle energy;
  • power usage effectiveness (PUE), which adds facility overhead;
  • grid location, time, and carbon-accounting method;
  • direct and electricity-related water boundaries;
  • the date and distribution, including median and high-percentile requests.

Without those details, comparisons create false precision.

AI's carbon footprint depends on where and when it runs

Two identical jobs can consume the same electricity and produce different emissions. The result depends on the generators serving the data center at that time. Carbon-free power contracts and annual renewable-energy matching may improve a company's accounting position, while the local grid can still supply fossil generation during some hours.

The IEA's emissions analysis estimates that electricity used by all data centers caused about 180 million metric tons of indirect carbon dioxide emissions in 2024, around 0.5% of global fuel-combustion emissions. AI is a subset of that total. The IEA also expects data-center emissions to rise through this decade in its central scenario even as clean electricity and efficiency improve, because demand is growing quickly.

Carbon accounting should therefore show both energy and emissions. For location-based reporting, multiply measured energy by the relevant grid emissions factor. For market-based reporting, explain contractual clean-energy instruments separately. Neither method captures chips, buildings, backup generators, or end-of-life equipment unless the boundary adds them.

AI's water footprint is local and easy to misstate

Water enters the system in two main places. A data center may consume water directly through evaporative cooling. Power plants can also consume water while producing the electricity used by the facility. Cooling designs that reduce on-site water can require more electricity, so optimizing one metric alone can shift the burden elsewhere.

The Berkeley Lab estimate finds that US data centers consumed 66 billion liters of water directly in 2023. Their indirect water footprint from electricity generation was nearly 800 billion liters. Again, these are data-center totals, not AI-only figures. National totals also say little about risk in a water-stressed watershed, where one facility's timing and source can matter more than its share of US consumption.

The widely repeated “bottle of water” claim comes from a modeled scenario, not a meter attached to each ChatGPT response. The study summarized by GAO estimated 0.5 liters for roughly 10 to 50 queries under particular model, data-center, grid, location, and time assumptions. Google's 0.26 milliliters per median Gemini text prompt covers a different system and water boundary. Both figures are conditional. They cannot be swapped into claims about every AI request.

Good water reporting identifies the watershed, cooling method, water source, seasonal consumption, water stress, and both site and source water usage effectiveness (WUE).

Hardware, construction, and e-waste belong in the calculation

Accelerators do not appear without environmental cost. Semiconductor fabrication uses energy, water, chemicals, and raw materials. Data centers add steel, concrete, power equipment, backup systems, and networking hardware. Short replacement cycles can increase embodied emissions and electronic waste even when each new accelerator performs more work per watt.

These impacts remain underreported. GAO found insufficient information for a complete assessment of generative AI infrastructure construction and end-of-life equipment. That gap is a reason to label boundaries clearly, rather than treating operational electricity as a full life-cycle result.

The practical metric is an allocated share of embodied impact per completed unit of useful work. It should use the actual hardware lifetime and utilization, then include reuse, refurbishment, and recycling outcomes. A more efficient chip can still increase total impact when lower cost drives enough additional demand, an effect known as rebound.

AI can reduce environmental impact, but benefits need a baseline

AI can help detect methane leaks, forecast renewable generation, optimize industrial processes, route vehicles, and control building systems. Those benefits are application effects. They do not erase the infrastructure footprint automatically.

An IEA adoption scenario estimates that widespread use of existing AI applications could reduce 1.4 billion metric tons of carbon dioxide emissions in 2035. The IEA also says current momentum is insufficient to deliver that scenario and warns that rebound effects could negate some gains. The net result depends on deployment, incentives, and what happens without the AI system.

A credible benefit claim compares two defined scenarios:

  1. Measure the full process without AI.
  2. Measure the same outcome with AI, including compute and infrastructure.
  3. Hold service quality and demand assumptions constant.
  4. Test for additional demand caused by lower cost or greater convenience.
  5. Report uncertainty and the time period.

How technical teams can reduce AI's footprint

The largest levers usually sit in product and architecture decisions, well before anyone buys offsets.

  1. Use the least resource-intensive method that meets the requirement. Rules, search, a small classifier, or conventional software may solve some tasks. For generative work, route simple requests to smaller models and reserve larger models for cases that need them.
  2. Measure quality and resources together. Energy per request is incomplete if a low-energy model fails and triggers retries or human rework. Track energy, latency, completion rate, and task quality for the same test set.
  3. Reduce unnecessary computation. Cap output length, trim context, cache stable results, batch compatible requests, prevent retry loops, and stop tool-using agents when the goal is complete. Quantization and efficient serving engines can reduce inference cost when quality remains acceptable.
  4. Avoid repeating training work. Start from an existing model where appropriate. Parameter-efficient fine-tuning, retrieval, and better data can replace a full training run for many applications. Record experiments so failed configurations are not rerun.
  5. Raise hardware utilization. Idle accelerators still draw power and carry embodied impact. Autoscale carefully, match accelerators to workload size, and consolidate low-utilization services without breaking latency or reliability requirements.
  6. Choose location and time with carbon and water in view. Favor lower-carbon electricity and avoid water-stressed locations or seasons where workload constraints allow. Review PUE and WUE together because cooling choices create tradeoffs.
  7. Require vendor-level evidence. Ask for energy per defined workload, PUE, direct and indirect water, location-based emissions, hardware lifetime, and system boundaries. Annual corporate totals cannot substitute for product-level measurements.

For production voice AI, “one prompt” is an especially weak unit. A live conversation can include telephony and media transport, speech recognition, turn detection, a large language model (LLM), retrieval and tools, speech synthesis, monitoring, and storage. Our voice AI stack guide explains those layers. Measure the whole path per completed conversation minute or resolved task, and count failed calls and retries.

The International Telecommunication Union guidance formalizes the same basic discipline: define the functional unit, perform a life-cycle assessment, and evaluate first-order footprint alongside the effects created by using the system.

Frequently asked questions

Is one ChatGPT question bad for the environment?

Every request has some physical footprint, but there is no universal figure for one ChatGPT question. Model routing, tokens, tools, modality, hardware, utilization, data-center overhead, grid mix, and cooling all change the result. The more useful question is the measured footprint per completed task across a representative workload.

Does AI training or inference use more energy?

It depends on lifetime use. Training is a large, bounded event. Inference repeats for every request and can eventually exceed training for a heavily used service. Include experimentation, fine-tuning, evaluations, and model updates in the training side of the ledger.

Does renewable energy solve AI's environmental impact?

Lower-carbon electricity can reduce operational emissions substantially. It does not remove water, chip manufacturing, construction, mining, backup power, or e-waste impacts. Annual renewable matching also differs from receiving carbon-free electricity in the same location and hour as the workload.

Which metric should an AI team track first?

Start with energy per useful outcome, such as Wh per 1,000 successful tasks, and report quality beside it. Then convert energy to location-based CO2e, add direct and indirect water, and allocate embodied hardware and infrastructure impacts. A transparent boundary is more useful than a precise-looking number with hidden exclusions.

Environmental performance is an architecture and operations question. If you are building a production conversational AI product, evaluate Dasha as a managed runtime with APIs, telephony, integrations, testing, monitoring, and large-scale call execution.

Share

Subscribe

Sign up to our e-mail list to get the best of the Dasha blog sent directly to your inbox.

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.