Section 6 of 8
6. Flexible and Curtailable Load — Potential and Implementation Challenges#
The issue. Every other section describes a constraint; this one describes a source of relief. Because a large computational load can often give up a small share of its annual consumption — a quarter to a half of one percent — without harming its core function, curtailable service could let the existing grid absorb far more demand than a firm-service forecast implies — a modelled 76 to 98 GW nationwide, under defined participation and curtailment assumptions. Machinery blocks it rather than physics or economics: the industry has not yet built the contractual, operational, and compensation terms needed to deploy flexibility at scale, and a lever of that size sits idle until it does. The grid carries sizing for a peak that occurs a few dozen hours a year. Duke University's Nicholas Institute quantified the implication: across the 22 largest balancing authorities, covering about 95% of U.S. load, roughly 76 GW of new load could be absorbed on the existing system at a curtailment rate of 0.25%, 98 GW at 0.5%, and 126 GW at 1%. Those percentages are shares of energy, not of time — the study defines the rate as annual curtailed megawatt-hours divided by the new load’s maximum potential annual consumption — and the distinction matters to anyone signing the contract. Some curtailment occurs in 85 hours a year at the 0.25% rate, 177 at 0.5%, and 366 at 1%. Events average 1.7, 2.1 and 2.5 hours respectively, and most leave the load largely running: across the hours in which any curtailment is called, 88% retain at least half the new load, 60% at least three-quarters, and 29% at least nine-tenths. A 1–2% reduction in data-center peak demand tracks a 0.5–2.8% reduction in electricity rates.

Figure 5 — Curtailment-enabled headroom: additional load the existing network could serve if that load accepts a limited number of curtailment hours each year. The quantity runs large because networks carry sizing for a peak that occurs in a few dozen hours, so a modest, well-timed reduction releases capacity out of proportion to its size. A planner should conclude that flexibility operates on the load’s own timescale, where new generation and transmission do not. Three qualifications attach. The figure states a modelled result; it depends on the participation and curtailment assumptions given in Section 6; and it forms an upper bound in one specific respect, because the study does not model network constraints and lists them first among the factors that would reduce it.
The evidence base, and what it agrees on#
The **Duke** estimate is the most cited figure in this debate and should not carry the argument alone, so it is worth setting beside the other work. **EPRI’s DCFlex** initiative — a consortium of hyperscalers, utilities, grid operators and equipment suppliers — reports from its pilots that 10 to 40 percent load modulation is achievable at a data center without breaching customer service-level agreements, which is a facility-level measurement rather than a system-level model. **ICF**’s analysis anchors on roughly 80 hours of curtailment a year for flexible loads while noting a tail past 300 to 400 hours in extreme conditions. A 2025 **GridLab and Telos** Energy study of NV Energy found that if utilities account for flexible data center operations during long-term resource planning, rather than after investment decisions have been made, they can significantly reduce system costs. By modeling 1–2 GW of data center load that can be temporarily curtailed or shifted during periods of peak demand, the study estimated approximately $300–400 million in net present value savings over 2025–2050. Most of these savings come from avoiding or delaying investments in new generation, transmission, and other grid infrastructure, while requiring data center flexibility for only a limited number of hours each year.
Four separate studies, measuring four different things. The **Duke** estimate is a system figure: how much additional load the existing network could carry if that load accepts a limited number of curtailment hours a year. **EPRI’s DCFlex** pilots measure something narrower — the share of its own draw a single facility can shed without breaching its service-level agreements. **ICF** estimates the hours such a commitment would require in practice, anchoring near 80 hours in a normal year with a tail past 300 to 400 hours in extreme conditions. **GridLab and Telos** convert the idea into money, estimating the system-cost savings that follow when flexibility is planned for in advance rather than treated as an afterthought. Because each answers a different question — network headroom, facility capability, hours required, and dollars saved — the four results are not additive, and a reader who stacks them will double-count.
The four agree on direction and rough order of magnitude: a modest, well-timed reduction in demand releases capacity out of proportion to its size, because the grid is built for a peak that occurs only a few dozen hours a year. They diverge on the tail — the number of curtailment hours a flexible load might face in a stressed year — and a developer’s counsel reads the tail most closely. A figure near 85 to 177 hours of mostly partial reduction describes the average year, and a developer can plan around it; an obligation that is effectively unbounded in a bad year is a different contract entirely, and the service terms have to resolve the distance between those two readings.
Who directs the curtailment, and what that decides#
Two different things travel under the word flexibility, and the studies above measure the first while most of the industry’s own examples describe the second. In the first, the operator of the system calls the event: a dispatch instruction, a tariff trigger, a contractual notice, arriving when the system is short and not otherwise. In the second, the operator of the load schedules around the grid on its own initiative — deferring a training run, moving a batch job to a region with headroom, shifting work off the local peak. The second costs far less. Training and batch inference checkpoint cleanly and tolerate delay, the software to orchestrate them exists and deploys in weeks rather than permitting cycles, and the load surrenders throughput on its own schedule rather than someone else’s. It also gives the system almost nothing it can plan against. A planner can count a resource only if it is obligated rather than voluntary, verifiable after the fact, available at the system’s peak rather than at the load’s convenience, and deliverable at the point where the constraint binds. Self-directed deferral satisfies none of those four by construction. The relation runs close to inverse: the less a flexibility costs the load, the less the system can rely on it. The contract exists to resolve that, converting an unobligated capability into an obligated one at a price both parties accept.
Two consequences follow for anyone writing that contract. The flexible share covers a fraction of the campus rather than the campus: interactive inference runs under a latency bound and cannot wait, while training, batch inference and offline processing carry the deferrable part, so a commitment sized against today’s training-heavy mix may bind differently on a site that later shifts toward serving. And the same megawatts carry different flexibility depending on who holds the control rights. An operator running its own models can defer its own training; a colocation or cloud provider cannot pause a tenant’s workload, and reaches the same obligation through batteries, on-site generation, and the shape of its own contract portfolio instead. That mechanism drives the fleet-depth asymmetry noted earlier: a term costing a global operator almost nothing, because it has somewhere else to send the work, becomes a firm curtailment obligation for a single-site provider with nowhere to send it, and a tariff modelled on the first will misprice the second.
Interruptible load is not a new idea#
Curtailable service is neither an untested proposition nor an invention of the data-center era. Aluminum smelters, chlor-alkali plants and other electro-intensive industry have taken interruptible rates for decades, in exchange for prices unavailable to firm customers. ERCOT has dispatched load as a frequency resource since the 2000s through Loads Acting as Resources: on February 26, 2008, roughly 1,100 MW of contracted load was deployed automatically after a frequency decline — load acting as the remedy rather than the disturbance, and the mirror image of every event in Table 4A. Cryptocurrency mining has since become the largest voluntarily price-responsive load class on any North American grid, idling within seconds when prices spike.
The scale is not marginal either. Demand response accounts for about 5% of the capacity cleared in PJM’s most recent auctions — comparable to the entire hydro fleet and larger than wind and solar combined in that market. The movement in that figure is instructive. Between the 2026/27 and 2027/28 auctions, cleared demand response rose from 5,531 MW to 7,299 MW with essentially no change in the megawatts offered. The increase came from an accreditation reform approved by FERC in May 2025 that raised demand response’s effective load carrying capability. Nearly 1,800 MW of capacity appeared because the rules for counting it changed, which demonstrates the section’s argument directly: the binding constraint on flexibility is how it is valued and contracted, not whether the capability exists.
What each region is actually building#
The instruments differ by region, and the differences follow from market structure rather than from disagreement about the physics. Regions with centralised capacity markets already have a venue in which flexibility can be paid — so their work is accreditation reform. Regions without one are writing bespoke service products instead.
| Region | Instrument | How it works | Where it stands |
|---|---|---|---|
| PJM | Demand response in the capacity market (RPM); Non-Firm Contract Demand for co-located load | DR bids into the Base Residual Auction and is accredited by ELCC like any other resource; non-firm service caps grid withdrawals by contract. | DR is about 5% of the cleared supply mix. An accreditation reform approved by FERC in May 2025 raised DR’s ELCC and lifted cleared DR from 5,531 MW to 7,299 MW between the 2026/27 and 2027/28 auctions — with essentially no change in the megawatts offered. |
| ERCOT | Controllable Load Resource; Emergency Response Service; four-coincident-peak price response | A CLR registers as dispatchable and accepts ERCOT dispatch and ramp-rate instructions; ERS is a contracted emergency reserve; 4CP exposure gives every large load an unpriced incentive to curtail at the annual peaks. | CLR pathway formalised under NPRR1188 and NPRR1325. Cryptocurrency load has been the practical proving ground: highly price-responsive, and able to idle within seconds. |
| MISO | Load Modifying Resources; Expedited Resource Addition Study on the supply side | LMRs are registered demand-side resources counted toward the planning reserve margin and callable in emergencies. | MISO tightened LMR notification and availability requirements after the 2022–23 winter events, trading volume for dependability. |
| SPP | Conditional High Impact Large Load Service (CHILLS) | Long-term transmission service granted ahead of the upgrades that would normally be required, conditioned on curtailability and separate telemetry. | Accepted by FERC on June 5, 2026. The closest instrument in force to a bounded, priced curtailment product. |
| NYISO / ISO-NE / CAISO | Special Case Resources; Active Demand Capacity Resources; Proxy Demand Resource and the Demand Response Auction Mechanism | Aggregator-mediated programs that let retail customers offer curtailment into wholesale capacity and energy markets. | Long-established and modest in scale relative to peak; built for aggregations of small commercial load rather than for single gigawatt-scale customers. |
| Bilateral utility contracts | Custom demand-response terms inside a large-load service agreement | The utility and the customer negotiate curtailment terms directly, outside any market product. | The fastest-moving venue. Google has integrated about 1 GW across Indiana Michigan Power, TVA, Entergy Arkansas, Minnesota Power and DTE. |
Table 6A — Flexible and curtailable load instruments by region. Program scale is given where operators publish it; several of the smaller programs report on different bases and are described qualitatively rather than compared numerically.
Sources: PJM; ERCOT; MISO; SPP; NYISO; ISO-NE; CAISO tariff and auction materials. Program scale given where the operator publishes it.
The bilateral row sets the terms first, and one contract repays close reading because it tests this section’s central proposition. In July 2025 Indiana Michigan Power filed a special contract with Google covering a data center campus at Fort Wayne, pairing a clean-capacity arrangement with a demand-response product built for machine-learning workloads. Most commercial terms are confidential, but the filed structure is public: an unlimited number of curtailments, each capped at a maximum duration of 12 hours in summer and 15 in winter. Google has since extended the model to TVA, Entergy Arkansas, Minnesota Power and DTE, reaching about 1 GW of contracted demand response, and reports that the approach lets new campuses connect faster than a firm-service queue position would allow.
That structure is the inverse of the product this section has been describing. This report has argued for a cap on annual hours with bounded event duration and frequency; the I&M contract bounds duration but leaves frequency open. Both are bounded obligations — a customer can price either — and the divergence suggests the market is converging on boundedness as the requirement rather than on any particular parameter. That matters for tariff design: a regulator writing a flexibility product should be less concerned with matching some canonical hour count than with ensuring that every dimension of the obligation is closed and stated. What developers decline is not curtailment but discretion.
Why unlimited events can be the cheaper side of the bargain#
For most industrial customers an unlimited number of curtailments would be the harsher of the two terms. That ordering reverses for a workload that is deferrable and measured in total throughput rather than in delivery time. A training run cares about the compute delivered over weeks; whether a given hour of it happens on Tuesday or Thursday changes little, provided the interruption is short enough that checkpointing absorbs it and the cooling plant and the schedule recover. What such an operator must underwrite is the worst single event — the outage long enough to break a delivery commitment or strand very expensive accelerators for a shift. Cap the duration and that risk is bounded. The number of events then bears on cost roughly in proportion to lost throughput, which is recoverable, rather than to risk, which is not.
The count is also bounded in practice even where the contract leaves it open, because the counterparty has no reason to call events indiscriminately. I&M serves its PJM capacity obligation under the Fixed Resource Requirement rather than by buying in the capacity auction, so the value of a curtailment is concentrated in the handful of hours that set its peak obligation. A utility in that position calls the hours it must and leaves the rest alone. An unlimited-count term granted to a counterparty whose own economics confine it to peak hours is a smaller concession than it appears on the page.
The term is affordable to this signatory for a further reason, which does not generalise. Google operates a global fleet and has been shifting workloads between regions and across hours since well before these agreements — the utility contracts followed a demonstration with Omaha Public Power District in which machine-learning demand was reduced across three grid events. For an operator who can move the work, a curtailment at one campus is largely absorbed elsewhere, and the marginal cost of one more event approaches zero. For a single-site colocation provider running one tenant’s inference workload under a service-level agreement, the same term is difficult to accept at any price, because the work has nowhere to go.
The asymmetry repeats a pattern identified earlier in this report. Section 1 noted that a $50,000/MW security deposit screens for balance-sheet depth as much as for project credibility. Flexibility terms screen for fleet depth in the same way. A curtailment product written around what a hyperscaler with a global fleet can accept is not a neutral template for the rest of the market, and a regulator who takes the Google contracts as the model may write a tariff that only three or four firms in the country can use. The bounded-hours structure this section recommends is more portable precisely because it does not assume the customer has somewhere else to run.
What the flexibility is actually worth, and to whom#
Compensation is only partly documented. The I&M contract, approved by the Indiana Utility Regulatory Commission in 2026 under Cause No. 46276, compensates Google through credits — one associated with the demand-response structure and one with a parallel Clean Capacity Arrangement, under which Google transfers long-term accredited clean capacity to the utility and bears a penalty if it fails to deliver. The magnitude of those credits, and the contract capacity in megawatts, are redacted.
The funding source is identifiable. As an FRR entity, I&M must hold capacity against its own peak rather than buy it in the auction. Every megawatt of peak the customer removes therefore spares the utility a megawatt of generation it would otherwise procure or build — at a time when the marginal price of that capacity in PJM has cleared at the cap for three consecutive auctions. The utility’s stated case is exactly this: the arrangement reduces its long-term generation requirements and financial commitments, and the savings are shared with other customers rather than accruing only to the signatory. Whether the split is fair is the question the redactions make unanswerable from outside.
The credit is probably not the largest component of the compensation. Google’s own account of why it signs these agreements is that flexibility lets new campuses connect to local grids faster, bridging the gap between load growth and the longer timelines required for new generation. Where time to energization binds and the alternative runs to a queue position measured in years, an earlier energisation date carries more value than a per-megawatt payment. SPP’s CHILLS makes the same trade explicit, and PJM’s Non-Firm Contract Demand offers it as a tariffed product. Flexibility buys speed; the credit follows as a secondary term.
This conclusion raises a governance problem. Should bilateral utility contracts become the principal venue for flexible service — and on current evidence they outpace any tariff or market product — confidential filings will set the terms of the most important reform available. The Citizens Action Coalition intervened in the Indiana docket to argue precisely this, objecting that the redactions extended to contract capacity and to the penalty provisions that protect ratepayers, and that no public ratepayer-impact analysis accompanied the filing. The objection concerns evaluability rather than merit: the public cannot assess a bargain struck on its behalf. A product this report recommends should not arrive in a form that only its two signatories can read (Sections 3 and 7).
For context on why this matters so much: U.S. data-center demand should reach roughly 66 GW in 2027, up from 31 GW in 2025, with the data-center share of summer peak rising from about 4.1% to 8.5%. Berkeley Lab's range spans 74–132 GW of growth and 6.7–12% of U.S. electricity by 2028. Absent flexibility, that translates into an estimated 25–50 GW of incremental fossil generation by 2030.
Proposed solutions#
- Non-firm / curtailable interconnection service. The core product. SPP's HILL-adjacent service and PJM's NonFirm Contract Demand both offer a faster energization date in exchange for interruptibility. FERC's Category 4 requires every RTO to address transmission service for flexible large loads.
- ERCOT Controllable Load Resource (CLR) and ramp-rate obligations. Large loads registering as CLRs become dispatchable in the ERCOT market and accept dispatch ramp rates as a condition of that participation; NPRR1325 imposes registration and dispatch obligations on large loads generally, but ramp-rate discipline follows the CLR election rather than applying to every large load.
- Flexibility credited in planning, not treated as an afterthought. GridLab/Telos' NV Energy case study found 1– 2 GW of data-center flexibility produces roughly $300–400 million NPV savings over 2025–2050 when it is embedded in integrated resource planning. Most utilities still model large loads as firm and inflexible.
- Cheaper flexibility from elsewhere on the system. Headroom does not have to come from the data center. Public Service Co. of Oklahoma procured demand-response capacity at about $32/kW-year, versus $100+/kW-year for behind-the-meter generation at a data center. Expanding conventional DR often delivers the same MW of headroom at lower cost.
- Technical flexibility levers inside the facility. Thermal storage (ice banks, chilled-water tanks) can shift 20– 40% of cooling load for 4–8 hours; DVFS modulates GPU/CPU draw; spatial shifting and 'compute peering' move workloads between regions; on-site BESS bridges short events without touching compute at all.
- Transparency as an enabler. Publishing transmission availability, risk assessments, and the specific conditions under which curtailment will be called lets developers price the risk instead of refusing it.