Back to resources

April 27, 2026

The State-of-Charge Value Curve: From Dynamic Program to Offer Curve

battery storageoptimizationdynamic programmingbid curvesancillary servicesERCOT

Between the forecast and the market sits the control problem: given a distribution over tomorrow's energy and ancillary prices, what should a battery do, and, since the real-time dispatch engine acts mechanically against standing offer curves every five minutes, what curves should it file so that mechanical dispatch implements the intended behavior? This reference develops the answer: a value function over state of charge, its derivative as the offer curve, and the operating loop that keeps the curves fresh.

The central object

Work backward through the day. The value of holding a given amount of energy at hour t depends on what that energy could earn at hour t+1 and beyond, averaged over the price scenarios the forecast assigns probability to. The recursion bottoms out at an end-of-horizon value for whatever energy remains, and rolling it backward yields a value function: the worth of the battery's position at every hour and every state of charge.

The quantity that matters operationally is its slope, the marginal value of stored energy, written here as mu(s): what one additional megawatt-hour in the tank is worth at state of charge s. Two structural properties do most of the work:

  • The value function is concave in state of charge. Each stage maximizes a linear payoff over a convex feasible set, and expectation preserves concavity, so by backward induction the property holds everywhere.
  • Therefore mu is non-increasing in state of charge. The fuller the battery, the less another megawatt-hour is worth. This monotonicity is not a modeling convenience: it is exactly the property that makes the induced offer curve monotone in price, which is the shape the protocol requires a submission to have.

The economic reading: mu is the shadow price of the state-of-charge constraint, the internal transfer price at which the battery's present desk and future desk trade energy with each other. Every decision in the system reduces to a comparison against mu.

Solving it numerically

Grid the state-of-charge axis, 50 to 200 points, finer near the boundaries where mu moves fastest, and represent the value function by piecewise-linear interpolation, which preserves concavity exactly and makes each stage problem a small linear program. Cubic splines look smoother but can violate concavity between knots.

The price uncertainty can enter three ways, in increasing fidelity and cost: a scenario tree or lattice built from the forecast layer's scenario set (transparent, fast, the right starting point); a regime-augmented state that compresses price information into a low-dimensional variable such as spike risk for the remaining evening; or re-solving often, model-predictive-control style, on freshly conditioned scenarios. In practice the winning combination is scenario structure for the day-ahead solve plus frequent re-solving for freshness, since much of the value of an elaborate price state is captured operationally by re-solving with updated forecasts.

Two traps deserve names:

  • The terminal condition. The end-of-horizon value prices energy left at the end of the day. Too low and the model dumps energy before midnight; a hard zero manufactures artificial end-of-day discharge. Extend the horizon 12 to 24 hours past the decision day, or set the terminal value from a rolling estimate of overnight marginal value.
  • Tail-erasing scenario reduction. The lattice must preserve what the battery is paid to respect: the tail and its temporal shape. A reduction that washes out multi-hour spike episodes will undervalue duration and misprice mu in the evening hours. Verification is behavioral, not statistical: reprice a reference battery on the full scenario set and the reduced one, and accept only if the value gap is within tolerance.

The sanity atlas

Qualitative fingerprints of a correct mu surface, worth checking visually every day the system runs:

  • In state of charge: decreasing everywhere; near-flat and low when full; steepening sharply as the level approaches what the evening's spike exposure requires.
  • In time: rising through the afternoon toward the evening peak, collapsing after the last expensive hour, near-constant overnight.
  • In the forecast: the upper region scales with spike probability. A day whose spike risk doubles should visibly lift mu at low-to-mid state of charge; if reforecasting does not move the surface, the pipeline is disconnected.
  • The efficiency wedge: the charge-side valuation sits strictly below the discharge-side valuation, separated by round-trip losses plus cycling cost. That wedge is the minimum spread the battery should ever transact across, and any behavior implying trades inside it is a bug with a money leak attached.

From mu to the energy offer curve

The dispatch engine treats the battery as a price taker against its filed curve: it dispatches the offered quantities whose prices sit below the realized nodal price. The battery should therefore be willing to sell its marginal megawatt-hour exactly when the price exceeds that megawatt-hour's value if kept:

discharge offer price at fill s  =  mu(s) / discharge efficiency + cycling cost
charge bid price at fill s       =  charge efficiency * mu(s)

Sweeping the state of charge from the current fill downward traces the discharge offer curve: the first megawatt-hours out, moving along mu's flat high-fill region, are offered cheap; successive megawatt-hours, drawing down toward the precious low-fill region, are offered progressively higher. Because mu is non-increasing, the induced schedule is monotone in price, precisely the submittable shape. Sweeping upward traces the charge bid curve symmetrically.

The deep point bears repeating: filing these curves makes the dispatch engine, maximizing mechanically against them 288 times a day, implement the dynamic program's optimal policy, spike-chasing and hoarding included, with no intra-hour human action. The optimal bid curve is the derivative of a value function.

Compliance shaping is a small approximation problem: the protocol accepts a limited number of monotone piecewise-linear segments within floor and cap bounds (read the exact counts and bounds from the current protocols, not from memory). Choose breakpoints to minimize expected revenue loss under the forecast distribution, which concentrates them where dispatch probability mass lives and where mu bends fastest, at the low-fill knee.

The committed-hour overlay

Hours carrying a day-ahead sale are not curve-neutral. With a quantity already sold day-ahead, real-time settlement rides on the deviation between delivery and the commitment, and the risk calculus changes: an aggressive high-priced curve that brilliantly captures spikes on open capacity becomes dangerous against a committed sale, because being priced out of dispatch turns the award into a forced short position in a rising market.

The doctrine expressed as curve construction: offer the committed quantity at or below a conservative price, so it dispatches readily and closes the position, and reserve the aggressive high-mu pricing for quantities beyond the commitment. Symmetrically, a committed purchase should be bid to clear regardless. The commitment decision and the real-time curve strategy cannot be optimized separately; the recourse behavior of the curves must be the model inside the commitment optimizer, so the promise made at 10 AM is priced by exactly the policy that will be dispatched.

Ancillary offers from the same surface

Each reserve product is priced by opportunity cost, all computable from mu. Upward products (regulation up, responsive reserve, ECRS, non-spin) consume discharge headroom and encumber state of charge, since energy must stand behind the award. The offer price for a marginal megawatt of reserve is the sum of three terms: the foregone energy margin from the displaced dispatch, the internal price of the encumbered energy (mu evaluated along the planned trajectory), and the expected deployment cost.

In calm hours the foregone margin is near zero and reserves are offered cheap; ahead of spike risk, mu inflates the encumbrance term and the same megawatt of reserve is offered dear, exactly the behavior that keeps the battery's powder dry for energy when energy is where the money is. Downward products symmetrically consume charge headroom and add energy on deployment, often a benefit when mu is high, which is why regulation-down can rationally be offered near zero ahead of expensive evenings.

Since real-time co-optimization went live, the market engine performs the energy-versus-reserves trade-off every five minutes against these offers jointly. The design consequence worth internalizing: with co-optimization in the engine, mispricing your own opportunity costs is the only way to lose the allocation game.

Degradation: the shadow price that reprices everything

A linear throughput cost in dollars per megawatt-hour discharged is the workhorse approximation. Two upgrades earn their complexity. Warranty and augmentation contracts cap annual cycles or throughput; the clean treatment is a budget multiplier added to the throughput cost, tuned so induced annual cycling meets the budget, and revisited monthly: a battery ahead of its cycle budget in a mild spring should cheapen its throughput price into a volatile summer. Where cycle depth matters, a convex throughput cost approximates detailed degradation models at a fraction of the complexity while preserving the optimization structure.

The estimation of the cost itself is economics, not physics: replacement cost per usable megawatt-hour divided by cycle life, sanity-checked against the warranty's implicit price of a cycle. A degradation parameter last updated when cells cost twice as much is silently distorting every bid.

The operating loop

The pieces assemble into a daily cycle. Before the day-ahead close: condition scenarios on the morning information set, solve the dynamic program, and submit offers embodying its output. After awards publish: re-solve with awards fixed as commitments and file initial real-time curves. Through the adjustment period: reforecast on the growing information set, re-solve from the current telemetered state of charge (dispatch will have moved it off plan), and refile curves for the remaining hours. During delivery: no decisions, only monitoring, with an off-cycle re-solve triggered when realized prices or dispatch depart far enough from plan that the standing curves no longer represent current mu.

Re-solving costs little; stale curves cost real money precisely in the volatile hours when refreshing matters most. The engineering target is a loop fast and reliable enough that curve staleness is never the binding error term.

Validating the control layer

Separate the policy's quality from the forecast's, or improvement effort will be misdirected. The hindsight-optimal solve on realized prices gives each day's perfect-foresight value, an upper bound. The gap between it and realized policy value decomposes into forecast error and genuine policy loss (discretization, curve compression, staleness), and tracking both components answers the management question: should the next month of work go into the forecaster or the controller?

Below that, simulation invariants: assert the sanity atlas on every solve; check conservation (settled energy equals dispatched energy through the efficiency chain); and run adversarial scenarios (feed a guaranteed-spike scenario set and verify the system hoards; feed flat prices and verify it sits still, because a system that trades on flat prices is monetizing its own noise). Then walk-forward end to end: frozen decision-time information, curves generated, dispatch simulated against realized prices, settlement computed with two-settlement arithmetic, and P&L attributed among forecast edge, policy quality, and luck. That attribution triad is the management dashboard of the whole enterprise.