An arXiv preprint dated Aug. 20, 2026, describes a computational framework that brings three tasks under one model: sizing a heat pump or air conditioner, simulating its dynamics and supporting predictive control. The study asks whether a single end-to-end differentiable finite-volume residual—the shared set of physics equations at the heart of the framework—can handle that workflow.
The reported evidence is computational: algebraic sizing inversions, simulated closed-loop runs and comparisons with public benchmark data. The cases are heterogeneous and are not pooled into one analytic sample.
One engine for several jobs
The plant represents a single-stage, subcritical, two-phase air-source heat pump or air conditioner, with finite-volume coils, a compressor, an electronic expansion valve, a lumped zone, optional humidity and frost, and several controllers. In practical terms, it models a liquid-and-vapor refrigerant system while breaking the coils into calculation sections and treating the conditioned space as one zone.
To keep the thermodynamics consistent, it uses pre-flashed pressure-and-enthalpy tables—enthalpy is a measure of a refrigerant’s energy—and a pressure equation that tracks both density derivatives for mass conservation in two-phase coils.
For sizing, the framework works backward from a stated thermal duty, directly inverting for compressor displacement, electronic-expansion-valve area and heat-exchanger tube counts. It combines a four-point cycle calculation with heat-exchanger matching and uses the same polytropic compressor map.
The sizing case studies contain nine algebraic inversions. In one R32 example, the target was 5.5 kW at 0°C outdoors and 20°C indoors; the model returned a compressor displacement of 25.9 cm³/rev, a 1.20 mm² EEV area, and 67 indoor and 63 outdoor tubes. Modeled capacity crossed the envelope load at −1.28°C.
The benchmark picture is mixed
Across 16 Ramírez mini-split runs, predicted cooling capacity had a 7.37% mean absolute percentage error (MAPE)—the average size of the percentage miss—and a largest absolute error of 19.23%. The corresponding MAPE was higher for compressor power, at 19.34%, and COP, or coefficient of performance, at 18.14%. No confidence intervals or formal uncertainty estimates were reported.
On four NREL on-period hardware-in-the-loop conditions, cooling-capacity errors were −1.62% at 35°C outdoor and −1.19% at 23.9°C. Heating errors were 19.56% at 7.2°C and −11.89% at −15°C. The comparison assumed R410A because the catalog did not identify the unit’s refrigerant.
The figures should not be read as one pooled performance score. The paper treats the sizing inversions, closed-loop simulations and benchmark subsets as separate computational cases rather than combining them into a single analytic sample.
Three closed-loop runs used an ISA PID controller and a superheat EEV under a one-hour quasi-steady reduction, and none was compared with a laboratory time series. In the R32 heating run, after 60 minutes, the zone ended at 20.12°C with 0.12 K absolute error. An R410A cooling run ended at 23.53°C.
The control result is still numerical
In a separate 90-second numerical comparison, PID ended at 17.30°C, 2.70 K from the setpoint; a linear model-predictive controller (MPC), which uses the model to anticipate the system’s response, ended at 19.34°C, 0.66 K away. That single demonstration put MPC closer to the setpoint in that run, but it was not a statistical comparison or a laboratory test.
A separate nine-setpoint Lee compressor-map check found largest JAX-versus-NumPy relative differences below 10⁻¹⁶ for both power and mass flow. That tests agreement between software implementations, not a measured-compressor comparison.
What the preprint does—and does not—show
The model’s operating scope is narrow. It excludes automatic defrost schedules, ducts, multi-zone buildings, transcritical CO₂, flash tanks, economizers, oil and piping inertia; it assumes acoustic equilibrium, thermodynamic-equilibrium slip, an oil-free refrigerant and a single lumped zone. Frost is represented as a lumped mass without an automatic defrost schedule, and the examples are dry.
The benchmark comparisons have their own caveats. Ramírez electrical power was calculated as current multiplied by 120 V without a power factor, while NREL predicted shaft power excludes fans and auxiliary heat, limiting direct COP comparisons. The benchmark scores are unfitted algebraic steady closes rather than integrations against 1 Hz laboratory traces.
Property-table interpolation can reconstruct states that do not necessarily coincide with a single Helmholtz flash. The reported COP is diagnostic rather than a catalog rating, and air-side zone heat need not equal refrigerant condenser duty at part load.
The authors present the framework as a foundation for automated machine synthesis, dynamic grid orchestration and hardware-control co-design. The supplied results do not establish deployed-system grid responsiveness, energy use or emissions, certified ratings or seasonal performance, or general performance in multi-zone buildings and other excluded cycle configurations.
Further work would need identified hardware with measured refrigerant, geometry, fan and auxiliary power and electrical power factor, plus longer full-DAE and hardware-in-the-loop control experiments. The study also leaves defrost, moisture and frost dynamics and other excluded configurations for future testing.
Paper data and sources
Original title: An end-to-end differentiable transient vapor-compression framework for automated machine sizing and unified optimal control
Authors: Sam Yang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text