Home / Benchmarking

Benchmarking

How well each layer performs, measured on three axes - and stated plainly, including where the Almanac is weaker than the alternatives.

7 property areas 3 evidence axes 41 years tested v2026.1

The niche this is built for

There are already many single-epoch geospatial datasets, usually covering one or a few recent years and one or a few properties - biomass (GEDI), water (Forests-to-Faucets), disturbance (Hansen), fuels and wildfire (LANDFIRE), vegetation type (RAP). Many offer high spatial accuracy and rest on foundational work that gave the Almanac excellent starting points.

Improving on their spatial accuracy is not the aim. The field already knows a great deal about mapping near-current biomass, water balance, fuel and wildfire, and highly capable groups are working on it. The aim is to extend those records across the full Landsat archive with enough temporal precision to say reliably how conditions changed over four decades at 30 m, from single pixels to CONUS, and to difference years to show where and when. Confirming that recent-year patterns agree with established products is a necessary baseline for that - but it is not sufficient on its own, and it is not what this record is for.

Three axes

External skill

Spatial and temporal accuracy against real observations - river gauges, fire perimeters, field plots. Direct comparison with high-priority real-world conditions.

Internal coherence

Stability within a constant cohort of never-disturbed pixels. Ensures multi-year observations see real change rather than dataset artefacts, noise or drift.

Cross-dataset consistency

Agreement with established datasets such as LANDFIRE, GEDI and RCMAP. A minimum quality threshold - these products are widely used and accepted.

The first two axes are the ones that matter most here. Both speak directly to temporal precision. The never-disturbed-cohort test asks one specific question: across 41 years and four Landsat sensor changeovers, does an undisturbed pixel stay put? A seam at a satellite transition is invisible in any single-epoch comparison, and fatal to measuring change. Together the three axes ask whether the Almanac is spatially accurate enough to sit alongside established datasets, and whether it holds together across four decades.

Results at a glance

Every property carries both a high-confidence and a lower-confidence verdict. Follow a row for the measurements behind it.

PropertyHigh confidenceLower confidence
Aboveground biomassMagnitude of AGB; rate of growth in young stands; abrupt disturbance lossDiffuse mortality and closed-canopy growth
Cover and structureTree-cover level and pattern; loss scales with severity; precise in timeNo independent change accuracy; non-tree fractions and diffuse change
DisturbanceFire location, timing and severityNon-fire-agent magnitude and small-event detection
Drought dieoffHeight x multi-year deficit interaction locates dieoffPer-pixel severity and absolute totals
WaterRunoff in space and time; AET magnitude; long-term precision and stabilityWater response to disturbance; possible saturation at high AET
Fire: flame lengthIdentifies pixels at elevated risk of high-severity fire, and tracks them precisely for four decadesWeaker within shrub and timber-understory fuel classes
Fire: burn probabilityMaps burn probability across the landscape and tracks elevated risk for four decadesMoran and Pyrologix predict subsequent burning better; WA is not calibrated to absolute probability

Where the Almanac is strong, and where it falls short

What the measurements support, and where they do not. Where another product does better, it is named.

Well supported

Weaker, and by how much

On circularity. Several comparisons here are not independent, and are labelled as such rather than presented as validation. All gridded cover products are Landsat-derived; USFS TCC trains the Almanac's tree fraction and RAP trains shrub, herb and bare, so only RCMAP is independent for all four. MTBS and Hansen share Landsat lineage with the disturbance layer. Runoff comparisons are driven on both axes by precipitation of common origin - which is why the AET comparisons, which subtract that out, are the stronger evidence. Agreement under circularity is a consistency check, not proof.

By property

A work in progress

The pipeline reprocesses every year and every property at once, producing a complete new version while previous versions are archived. Full reprocessing combined with benchmarking creates a continuous improvement cycle: it allows updates with minimal latency; it ties the properties to one another, which is what makes cross-theme tradeoffs quantifiable; testing one property informs the others; and it removes any hesitancy to make improvements that might otherwise break the time series.

These results should be read as one turn of that cycle. They document the state of the dataset now, and they identify where the next round of development is aimed.