Home / Benchmarking
How well each layer performs, measured on three axes - and stated plainly, including where the Almanac is weaker than the alternatives.
There are already many single-epoch geospatial datasets, usually covering one or a few recent years and one or a few properties - biomass (GEDI), water (Forests-to-Faucets), disturbance (Hansen), fuels and wildfire (LANDFIRE), vegetation type (RAP). Many offer high spatial accuracy and rest on foundational work that gave the Almanac excellent starting points.
Improving on their spatial accuracy is not the aim. The field already knows a great deal about mapping near-current biomass, water balance, fuel and wildfire, and highly capable groups are working on it. The aim is to extend those records across the full Landsat archive with enough temporal precision to say reliably how conditions changed over four decades at 30 m, from single pixels to CONUS, and to difference years to show where and when. Confirming that recent-year patterns agree with established products is a necessary baseline for that - but it is not sufficient on its own, and it is not what this record is for.
Spatial and temporal accuracy against real observations - river gauges, fire perimeters, field plots. Direct comparison with high-priority real-world conditions.
Stability within a constant cohort of never-disturbed pixels. Ensures multi-year observations see real change rather than dataset artefacts, noise or drift.
Agreement with established datasets such as LANDFIRE, GEDI and RCMAP. A minimum quality threshold - these products are widely used and accepted.
Every property carries both a high-confidence and a lower-confidence verdict. Follow a row for the measurements behind it.
| Property | High confidence | Lower confidence |
|---|---|---|
| Aboveground biomass | Magnitude of AGB; rate of growth in young stands; abrupt disturbance loss | Diffuse mortality and closed-canopy growth |
| Cover and structure | Tree-cover level and pattern; loss scales with severity; precise in time | No independent change accuracy; non-tree fractions and diffuse change |
| Disturbance | Fire location, timing and severity | Non-fire-agent magnitude and small-event detection |
| Drought dieoff | Height x multi-year deficit interaction locates dieoff | Per-pixel severity and absolute totals |
| Water | Runoff in space and time; AET magnitude; long-term precision and stability | Water response to disturbance; possible saturation at high AET |
| Fire: flame length | Identifies pixels at elevated risk of high-severity fire, and tracks them precisely for four decades | Weaker within shrub and timber-understory fuel classes |
| Fire: burn probability | Maps burn probability across the landscape and tracks elevated risk for four decades | Moran and Pyrologix predict subsequent burning better; WA is not calibrated to absolute probability |
What the measurements support, and where they do not. Where another product does better, it is named.
Tracks FIA to about 200 Mg/ha then saturates; agrees closely with GEDI lidar; misses diffuse mortality.
Tree cover matches NLCD closely and tracks burn severity; independent evidence for change accuracy is the weak point.
Finds and dates fire - 92% of high-severity area, 92% within a year - but under-registers diffuse and partial disturbance.
Locates dieoff through a super-additive height x deficit interaction, and leads aerial survey by about a year - but it is a screening layer, not a per-pixel count.
The strongest external test in the dataset: runoff matches 122 gauged basins in both space and time, and AET holds steady beneath a five-fold runoff swing.
Identifies where severe fire later occurred with consistent skill across the whole 39-year record, and shows an ecologically plausible fuel build-up in undisturbed forest.
Ranks where fire occurs well above chance across four decades - but an annually-updated FSim product beats it on every overlapping year.
The pipeline reprocesses every year and every property at once, producing a complete new version while previous versions are archived. Full reprocessing combined with benchmarking creates a continuous improvement cycle: it allows updates with minimal latency; it ties the properties to one another, which is what makes cross-theme tradeoffs quantifiable; testing one property informs the others; and it removes any hesitancy to make improvements that might otherwise break the time series.
These results should be read as one turn of that cycle. They document the state of the dataset now, and they identify where the next round of development is aimed.