- Home
- Documentation
- API reference
- Corrections
- hyperproc.correct.pipeline
hyperproc.correct.pipeline¶
Source-derived reference
Generated from the current hyperproc 0.1.2 checkout.
Implementation: hyperproc/correct/pipeline.py. Signatures, defaults, docstrings, and expandable source are extracted statically; the module is not imported or executed. Names beginning with _ are implementation details, not a stable public API.
Use the function signature as the authority for individual parameter defaults and return annotations. Original docstrings sometimes group parameter names or wrap return descriptions across lines; these descriptions are preserved rather than inferred or rewritten.
Run topographic and BRDF corrections on the cubes the readers return.
The flow is sample -> fit -> apply -> export, and each step is a function you can call on its own:
- :func:
sample_imagereads a spread of whole reader chunks from one cube (all bands, ~10% of the image, capped) and keeps everything a fit needs: reflectance, geometry, NDVI, the six Zhai index bands, the swath footprint. Chunk-aligned reads matter: the cubes live on slow disks and a request that straddles two compressed chunks costs both. - :func:
fit_topofits one image's SCS+C (or cosine / C / SCS) coefficients and runs the illumination diagnostic whose verdict says whether the correction is warranted for that product. - :func:
fit_brdfpools the samples of a group of images - the lines of one site and flight day - and fits FlexBRDF, after checking that the group's view geometry is diverse enough for the coefficients to be identifiable.
This module is for airborne imaging spectroscopy (NEON AOP, the AVIRIS
family): the corrections rely on per-pixel terrain layers and on the
across-track view-angle diversity of flightline swaths. Satellite products
(EMIT, EnMAP, PRISMA, DESIS, PACE, Tanager) are refused at every entry point.
* :func:apply builds the corrected cube lazily (dask), block by block, so
export streams; :func:export writes GeoTIFF + band table + provenance.
* :func:view_dependence and :func:overlap_agreement are the real-data
consistency checks the notebooks report.
Nothing here needs a coefficient file from anywhere else; the fits are made
from the cubes themselves and written with :mod:hyperproc.correct.coefficients.
TOPO_CALC = {'ndvi': [0.1, 1.0], 'slope_min_deg': 5.0, 'cos_i_min': 0.12, 'cloud': {'T1': 1.0, 't2': 0.5, 't3': 1 / 3, 't4': 2 / 3, 'T7': 15, 'T8': 15}}
module-attribute
¶
TOPO_APPLY = {'ndvi': [0.1, 1.0], 'slope_min_deg': 5.0, 'cos_i_min': 0.12}
module-attribute
¶
BRDF_CALC = {'ndvi': [0.1, 1.0], 'vza_min_deg': 2.0, 'edge_px': 30, 'cloud': {'T1': 0.01, 't2': 0.1, 't3': 0.25, 't4': 0.5, 'T7': 9, 'T8': 9}}
module-attribute
¶
BRDF_APPLY = {'ndvi': [0.05, 1.0]}
module-attribute
¶
INDEX_BANDS = {'blue': 440.0, 'green': 550.0, 'red': 660.0, 'nir': 850.0, 'swir1': 1570.0, 'swir2': 2110.0}
module-attribute
¶
GEOMETRY = ('sza', 'saa', 'vza', 'vaa', 'slope', 'aspect')
module-attribute
¶
AIRBORNE_SENSORS = ('NEON', 'AVIRIS', 'AVIRIS-3', 'AVIRIS-NG', 'AVIRIS-CLASSIC', 'AVIRIS-5')
module-attribute
¶
Sample
dataclass
¶
What a fit needs from one image, at a spread of sampled pixels.
Source code in hyperproc/correct/pipeline.py
132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 | |
stem: str
instance-attribute
¶
sensor: str
instance-attribute
¶
level: str
instance-attribute
¶
path: str
instance-attribute
¶
wavelength: np.ndarray
instance-attribute
¶
good: np.ndarray
instance-attribute
¶
rho: np.ndarray
instance-attribute
¶
sza: np.ndarray
instance-attribute
¶
saa: np.ndarray
instance-attribute
¶
vza: np.ndarray
instance-attribute
¶
vaa: np.ndarray
instance-attribute
¶
slope: np.ndarray
instance-attribute
¶
aspect: np.ndarray
instance-attribute
¶
cos_i: np.ndarray
instance-attribute
¶
raa: np.ndarray
instance-attribute
¶
ndvi: np.ndarray
instance-attribute
¶
rows: np.ndarray
instance-attribute
¶
cols: np.ndarray
instance-attribute
¶
edge_ok: np.ndarray
instance-attribute
¶
edge_px: int
instance-attribute
¶
index: dict
instance-attribute
¶
blocks: list
instance-attribute
¶
all_ndvi: np.ndarray
instance-attribute
¶
cloud_stats: dict
instance-attribute
¶
fraction: float
instance-attribute
¶
seed: int
instance-attribute
¶
n_valid_total: int
instance-attribute
¶
elapsed: float = 0.0
class-attribute
instance-attribute
¶
strategy: str = 'chunks'
class-attribute
instance-attribute
¶
topo_sums: dict | None = None
class-attribute
instance-attribute
¶
cos_i_calc: np.ndarray | None = None
class-attribute
instance-attribute
¶
n_calc_topo: int = 0
class-attribute
instance-attribute
¶
_cloud_cache: dict = field(default_factory=dict, repr=False)
class-attribute
instance-attribute
¶
_cloud_precomputed: dict = field(default_factory=dict, repr=False)
class-attribute
instance-attribute
¶
n: int
property
¶
__init__(stem: str, sensor: str, level: str, path: str, wavelength: np.ndarray, good: np.ndarray, rho: np.ndarray, sza: np.ndarray, saa: np.ndarray, vza: np.ndarray, vaa: np.ndarray, slope: np.ndarray, aspect: np.ndarray, cos_i: np.ndarray, raa: np.ndarray, ndvi: np.ndarray, rows: np.ndarray, cols: np.ndarray, edge_ok: np.ndarray, edge_px: int, index: dict, blocks: list, all_ndvi: np.ndarray, cloud_stats: dict, fraction: float, seed: int, n_valid_total: int, elapsed: float = 0.0, strategy: str = 'chunks', topo_sums: dict | None = None, cos_i_calc: np.ndarray | None = None, n_calc_topo: int = 0, _cloud_cache: dict = dict(), _cloud_precomputed: dict = dict()) -> None
¶
cloud_flags(params: dict) -> np.ndarray
¶
Zhai cloud/shadow flag per sampled pixel (True = bad), evaluated in 2-D on each sampled block with the scene-wide statistics (or, for the whole-image strategy, taken from the mask computed over the full image).
Source code in hyperproc/correct/pipeline.py
topo_calc_mask(spec: dict = TOPO_CALC) -> np.ndarray
¶
Source code in hyperproc/correct/pipeline.py
topo_apply_mask(spec: dict = TOPO_APPLY) -> np.ndarray
¶
Source code in hyperproc/correct/pipeline.py
brdf_calc_mask(spec: dict = BRDF_CALC, volume='ross_thick', geometric='li_dense_r', b_r=1.0, h_b=2.0) -> np.ndarray
¶
Source code in hyperproc/correct/pipeline.py
summary() -> str
¶
Source code in hyperproc/correct/pipeline.py
is_airborne(sensor) -> bool
¶
True for the airborne sensors this module serves.
_require_airborne(sensor, what: str) -> None
¶
_main_var(ds: xr.Dataset) -> str
¶
_index_bands(wavelength: np.ndarray) -> dict
¶
Band index of each index band present in this cube (within 40 nm).
Source code in hyperproc/correct/pipeline.py
_cloud_stats(index: dict, valid=None) -> dict
¶
Zhai scene statistics, or {"n": 0} (no cloud test) when the cube has
no blue/green index band - a wl_range starting above 480 nm, say.
Source code in hyperproc/correct/pipeline.py
_cloud_mask(index: dict, valid, stats, params) -> np.ndarray
¶
Zhai cloud/shadow flags, all False when the statistics are empty.
Source code in hyperproc/correct/pipeline.py
_geometry_2d(ds: xr.Dataset, keys=GEOMETRY, need=GEOMETRY) -> dict
¶
Geometry layers as float64 radians (NaN where missing). Layers in
need must exist; the others are filled with NaN when absent, so a
scene with angles but no terrain layers can still take a BRDF
normalisation (not a topographic correction).
Source code in hyperproc/correct/pipeline.py
sample_image(ds: xr.Dataset, fraction: float = 0.1, max_pixels: int | None = None, seed: int = 0, edge_px: int = 30, min_blocks: int = 4, strategy: str = 'pixels', topo_calc: dict = TOPO_CALC, brdf_calc: dict = BRDF_CALC) -> Sample
¶
Build the sample the fits are made from.
strategy="pixels" (default) is the reference procedure: the whole
flightline is read once to build every mask on every pixel (valid,
NDVI, Zhai cloud/shadow with scene statistics, swath edge); a random
fraction of the pixels passing the valid and NDVI masks is then drawn
and its full spectra read for the BRDF fit; and the topographic C is later
fitted on all pixels of the topo calc mask, from per-band regression
sums accumulated while reading (exactly the all-pixel NNLS/OLS result,
without holding the cube in memory). Cost: two passes over the cube.
max_pixels caps the random draw (None = no cap; memory is
n x bands x 4 bytes).
strategy="chunks" is the cheap alternative: fraction of the reader
chunks, spaced evenly over the chunk grid, are read and at most
max_pixels (default 250,000) random pixels are kept from them; cloud
statistics and the NDVI population come from the chunks read only. One
tenth of a pass instead of two - use it for quick looks.
Source code in hyperproc/correct/pipeline.py
_sample_chunks(ds: xr.Dataset, fraction: float, max_pixels: int, seed: int, edge_px: int, min_blocks: int) -> Sample
¶
Source code in hyperproc/correct/pipeline.py
_iter_chunks(da)
¶
Yield (ys, xs, block) for every reader chunk of a (y, x, wavelength) dask array, in order.
Source code in hyperproc/correct/pipeline.py
_sample_pixels(ds: xr.Dataset, fraction: float, max_pixels, seed: int, edge_px: int, topo_calc: dict, brdf_calc: dict) -> Sample
¶
Whole-image masks, random pixel draw, all-pixel topographic sums (two passes).
Source code in hyperproc/correct/pipeline.py
329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 | |
merge_samples(samples, stem: str | None = None) -> Sample
¶
Pool the samples of several cubes that belong together - the chunks of
one AVIRIS-5 flightline - into one :class:Sample, so a topographic C
is fitted per flightline rather than per chunk. Cloud statistics are
recomputed over all pooled blocks.
Source code in hyperproc/correct/pipeline.py
per_block_effects(sample: Sample, wavelengths=(549, 659, 849, 1651, 2202), spec: dict = TOPO_CALC, min_pixels: int = 500, split: int = 1) -> list
¶
The illumination effect fitted separately in each sampled block.
A single C per image assumes one relation between reflectance and cos i
over the whole image. This shows whether that holds: blocks that agree in
sign and size support a per-image C; blocks that disagree mean the pooled
slope is driven by which regions were pooled, not by illumination, and
the verdict should be read with that in mind. split divides every
sampled block into split x split sub-blocks (no extra reading), which
gives the consistency statistic more, smaller regions to compare.
Source code in hyperproc/correct/pipeline.py
fit_topo(sample: Sample, method: str = 'scs+c', fit: str = 'nnls', calc: dict = TOPO_CALC, apply_spec: dict = TOPO_APPLY, diagnostic_bands: int = 40, min_samples: int = 100, block_agreement: float = 0.7, block_min_pixels: int = 500, block_split: int = 2, block_t: float = 2.0) -> TopoCoefficients
¶
Fit one image's topographic coefficients and judge whether to use them.
Every good band gets a :func:hyperproc.correct.topo.fit_c - by default
NNLS, which keeps the intercept and therefore C non-negative (an OLS fit
on a product with an additive offset gives C < 0 and a singular factor;
with fit="ols" such bands come back negative_intercept and are not
corrected). The illumination diagnostic runs on up to diagnostic_bands good bands and
its verdict (correct | skip | refuse | inconclusive) is stored with the
coefficients. A correct or refuse verdict additionally requires
the effect to be spatially consistent - at least block_agreement of the
sampled blocks, each split into block_split x block_split sub-blocks
with >= block_min_pixels fit pixels and a significant slope of their
own (|t| >= block_t), must share the pooled sign and their median
effect must exceed min_effect - otherwise
it is downgraded to inconclusive with the reason recorded. Bands the
reader flags bad get status bad_band and no C.
Source code in hyperproc/correct/pipeline.py
502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 | |
angular_diversity(sza, vza, raa, mask=None, volume='ross_thick', geometric='li_dense_r', b_r=1.0, h_b=2.0, span_min_deg=8.0, cond_max=2000.0) -> dict
¶
Can a kernel fit tell f_vol and f_geo apart on this geometry?
Reports the p05-p95 span of view zenith, the circular spread of relative
azimuth, and the condition number of the [k_vol, k_geo, 1] design
matrix. Verdict ok needs a view-zenith span of at least
span_min_deg and a condition number below cond_max; narrow-swath
narrow-swath scenes (2-3 degrees of view zenith) fail this, airborne swaths
(15-20 degrees) pass.
Source code in hyperproc/correct/pipeline.py
_topo_correct_sample(s: Sample, tc: TopoCoefficients | None, wl: np.ndarray) -> np.ndarray
¶
The sample's spectra topographically corrected with tc, on the
sample's own wavelength axis (wl must be s.wavelength; a group
member's axis may differ from the first sample's by up to 1 nm, more than
the 0.5 nm coefficient alignment tolerates).
Source code in hyperproc/correct/pipeline.py
fit_brdf(samples, topo=None, calc: dict = BRDF_CALC, apply_spec: dict = BRDF_APPLY, volume: str = 'ross_thick', geometric: str = 'li_dense_r', b_r: float = 1.0, h_b: float = 2.0, sza_ref='group', num_bins: int = 18, ndvi_min: float = 0.05, ndvi_max: float = 1.0, perc_min: float = 10, perc_max: float = 95, second_split: bool = True, group_id: str | None = None, force: bool = False, force_topo: bool = False, **diversity_kw) -> BRDFCoefficients
¶
Fit FlexBRDF on a group of samples (one site, one flight day) - airborne only.
If topo (a list of :class:TopoCoefficients, one per sample, or
None entries) is given, each sample is topographically corrected first, so
the BRDF fit sees the reflectance it will later be applied to. sza_ref
is "group" (mean solar zenith of the fitted pixels) or degrees. The
angular-diversity check runs first; an insufficient verdict raises
unless force=True, because coefficients fitted on a narrow view range
are noise.
A sample is topographically corrected only when its coefficients' verdict
is "correct"; for skip/refuse/inconclusive the sample is
used as delivered (with a warning), so the BRDF fit matches what
:func:apply will do with the same default. force_topo=True applies
the coefficients regardless - then pass force_topo=True to apply
as well, so fit and product agree.
Source code in hyperproc/correct/pipeline.py
apply(ds: xr.Dataset, topo: TopoCoefficients | None = None, brdf: BRDFCoefficients | None = None, block_bytes: float = 200000000.0, notes: dict | None = None, force_topo: bool = False, brdf_ratio_max: float | None = 5.0) -> xr.Dataset
¶
The corrected cube, lazily: topo first (if given), then BRDF.
The topographic coefficients are applied only when their verdict is
"correct". Otherwise the topo stage is skipped with a warning and the
product is named for the stages actually applied, unless
force_topo=True (recorded in the provenance as topo_forced).
brdf_ratio_max bounds the BRDF factor per pixel and band (see
:func:hyperproc.correct.brdf.apply_flex); None disables the bound.
Returns a copy of ds whose cube is a dask array evaluated block by
block when written; geometry layers and coefficients are aligned to the
cube's wavelength axis here. attrs['stem'] gains _topo, _brdf
or _topo_brdf so exports never overwrite the input's name.
Source code in hyperproc/correct/pipeline.py
721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 | |
export(ds: xr.Dataset, out_dir, wavelengths=None, window=None, suffix: str = '', compress: str = 'deflate', workers: int = 4, overviews=None, overview_resampling: str = 'average', format: str = 'GTiff') -> Path
¶
Write the (corrected) cube as GeoTIFF or ENVI, + band CSV + provenance JSON.
wavelengths selects bands by nearest wavelength (nm); window is
(y0, y1, x0, x1) in pixels. Either makes the output a subset, and
suffix should then say so in the file name. The provenance JSON
records the stages, coefficient sources and any subsetting. workers
caps the dask threads while computing: every block holds a full-swath,
all-band slab, so a 14-thread default can exceed memory on 400-band cubes.
overviews (True or a list of factors) adds internal pyramids for GIS
display after the write, see :func:hyperproc.build_overviews.
Source code in hyperproc/correct/pipeline.py
footprint_bounds(ds: xr.Dataset)
¶
(x0, y0, x1, y1) map bounding box of the cube's pixel footprint,
rotation included (all four corners, not just two).
Source code in hyperproc/correct/pipeline.py
footprint_polygon(ds: xr.Dataset)
¶
The cube's pixel footprint as a shapely polygon (a rotated rectangle for flight-aligned grids). Bounding boxes overstate the overlap of rotated swaths - two lines that never touch can still have intersecting boxes.
Source code in hyperproc/correct/pipeline.py
footprint_overlap(ds_a: xr.Dataset, ds_b: xr.Dataset)
¶
Intersection of the two footprint polygons: (area, bounds) with
bounds = (x0, y0, x1, y1), or (0.0, None) when they do not touch.
Source code in hyperproc/correct/pipeline.py
_window_for_bbox(ds: xr.Dataset, bbox, max_rows: int | None)
¶
Source code in hyperproc/correct/pipeline.py
_window_bounds(ds: xr.Dataset, win)
¶
Source code in hyperproc/correct/pipeline.py
overlap_windows(ds_a: xr.Dataset, ds_b: xr.Dataset, max_rows: int = 1200)
¶
Pixel windows (y0, y1, x0, x1) in each cube covering one common map
region inside the overlap of their footprints, with roughly max_rows
rows in each (a rotated grid needs somewhat more rows to cover an
axis-aligned region). None when the footprints do not overlap.
Source code in hyperproc/correct/pipeline.py
find_overlapping_pair(dss, exclude_same=None)
¶
(i, j, area) of the pair of cubes whose footprint polygons overlap
most, or None. exclude_same(ds) returns a key; pairs with equal keys
are skipped (e.g. chunks of one flightline when the question is between
lines).
Source code in hyperproc/correct/pipeline.py
_common_grid(srcs, mode='union', res=None)
¶
A north-up grid (crs, transform, width, height) over the union or intersection of the sources' footprints, at the finest pixel size.
Source code in hyperproc/correct/pipeline.py
_reproject(src, crs, transform, W, H, resampling='nearest')
¶
One rasterio dataset's bands on the common grid, NaN where absent.
Source code in hyperproc/correct/pipeline.py
mosaic_geotiffs(paths, out_path, method: str = 'first', resampling: str = 'nearest', res: float | None = None, overviews=None, overview_resampling: str = 'average') -> Path
¶
Merge GeoTIFFs into one north-up file.
Flight-aligned (rotated) grids - every AVIRIS line - are reprojected onto a
common north-up grid at the finest pixel size first, so lines with
different rotations can be mosaicked. method="first" keeps the first
file's value where they overlap (seams stay visible, which is what a
seam check wants); "mean" averages the overlap. overviews adds
internal pyramids to the result (see :func:hyperproc.build_overviews).
Source code in hyperproc/correct/pipeline.py
overlap_agreement_tifs(path_a, path_b, resampling: str = 'nearest', min_value: float = 0.005) -> dict
¶
Do two GeoTIFFs agree where they overlap? Both are put on a common
north-up grid over the intersection; per band: median and p90 absolute
relative difference, median ratio, correlation. The reprojected arrays
(a, b, valid) come back too, for difference maps.
Source code in hyperproc/correct/pipeline.py
seam_check(ds_a: xr.Dataset, ds_b: xr.Dataset, cor_a: xr.Dataset, cor_b: xr.Dataset, out_dir, wavelengths=(450, 550, 650, 850, 1650, 2200), max_rows: int = 1200, tag: str = 'seam', overviews=None) -> dict
¶
Do two adjacent images still mismatch where they overlap after correction?
Exports the overlap windows of both images at a few wavelengths, raw and corrected, mosaics each pair with the first image on top (so the seam is where any colour mismatch shows), and measures the agreement in the overlap for both. Returns the paths and the two agreement dicts.
Source code in hyperproc/correct/pipeline.py
view_dependence(samples, brdf: BRDFCoefficients | None = None, topo=None, wavelengths=(550, 660, 850, 1650, 2200), ndvi_classes=((0.3, 0.5), (0.5, 0.7), (0.7, 0.9)), calc: dict | None = None) -> dict
¶
How much does reflectance still depend on view geometry, per NDVI class?
Within each NDVI class (a proxy for one cover type) the least-squares
slope of reflectance on the volume kernel, times the kernel's p05-p95
span, over the mean reflectance: the view-driven change as a fraction of
the mean. Reported before and, if brdf is given, after normalisation.
A working correction should cut it several-fold. Pixels come from the
BRDF calc mask - the one the coefficients were fitted with, or calc.
Source code in hyperproc/correct/pipeline.py
overlap_agreement(ds_a: xr.Dataset, ds_b: xr.Dataset, wavelengths=(550, 660, 850, 1650, 2200), n_points: int = 20000, n_strips: int = 4, seed: int = 0) -> dict
¶
Do two overlapping north-up cubes on one grid agree where they overlap?
Samples map coordinates inside the intersection (in a few row strips, to keep the chunk reads bounded), takes the nearest pixel from each cube and reports the median absolute relative difference and correlation per band. Run it on the raw cubes and on the corrected ones: a real BRDF correction must bring the two views of the same ground closer together.