geovalidate.LocalBootstrap

class geovalidate.LocalBootstrap(n_bootstraps=100, bandwidth=None, k=None, kernel='gaussian', graph=None, random_state=None)[source]

Locally-weighted bootstrap – resampling with spatial (or temporal) proximity weighting.

Generates n_bootstraps resampled datasets each of size n. At every site i the observation placed there is drawn with replacement from all n observations; the probability is proportional to a kernel weight K(d_{ij} / bandwidth):

p_{ij} ∝ K(d_{ij} / bandwidth)

so nearby observations are drawn more often.

Works for spatial data (GeoDataFrame / (n, 2) coordinates) and for time-series data (pass a 1-D array of time indices; distance is then the absolute time lag).

Passing a pre-built libpysal.graph.Graph via graph overrides bandwidth and kernel.

Parameters:
n_bootstraps : int, default 100

Number of bootstrap resamples to generate.

bandwidth : float or None

Kernel bandwidth in the same units as the input distances. Required unless graph is provided.

kernel : str, default 'gaussian'

One of 'gaussian', 'exponential', 'bisquare', 'triangular', 'uniform', 'parabolic'.

graph : libpysal.graph.Graph or None

Pre-built spatial weights (must expose .sparse). Overrides bandwidth / kernel.

random_state : int, RandomState instance, or None

Notes

When input is a GeoDataFrame/GeoSeries and a supported kernel is given, a libpysal Graph is built internally – the weight matrix stays sparse. The dense O(n**2) path is used only for raw array inputs or when kernel is 'exponential' (not supported by libpysal).

Explored initially in

Statham, Thomas A. The Global Inconsistencies of Gridded Population Data at Different Spatial Scales. Diss. University of Bristol, 2024.

Examples

Spatial use (GeoDataFrame):

>>> lb = LocalBootstrap(n_bootstraps=200, bandwidth=5_000,
...                     kernel='bisquare', random_state=0)
>>> for indices in lb.sample(gdf):
...     boot = gdf.iloc[indices].copy()
...     boot.geometry = gdf.geometry.values   # restore original locations
...     model.fit(boot[features], y[indices])

Time-series use (1-D array of time steps):

>>> t = numpy.arange(len(df))
>>> lb = LocalBootstrap(n_bootstraps=100, bandwidth=12, random_state=0)
>>> for indices in lb.sample(t):
...     model.fit(X[indices], y[indices])
__init__(n_bootstraps=100, bandwidth=None, k=None, kernel='gaussian', graph=None, random_state=None)[source]

Methods

__init__([n_bootstraps, bandwidth, k, ...])

fit(X[, y])

Build the spatial weight graph.

get_metadata_routing()

Get metadata routing of this object.

get_params([deep])

Get parameters for this estimator.

sample(X[, donor])

Yield locally-weighted bootstrap samples.

set_params(**params)

Set the parameters of this estimator.

fit(X, y=None)[source]

Build the spatial weight graph.

Parameters:
X : GeoDataFrame | GeoSeries | (n, 2) ndarray | (n,) or (n, 1) ndarray

Locations.

y : array-like or None

Response variable; required when bandwidth='auto' or k='auto'.

Return type:

self

get_metadata_routing()[source]

Get metadata routing of this object.

Please check User Guide on how the routing mechanism works.

Returns:

routing – A MetadataRequest encapsulating routing information.

Return type:

MetadataRequest

get_params(deep=True)[source]

Get parameters for this estimator.

Parameters:
deep : bool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:

params – Parameter names mapped to their values.

Return type:

dict

sample(X, donor=None)[source]

Yield locally-weighted bootstrap samples.

Parameters:
X : GeoDataFrame | GeoSeries | (n, 2) ndarray | (n,) or (n, 1) ndarray

Target locations at which to draw samples. Pass a 1-D array of time indices for time-series data.

donor : GeoDataFrame, GeoSeries, ndarray, or None

Donor pool from which observations are drawn (same types as X). When None (default), observations are drawn from X itself (standard bootstrap).

When provided, the weight matrix is (|X|, |donor|): row i weights every donor observation by its kernel distance from X[i], so each new location draws locally from the donor pool.

If both X and donor are GeoDataFrames, each yielded value is a GeoDataFrame of donor rows with index and geometry overwritten to match X – i.e. donor attributes placed at new locations. Otherwise each yielded value is an array of donor index labels (or positional integers when donor has no pandas index).

Incompatible with a pre-built graph.

Yields:

result (GeoDataFrame or ndarray of shape (n,)) – When donor is a GeoDataFrame and X is a GeoDataFrame: a GeoDataFrame of selected donor rows with X’s index and geometry. Otherwise: result[i] is the label (or position) of the donor row drawn for location i.

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**params : dict

Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance