geovalidate.HilbertKFold

class geovalidate.HilbertKFold(n_splits=5, level=16, random_state=None)[source]

Spatially balanced k-fold cross-validator via Hilbert curve ordering. [1]

Partitions observations into n_splits folds such that each fold is a spatially spread-out subsample covering the entire study area. Points are sorted along the Hilbert curve (via GeoSeries.hilbert_distance); every k-th point in this ordering is assigned to the same fold.

Because the Hilbert curve is continuous with no quadrant-boundary jumps, spatially nearby points always receive consecutive codes, so cycling through the sorted order reliably sends neighbours to different folds. The result is analogous to a quasi-random spatial sample: each fold is a low-discrepancy subsample of the full dataset rather than a compact geographic block.

This maximises the mean within-fold nearest-neighbour distance and ensures each fold’s training set is geographically representative of the whole region – the spatial equivalent of stratified random sampling.

Parameters:
n_splits : int, default 5

Number of folds.

level : int (1-16), default 16

Hilbert curve precision. Higher values give finer spatial discrimination; 16 is sufficient for any practical dataset.

random_state : int, RandomState instance, or None

Controls within-band shuffle (breaks ties among points with the same Hilbert code).

Examples

>>> skf = HilbertKFold(n_splits=5, random_state=0)
>>> for train_idx, test_idx in skf.split(gdf):
...     model.fit(X[train_idx], y[train_idx])
...     score = model.score(X[test_idx], y[test_idx])
__init__(n_splits=5, level=16, random_state=None)[source]

Methods

__init__([n_splits, level, random_state])

get_metadata_routing()

Get metadata routing of this object.

get_n_splits([X, y, groups])

get_params([deep])

Get parameters for this estimator.

set_params(**params)

Set the parameters of this estimator.

set_split_request(*[, groups])

Configure whether metadata should be requested to be passed to the split method.

split(X[, y, groups])

Yield (train_indices, test_indices) for each spatial fold.

get_metadata_routing()[source]

Get metadata routing of this object.

Please check User Guide on how the routing mechanism works.

Returns:

routing – A MetadataRequest encapsulating routing information.

Return type:

MetadataRequest

get_n_splits(X=None, y=None, groups=None)[source]
get_params(deep=True)[source]

Get parameters for this estimator.

Parameters:
deep : bool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:

params – Parameter names mapped to their values.

Return type:

dict

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**params : dict

Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

set_split_request(*, groups='$UNCHANGED$')[source]

Configure whether metadata should be requested to be passed to the split method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to split if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to split.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
groups : str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED

Metadata routing for groups parameter in split.

Returns:

self – The updated object.

Return type:

object

split(X, y=None, groups=None)[source]

Yield (train_indices, test_indices) for each spatial fold.

Parameters:
X : GeoDataFrame | GeoSeries | (n, 2) ndarray | (n,) or (n, 1) ndarray

Locations. Pass a 1-D array of time indices for time-series data.

y : ignored, present for sklearn API compatibility.

groups : ignored, present for sklearn API compatibility.

Yields:
  • train (ndarray of int)

  • test (ndarray of int)