geovalidate.LeaveBallOut

class geovalidate.LeaveBallOut(radius)[source]

Buffered leave-one-out cross-validator.

For each observation i the test set is {i} and the training set is every observation farther than radius from i. Points inside the buffer zone are dropped entirely – they are neither in train nor in test.

Excluding nearby training points prevents spatial autocorrelation leakage: observations that are close enough to give the model an unfair advantage are withheld when predicting point i.

Parameters:
radius : float

Exclusion radius in the same units as the input coordinates. All points within this distance of the test point are excluded from the training set for that iteration.

Notes

For dense datasets or large radii, many training points will be dropped per iteration and some folds may have very small training sets. Consider BallKFold when a guaranteed minimum training size matters more than strict per-point buffering.

Examples

>>> lbo = LeaveBallOut(radius=2000)
>>> for train_idx, test_idx in lbo.split(gdf):
...     model.fit(X[train_idx], y[train_idx])
...     score = model.score(X[test_idx], y[test_idx])
__init__(radius)[source]

Methods

__init__(radius)

get_metadata_routing()

Get metadata routing of this object.

get_n_splits(X[, y, groups])

get_params([deep])

Get parameters for this estimator.

set_params(**params)

Set the parameters of this estimator.

set_split_request(*[, groups])

Configure whether metadata should be requested to be passed to the split method.

split(X[, y, groups])

Yield (train_indices, test_indices) for each observation.

get_metadata_routing()[source]

Get metadata routing of this object.

Please check User Guide on how the routing mechanism works.

Returns:

routing – A MetadataRequest encapsulating routing information.

Return type:

MetadataRequest

get_n_splits(X, y=None, groups=None)[source]
get_params(deep=True)[source]

Get parameters for this estimator.

Parameters:
deep : bool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:

params – Parameter names mapped to their values.

Return type:

dict

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**params : dict

Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

set_split_request(*, groups='$UNCHANGED$')[source]

Configure whether metadata should be requested to be passed to the split method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to split if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to split.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
groups : str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED

Metadata routing for groups parameter in split.

Returns:

self – The updated object.

Return type:

object

split(X, y=None, groups=None)[source]

Yield (train_indices, test_indices) for each observation.

Parameters:
X : GeoDataFrame | GeoSeries | (n, 2) ndarray

Locations.

y : ignored, present for sklearn API compatibility.

groups : ignored, present for sklearn API compatibility.

Yields:
  • train (ndarray of int) – All observations farther than radius from the test point.

  • test (ndarray of int) – Single-element array containing the index of point i.