geovalidate.BallKFold

class geovalidate.BallKFold(radius=None, n_splits=None)[source]

Spatially exclusive k-fold cross-validator.

Assigns observations to folds such that no two observations in the same fold are within radius r of each other. Each fold’s test set is a spatial independent set in the conflict graph whose edges connect pairs of points within r.

The number of folds is determined by greedy graph colouring and equals at most the maximum number of points within r of any single point plus one. Colouring uses a largest-degree-first ordering, which minimises the number of colours on most practical inputs.

Parameters:
radius : float or None

Exclusion radius in the same units as the input coordinates. Two points within this distance cannot share a fold. Mutually exclusive with n_splits.

n_splits : int or None

Target number of folds. The implied radius is computed as the minimum n_splits-th nearest-neighbour distance across all points (set by the densest region), stored as radius_ after split() is called. Mutually exclusive with radius.

Notes

When n_splits is given, the actual number of folds returned by split() may be less than n_splits if the conflict graph is sparse enough to colour with fewer colours. It will not exceed n_splits by construction (the radius choice guarantees max degree <= n_splits - 1, and greedy colouring uses at most max_degree + 1 colours).

__init__(radius=None, n_splits=None)[source]

Methods

__init__([radius, n_splits])

get_metadata_routing()

Get metadata routing of this object.

get_n_splits([X, y, groups])

get_params([deep])

Get parameters for this estimator.

set_params(**params)

Set the parameters of this estimator.

set_split_request(*[, groups])

Configure whether metadata should be requested to be passed to the split method.

split(X[, y, groups])

Yield (train_indices, test_indices) for each fold.

get_metadata_routing()[source]

Get metadata routing of this object.

Please check User Guide on how the routing mechanism works.

Returns:

routing – A MetadataRequest encapsulating routing information.

Return type:

MetadataRequest

get_n_splits(X=None, y=None, groups=None)[source]
get_params(deep=True)[source]

Get parameters for this estimator.

Parameters:
deep : bool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:

params – Parameter names mapped to their values.

Return type:

dict

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**params : dict

Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

set_split_request(*, groups='$UNCHANGED$')[source]

Configure whether metadata should be requested to be passed to the split method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to split if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to split.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
groups : str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED

Metadata routing for groups parameter in split.

Returns:

self – The updated object.

Return type:

object

split(X, y=None, groups=None)[source]

Yield (train_indices, test_indices) for each fold.

Parameters:
X : GeoDataFrame | GeoSeries | (n, 2) ndarray

Locations.

Yields:
  • train (ndarray of int)

  • test (ndarray of int)