geovalidate.CellStratifiedKFold¶
-
class geovalidate.CellStratifiedKFold(n_splits=
5, grid='h3', resolution=None, shuffle=False, random_state=None)[source]¶ Stratified k-fold cross-validator using a discrete global grid system.
Each observation is indexed to a DGGS cell. Cell membership becomes the stratum label for a standard StratifiedKFold split, so every fold receives a proportional mix of observations from across the study area.
- Parameters:¶
- n_splits : int, default 5¶
Number of folds.
- grid : {"h3", "a5", "healpix", "s2"}, default "h3"¶
DGGS backend. The corresponding package must be installed (
h3,pya5,healpy,s2sphere).- resolution : int or None, default None¶
Grid resolution. Meaning varies by backend:
h3: 0 (coarsest) – 15 (finest)
a5: 0 (coarsest) – 30 (finest)
s2: 0 (coarsest) – 30 (finest)
healpix: log2(nside), so 0 = nside 1, 1 = nside 2, etc.
If None, the coarsest resolution with at least n_splits occupied cells is chosen automatically and stored as
resolution_after callingsplit().- shuffle : bool, default False¶
- random_state : int or None, default None¶
Notes
Cells with fewer members than n_splits are handled by sklearn’s StratifiedKFold (the minority observations land in fewer folds). If this causes an error, increase resolution or reduce n_splits.
When
shuffle=False(the default), observations within each cell are assigned to folds in their original row order. If the dataset has a spatial sort order (common in GeoDataFrames), this can produce visually clustered fold maps even though the stratification is correct. Passshuffle=True, random_state=<int>for reproducible, spatially uniform within-cell fold assignment.Examples
>>> cv = CellStratifiedKFold(n_splits=5, grid="h3") >>> for train, test in cv.split(gdf): ... model.fit(X[train], y[train]) ... score = model.score(X[test], y[test])-
__init__(n_splits=
5, grid='h3', resolution=None, shuffle=False, random_state=None)[source]¶
Methods
__init__([n_splits, grid, resolution, ...])Get metadata routing of this object.
get_n_splits([X, y, groups])get_params([deep])Get parameters for this estimator.
set_params(**params)Set the parameters of this estimator.
set_split_request(*[, groups])Configure whether metadata should be requested to be passed to the
splitmethod.split(X[, y, groups])Yield
(train_indices, test_indices)for each fold.- get_metadata_routing()[source]¶
Get metadata routing of this object.
Please check User Guide on how the routing mechanism works.
- Returns:¶
routing – A
MetadataRequestencapsulating routing information.- Return type:¶
MetadataRequest
- set_params(**params)[source]¶
Set the parameters of this estimator.
The method works on simple estimators as well as on nested objects (such as
Pipeline). The latter have parameters of the form<component>__<parameter>so that it’s possible to update each component of a nested object.
-
set_split_request(*, groups=
'$UNCHANGED$')[source]¶ Configure whether metadata should be requested to be passed to the
splitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed tosplitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it tosplit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.