geovalidate.area_of_applicability

geovalidate.area_of_applicability(X_test, X_train, *, y_train=None, model=None, cv=None, feature_weights='permutation', metric='euclidean', threshold='tukey', permutation_kwargs=None, return_diagnostics=False)[source]

Estimate the Area of Applicability (AOA) for X_test relative to X_train.

For each test sample, computes the feature-weighted distance to its nearest training sample, normalises by the mean within-training distance to obtain a Dissimilarity Index (DI), and flags the sample as “applicable” if its DI does not exceed a cut-off calibrated from the distribution of training DIs. Method introduced by Meyer and Pebesma [2021].

Parameters:
X_test : (n_test, n_features) array

Feature matrix for points where the model would be applied.

X_train : (n_train, n_features) array

Feature matrix used to train the model.

y_train : (n_train,) array, optional

Training targets. Required when feature_weights='permutation'.

model : fitted sklearn estimator, optional

Required when feature_weights='permutation'; passed to sklearn.inspection.permutation_importance.

cv : sklearn CV splitter, optional

If provided, the training-DI distribution is computed by holding out each fold in turn – the held-out points get their distance to the in-fold training points. If None, each training point’s DI is its distance to its nearest OTHER training point.

feature_weights : {'permutation', 'uniform'} or (n_features,) array

How to weight features when computing distances.

  • ’permutation’ : sklearn permutation importance, normalised to sum to 1. Negative importances are clipped to zero.

  • ’uniform’ / False / None : all features weighted equally.

  • array : pre-computed non-negative weights; normalised to sum to 1 internally.

metric : str, default 'euclidean'

Passed to sklearn.metrics.pairwise_distances.

threshold : {'tukey', 'mad'} or float in (0, 1)

Rule for converting the training-DI distribution into a single cut-off.

  • ’tukey’ : 75th percentile + 1.5 * IQR.

  • ’mad’ : median + 3 * median absolute deviation.

  • float p : the p-quantile of the training DIs.

permutation_kwargs : dict, optional

Extra arguments forwarded to sklearn.inspection.permutation_importance. Defaults set n_repeats=10 and n_jobs=-1 if not overridden.

return_diagnostics : bool, default False

If False, returns the boolean applicable mask (cheap path: only the nearest training point per test sample is computed). If True, returns a sklearn.utils.Bunch with applicable, dissimilarity_index, cutpoint, feature_weights, and lpd (Local Point Density). The diagnostics path computes the full (n_test, n_train) distance matrix, so memory cost is O(n_test * n_train).

Returns:

  • applicable ((n_test,) bool ndarray) – True where the model is considered applicable.

  • OR a Bunch with attributes applicable, dissimilarity_index,

  • cutpoint, feature_weights, and lpd when

  • return_diagnostics=True.

References

[meyer2021]

Meyer, H. & Pebesma, E. (2021). Predicting into unknown space? Estimating the area of applicability of spatial prediction models. Methods in Ecology and Evolution, 12(9), 1620-1633. https://doi.org/10.1111/2041-210X.13650