geovalidate.area_of_applicability¶
-
geovalidate.area_of_applicability(X_test, X_train, *, y_train=
None, model=None, cv=None, feature_weights='permutation', metric='euclidean', threshold='tukey', permutation_kwargs=None, return_diagnostics=False)[source]¶ Estimate the Area of Applicability (AOA) for
X_testrelative toX_train.For each test sample, computes the feature-weighted distance to its nearest training sample, normalises by the mean within-training distance to obtain a Dissimilarity Index (DI), and flags the sample as “applicable” if its DI does not exceed a cut-off calibrated from the distribution of training DIs. Method introduced by Meyer and Pebesma [2021].
- Parameters:¶
- X_test : (n_test, n_features) array¶
Feature matrix for points where the model would be applied.
- X_train : (n_train, n_features) array¶
Feature matrix used to train the model.
- y_train : (n_train,) array, optional¶
Training targets. Required when
feature_weights='permutation'.- model : fitted sklearn estimator, optional¶
Required when
feature_weights='permutation'; passed tosklearn.inspection.permutation_importance.- cv : sklearn CV splitter, optional¶
If provided, the training-DI distribution is computed by holding out each fold in turn – the held-out points get their distance to the in-fold training points. If None, each training point’s DI is its distance to its nearest OTHER training point.
- feature_weights : {'permutation', 'uniform'} or (n_features,) array¶
How to weight features when computing distances.
’permutation’ : sklearn permutation importance, normalised to sum to 1. Negative importances are clipped to zero.
’uniform’ / False / None : all features weighted equally.
array : pre-computed non-negative weights; normalised to sum to 1 internally.
- metric : str, default 'euclidean'¶
Passed to
sklearn.metrics.pairwise_distances.- threshold : {'tukey', 'mad'} or float in (0, 1)¶
Rule for converting the training-DI distribution into a single cut-off.
’tukey’ : 75th percentile + 1.5 * IQR.
’mad’ : median + 3 * median absolute deviation.
float p : the p-quantile of the training DIs.
- permutation_kwargs : dict, optional¶
Extra arguments forwarded to
sklearn.inspection.permutation_importance. Defaults setn_repeats=10andn_jobs=-1if not overridden.- return_diagnostics : bool, default False¶
If False, returns the boolean
applicablemask (cheap path: only the nearest training point per test sample is computed). If True, returns asklearn.utils.Bunchwithapplicable,dissimilarity_index,cutpoint,feature_weights, andlpd(Local Point Density). The diagnostics path computes the full (n_test, n_train) distance matrix, so memory cost is O(n_test * n_train).
- Returns:¶
applicable ((n_test,) bool ndarray) – True where the model is considered applicable.
OR a Bunch with attributes
applicable,dissimilarity_index,cutpoint,feature_weights, andlpdwhenreturn_diagnostics=True.
References
[meyer2021]Meyer, H. & Pebesma, E. (2021). Predicting into unknown space? Estimating the area of applicability of spatial prediction models. Methods in Ecology and Evolution, 12(9), 1620-1633. https://doi.org/10.1111/2041-210X.13650