geovalidate.gearygram

geovalidate.gearygram(X, geometry=None, n_bins=15, max_distance=None, max_k=None, kernel=None, nonparametric=False)[source]

Compute a Geary correlogram for univariate or multivariate spatial data.

Three modes:

Bandwidth (default, n_bins): evaluates the weighted Geary’s C at n_bins increasing kernel bandwidths. At each bandwidth h every pair (i, j) receives weight \(K(d_{ij}/h)\). Compact kernels use a sparse distance matrix; non-compact kernels compute all pairwise distances. Defaults to kernel="gaussian".

kNN (max_k): evaluates Geary’s C for cumulative k-NN graphs at k = 1 … max_k. Each observation’s k-th nearest-neighbour distance is used as an adaptive bandwidth so kernel controls how the k included neighbours are weighted by distance. Defaults to kernel="uniform", which gives equal binary weights (standard kNN Geary).

Nonparametric (nonparametric=True): fits a LOWESS curve of the per-pair Geary contribution \(\|\mathbf{z}_i-\mathbf{z}_j\|^2/(2p)\) against squared spatial distance \(d_{ij}^2\), then evaluates the smooth at n_bins equally spaced distance values.

In all parametric modes the multivariate statistic Anselin [2019] is:

\[C_m = \frac{\sum_{(i,j)} w_{ij}\,\|\mathbf{z}_i - \mathbf{z}_j\|^2} {2\,W\,p}\]

where \(\mathbf{z}\) is column-standardised X, W is the weight sum, and p is the number of variables. For p = 1 this reduces to the standard weighted Geary’s C Geary [1954].

Parameters:
X : array-like of shape (n,) or (n, p)

Observed values. A 1-D array is treated as a single variable. If X is a GeoDataFrame with a geometry column and geometry is not provided, locations are read from that column.

geometry : GeoDataFrame | GeoSeries | (n, 2) ndarray or None

Locations. Required if X has no geometry attribute.

n_bins : int, default 15

Number of bandwidths (bandwidth/nonparametric mode). Ignored for kNN mode.

max_distance : float or None

Maximum bandwidth (bandwidth/nonparametric mode). Defaults to the maximum pairwise distance. Ignored for kNN mode.

max_k : int or None

If set, use kNN mode with k = 1 … max_k.

kernel : str or None

Kernel weighting. One of "gaussian", "exponential", "bisquare", "triangular", "uniform", "parabolic". Defaults to "uniform" for kNN mode and "gaussian" for bandwidth mode. Ignored for nonparametric mode.

nonparametric : bool, default False

If True, fit a LOWESS curve instead of computing kernel-weighted bins. Ignored when max_k is set.

Returns:

bin_centersndarray

Bandwidth values (distance units) or k indices.

Cndarray

Geary’s C per lag. Approaches 1 under spatial independence, < 1 for positive autocorrelation, > 1 for negative.

n_pairsndarray or None

Number of pairs per lag (None for nonparametric mode).

Return type:

sklearn.utils.Bunch

Examples

Bandwidth correlogram (Gaussian kernel):

>>> result = gearygram(Y, gdf)

kNN correlogram with Gaussian kernel weighting:

>>> result = gearygram(Y, gdf, max_k=20, kernel="gaussian")

Nonparametric LOWESS correlogram:

>>> result = gearygram(Y, gdf, nonparametric=True)

Notes

Bandwidth bins with fewer than 2 pairs are returned as NaN.