Open Lab/Local Modeling · Method

Local Partial Least Squares Regression

Local PLS / LW-PLS

Query-specific KNN-LWPLSR: select nearest calibration neighbors, assign rchemo-style observation weights, and fit weighted PLS regression for each new sample.

ChemometricsPLSLocal PLSLWPLSRKNN-LWPLSRJust-in-TimeRegressionPythonMATLAB

LW-PLS

01

What is Local PLS?

Local PLS is a family of query-specific PLS methods. Instead of fitting one global latent-variable regression on the whole calibration database, a local method constructs a PLS model from calibration objects that are treated as relevant to the sample being predicted.

This card implements one modern member of that family: KNN locally weighted PLS regression (KNN-LWPLSR), as described by Lesnoff, Metz and Roger 2020 and as taught for NIR work by Deryck, Fernandez Pierna, Baeten and Lesnoff 2026. Current rchemo::lwplsr documents itself as that KNN-LWPLSR pipeline. The public Python and MATLAB listings reproduce that pipeline, not every algorithm that has ever been called local PLS.

02

Why local models?

Lesnoff et al. 2020 present LWPLSR as useful when heterogeneity produces nonlinear relations, including curvature and clustering, between the response and the predictors. That is common in agronomic NIR collections that mix materials or origins. Deryck et al. 2026 give the same motivation for large, clustered spectroscopic databases: a single linear global PLS can degrade when the global – structure is no longer approximately linear.

Local modeling is a strategy. It does not prove that every nonlinear relationship is locally linear, and it does not guarantee improved prediction.

03

Query-specific regression

For each new observation , neighbors, observation weights, local latent directions, and regression coefficients can all change. The calibration database is part of the prediction machinery. Fitting one PLS model in fit and reusing it unchanged for every query is not KNN-LWPLSR.

Chemometric local-PLS papers often describe this query-specific construction as just-in-time modeling. This card keeps the Lesnoff / Deryck names LWPLSR and KNN-LWPLSR as the primary identity.

04

A family of methods

Deryck et al. 2026 state that local PLS is a family, not one algorithm. Variants differ in preliminary dimension reduction, neighborhood definition, and within-neighborhood weighting. Names in the literature are not automatic synonyms.

Name on this cardWhat it doesNot the same as
Global PLSROne latent regression on all calibration samplesAny query-specific local model
Uniform local PLS / KNN-PLSSelect k neighbors, equal observation influence, local PLSk-NN regression
LWPLSRWPLSR on the full calibration set with query-dependent observation weightsKNN-LWPLSR (no KNN preselection)
KNN-LWPLSRSelect k neighbors, then WPLSR on that neighborhoodLWPR; Shenk LOCAL; Sicard local PLS1
05

Measuring similarity

There is no scientific notion of nearby without a geometry. Locality depends on preprocessing, the dissimilarity, any dimension reduction used before neighbor search, the neighbor rule, and the weight function. Kim et al. 2013 and Hazama and Kano 2015 emphasize that the similarity definition can dominate local-model performance. Those adaptive or covariance-based similarities are extensions, not part of the public V1 engine.

If Euclidean distance is used in a stated representation,

Euclidean distance is scale sensitive. Current rchemo lwplsr accepts Euclidean or Mahalanobis dissimilarities. Direct Mahalanobis distance in collinear spectral X is often undefined; Deryck et al. note that a reduced score space is usually required. This page does not invent a new regularizer. The educational Mahalanobis option follows rchemo getknn: covariance with divisor n, Cholesky, and rchemo's 1e-5 ridge only if that factorization fails. That ridge is software behavior, not LWPLSR theory.

Current Jchemo also offers cosine, spectral-angle, and correlation dissimilarities. Its correlation distance is , not a unique historical definition of correlation distance. Current rchemo getknn inspected for this card exposes Euclidean and Mahalanobis only. Correlation is therefore documented as a Jchemo option, not as the implemented default.

06

Selecting neighbors

For KNN-LWPLSR the local calibration set is the k nearest eligible calibration samples to the query in the chosen distance space:

k is a hyperparameter. This card does not use a universal rule such as k = 20 or k = sqrt(n). Equal distances are broken by smaller original training index. That tie rule is an implementation detail, not a result in Lesnoff 2020.

07

Uniform local PLS

If the selected neighbors all receive the same observation influence, the method is Deryck et al.'s KNN-PLS: Weighting 1 only. A local PLSR is still fitted. This is not k-NN regression, which predicts from neighbor responses without a latent-variable model.

08

Locally weighted PLS

Lesnoff et al. 2020 define LWPLSR as a particular case of weighted PLS regression (WPLSR): each calibration observation receives a statistical weight different from the ordinary 1/n, and those weights depend on dissimilarity to the query. WPLSR uses those observation weights when calculating PLS scores, loadings, and predictions. It is not ordinary PLS followed by a weighted average of neighbor responses.

Full-database LWPLSR applies WPLSR to all n calibration samples. KNN-LWPLSR first restricts the set. Deryck et al. note that KNN-LWPLS is numerically the same idea as LWPLS with zeros outside the neighborhood, while the KNN step reduces cost.

09

KNN-LWPLSR

Current rchemo locw separates two operations. Weighting 1 is binary neighborhood selection. Weighting 2 is the optional continuous observation weights on selected neighbors. If Weighting 2 is omitted, neighbors are equally weighted. The public implementation is KNN-LWPLSR: both stages, with Weighting 2 given by current rchemo wdist.

10

Weighted PLS mathematics

Observation / locality weight controls the influence of calibration sample i for query q. PLS X-weight is a latent predictor direction. They are different objects.

The displayed weight function is the current rchemo / Jchemo convention used by the 2026 tutorial (Deryck et al., Equation 1), adapted there from Kim et al. 2011. It is not the universal definition of LWPLSR. Centner and Massart 1998 study uniform and cubic weights for locally weighted regression. Those are historical LWR options, not this implementation's kernel.

holds the k neighbor distances. MAD follows R stats::mad: 1.4826 times the median absolute deviation from the median. Current rchemo sets a weight to zero when (implemented as a strict less-than inclusion test; default cri = 4). Deryck et al. describe a 4-MAD cutoff of the same form. The public code follows the rchemo inequality, not a separately invented threshold. Remaining weights are divided by their maximum so the largest equals 1. If that arithmetic is NaN, rchemo replaces the vector by ones. Current rchemo predict.Lwplsr then floors weights below 1e-5. Large h flattens the decay toward equal neighbor weights when the cutoff does not intervene. Deryck et al. identify h = infinity with KNN-PLS inside the neighborhood.

Inside WPLSR, rchemo first maps observation weights to that sum to 1, then uses weighted means

The educational engine is rchemo plskern: Dayal and MacGregor improved kernel #1 on the weighted-centered local matrices, with local X scaling off. The weighted cross-product is after that centering. Prediction uses the local intercept and coefficient vector so that

Different queries can produce different . There is not necessarily one global coefficient or VIP vector for the complete predictor. This card does not average local VIPs.

11

Algorithm

Algorithm 1

KNN-LWPLSR

InputCalibration X, y; query x_q; k; A; distance space; h; cri

OutputPredicted response y-hat_q

  1. 01Require calibration , query , , , distance space, ,
  2. 02Fit any training-only preprocessing and optional global PLS score map ().
  3. 03Represent calibration samples and in the selected distance space.
  4. 04Compute dissimilarities to eligible calibration samples.
  5. 05Select (Weighting 1: neighbors included, others excluded).
  6. 06Compute rchemo observation weights among those neighbors (Weighting 2), or use equal weights.
  7. 07Fit WPLSR ( components) on the weighted local .
  8. 08Predict with the local intercept and coefficients.
  9. 09Repeat independently for every new query.

Special cases: uniform weights give KNN-PLS; k = n with uniform weights recovers global PLSR under the same centering and plskern conventions; nlv = 0 predicts the weighted neighbor mean of y (Deryck et al. 2026).

12

Choosing the distance space

Current rchemo can compute dissimilarities in X or after a global PLS with components. If that argument is 0, there is no such reduction. Deryck et al. stress that the subsequent local PLS is still run on the original local (X, y), not on the score map. Using global PLS scores for neighbor search makes Y influence the geometry. That global PLS must be fitted inside each training fold.

Shen et al. 2019 study local PLS based on global PLS scores as a distinct published variant. It is optional here, not the default.

13

Choosing the number of neighbors

Small k is more local and can be unstable. Large k uses more of the database and moves toward a more global neighborhood. Lesnoff et al. 2020 report that LW and KNN-LW had close prediction performance on their agronomic examples while KNN-LW was faster, and they recommend KNN-LW for large sets. That is not a claim that any k is universally best, and it is not a claim that local methods always beat global PLS.

14

Choosing the weighting strength

In the rchemo weight function, smaller h is sharper and larger h is flatter. k chooses who enters. h chooses relative influence inside that set. They are not the same control. There is no universal h.

15

Choosing latent variables

Each local model uses A latent variables. A is not universal. It interacts with k because a small neighborhood may not support a complex latent model. Deryck et al. tune k, h, A, metric, and on a calibration/validation split of the training data. Averaging predictions over several A values is a later Lesnoff pipeline, not core V1.

16

Prediction of a new sample

Apply training-fitted preprocessing. Map the query into the distance space. Select neighbors among calibration samples. Compute observation weights from those distances. Fit local WPLSR. Predict. The query response is never used. Then repeat for the next query.

A query far from every calibration sample still has poor local support. Local PLS does not solve arbitrary extrapolation. No universal maximum-distance threshold is defined here.

17

Validation without leakage

Validation must reproduce query-specific prediction. For each held-out query the object is absent from the candidate pool, preprocessing and any global PLS score map are training-only, neighbor search uses training objects only, and local WPLSR uses training responses only. In leave-one-out, sample i cannot be its own neighbor. Hyperparameters k, h, A, metric, and are not selected using external-test responses.

Do not compute one full-data neighbor graph if the distance representation itself was learned from the full data, then reuse it inside folds.

18

Local PLS versus global PLS

Global PLSR fits one latent regression for all future queries. Local PLS builds a query-dependent model. Lesnoff et al. 2020 report that both LW strategies overpassed standard PLSR on their three agronomic illustrations. Those results are dataset-specific. Neither method is universally better.

19

Local PLS versus k-NN regression

k-NN regression predicts from neighbor responses, typically by an average. KNN-LWPLSR uses neighbors to fit a local latent-variable regression. Uniform KNN-PLS is therefore not k-NN. Local PLS also differs from ordinary local linear regression by the PLS latent step, which is why it is used for collinear spectra. It is not automatically superior to local linear regression. Kernel PLS builds one kernel latent model. Local PLS builds query-specific local models. Those are different nonlinear strategies.

PropertyGlobal PLSRk-NN regressionKNN-LWPLSR
Model scopeOne global latent modelLocal neighbor responsesQuery-specific local PLS
Use of neighborsNoneDirect response combinationLocal calibration set for WPLSR
Latent-variable modelYesNoYes, local
Query-specific fittingNoYes (neighborhood)Yes (neighborhood, weights, PLS)
Distance dependenceNot for predictionDefines neighborsDefines neighbors and optional weights
Prediction costOne stored modelNeighbor searchNeighbor search plus a local PLS per query

LWPR, Locally Weighted Projection Regression (Vijayakumar, Schaal, and related work), is an incremental receptive-field algorithm. It is not the chemometric LWPLSR / KNN-LWPLSR pipeline. This card never uses LWPR as an abbreviation for locally weighted PLS regression. The NIR LOCAL algorithm of Shenk, Westerhaus and Berzaghi is a specific local calibration strategy, not a generic name for every local PLS method.

20

Python

The listing is an educational KNN-LWPLSR matching current rchemo: store the calibration database in fit; for each query select k neighbors, compute wdist weights, fit weighted plskern, and predict. The constructor default h=2 is an educational convenience matching current Jchemo winvs, not a universal scientific value. sklearn PLSRegression is not LWPLSR and is not assumed to implement observation weights.

import numpy as np  R_MAD_CONSTANT = 1.4826WEIGHT_FLOOR = 1e-5  def _validate_xy(X, y):    X = np.asarray(X, dtype=float)    y = np.asarray(y, dtype=float)    if X.ndim != 2:        raise ValueError("X must be a 2D array of samples by variables.")    if y.ndim == 2:        if y.shape[1] != 1:            raise ValueError("This educational implementation supports a univariate response.")        y = y[:, 0]    else:        y = y.reshape(-1)    if X.shape[0] != y.shape[0]:        raise ValueError("X and y must have the same number of observations.")    if min(X.shape) == 0:        raise ValueError("X must have at least one row and one column.")    if not np.all(np.isfinite(X)) or not np.all(np.isfinite(y)):        raise ValueError("X and y must contain only finite values.")    return X, y  def r_mad(d):    d = np.asarray(d, dtype=float).reshape(-1)    center = np.median(d)    return R_MAD_CONSTANT * np.median(np.abs(d - center))  def mweights(weights):    weights = np.asarray(weights, dtype=float).reshape(-1)    total = np.sum(weights)    if not np.isfinite(total) or total <= 0:        raise ValueError("Observation weights must have a positive finite sum.")    return weights / total  def weighted_colmeans(X, weights):    X = np.asarray(X, dtype=float)    delta = mweights(weights)    return delta @ X  def euclidean_distances(X, z):    X = np.asarray(X, dtype=float)    z = np.asarray(z, dtype=float).reshape(-1)    if z.shape[0] != X.shape[1]:        raise ValueError("Query z must have one value per predictor.")    if not np.all(np.isfinite(z)):        raise ValueError("Query z must contain only finite values.")    return np.sqrt(np.sum((X - z) ** 2, axis=1))  def mahalanobis_whiten(X_train):    X_train = np.asarray(X_train, dtype=float)    n, p = X_train.shape    sigma = np.cov(X_train, rowvar=False, ddof=0)    if p == 1:        sigma = np.array([[float(np.var(X_train[:, 0], ddof=0))]])    ridge = np.zeros((p, p))    try:        upper = np.linalg.cholesky(sigma).T    except np.linalg.LinAlgError:        ridge = 1e-5 * np.eye(p)        upper = np.linalg.cholesky(sigma + ridge).T    try:        uinv = np.linalg.inv(upper)    except np.linalg.LinAlgError:        uinv = np.linalg.inv(upper + 1e-5 * np.eye(p))    return X_train @ uinv, uinv  def knn_select(X_ref, z, k, distance="euclidean", uinv=None):    X_ref = np.asarray(X_ref, dtype=float)    n = X_ref.shape[0]    k = int(k)    if k < 1 or k > n:        raise ValueError("k must be an integer between 1 and n inclusive.")    if distance == "euclidean":        d = euclidean_distances(X_ref, z)    elif distance == "mahalanobis":        if uinv is None:            raise ValueError("Mahalanobis neighbor search requires a training whitening matrix.")        z = np.asarray(z, dtype=float).reshape(-1)        d = euclidean_distances(X_ref, z @ uinv)    else:        raise ValueError("distance must be 'euclidean' or 'mahalanobis'.")    order = np.lexsort((np.arange(n), d))    idx = order[:k]    return idx, d[idx]  def wdist_rchemo(d, h, cri=4.0, squared=False):    d = np.asarray(d, dtype=float).reshape(-1)    if d.size == 0:        raise ValueError("Distance vector must be non-empty.")    h = float(h)    if not np.isfinite(h) and not np.isinf(h):        raise ValueError("h must be a positive scalar or infinity.")    if h <= 0:        raise ValueError("h must be positive.")    if squared:        d = d ** 2    zmed = np.median(d)    zmad = r_mad(d)    inside = d < (zmed + cri * zmad)    with np.errstate(divide="ignore", invalid="ignore"):        w = np.where(inside, np.exp(-d / (h * zmad)), 0.0)        wmax = np.max(w)        w = w / wmax    w[np.isnan(w)] = 1.0    return w  def apply_weight_floor(w, floor=WEIGHT_FLOOR):    w = np.asarray(w, dtype=float).reshape(-1).copy()    w[w < floor] = floor    return w  def plskern_wpls(X, y, weights, n_components):    X = np.asarray(X, dtype=float)    y = np.asarray(y, dtype=float).reshape(-1, 1)    n, p = X.shape    n_components = int(n_components)    if n_components < 0:        raise ValueError("n_components must be a non-negative integer.")    if n_components > min(n, p):        raise ValueError("n_components exceeds min(n_local, n_features).")    delta = mweights(weights)    xmeans = delta @ X    ymeans = float(delta @ y[:, 0])    Xc = X - xmeans    Yc = y - ymeans    if n_components == 0:        return {            "R": np.zeros((p, 0)),            "C": np.zeros((1, 0)),            "W": np.zeros((p, 0)),            "P": np.zeros((p, 0)),            "T": np.zeros((n, 0)),            "xmeans": xmeans,            "ymeans": ymeans,            "weights": delta,        }    Xd = delta[:, None] * Xc    tXY = Xd.T @ Yc    T = np.zeros((n, n_components))    R = np.zeros((p, n_components))    W = np.zeros((p, n_components))    P = np.zeros((p, n_components))    C = np.zeros((1, n_components))    for a in range(n_components):        w = tXY.copy()        nrm = np.sqrt(np.sum(w * w))        if not np.isfinite(nrm) or nrm == 0:            raise ValueError("Weighted PLS extracted a zero X-weight vector.")        w = w / nrm        r = w.copy()        if a > 0:            for j in range(a):                r = r - float(np.sum(P[:, j] * w[:, 0])) * R[:, [j]]        t = Xc @ r        tt = float(np.sum(delta * t[:, 0] * t[:, 0]))        if not np.isfinite(tt) or tt == 0:            raise ValueError("Weighted PLS score variance is zero for a requested component.")        c = (tXY.T @ r) / tt        zp = (Xd.T @ t) / tt        tXY = tXY - (zp @ c) * tt        T[:, a] = t[:, 0]        P[:, a] = zp[:, 0]        W[:, a] = w[:, 0]        R[:, a] = r[:, 0]        C[:, a] = c[0, 0]    return {        "R": R,        "C": C,        "W": W,        "P": P,        "T": T,        "xmeans": xmeans,        "ymeans": ymeans,        "weights": delta,    }  def predict_plskern(model, X_new):    X_new = np.asarray(X_new, dtype=float)    if model["R"].shape[1] == 0:        return np.full(X_new.shape[0], model["ymeans"])    B = model["R"] @ model["C"].T    intercept = model["ymeans"] - model["xmeans"] @ B[:, 0]    return intercept + X_new @ B[:, 0]  class KNNLWPLSR:    def __init__(        self,        n_neighbors,        n_components,        distance="euclidean",        nlvdis=0,        h=2.0,        cri=4.0,        weighting="rchemo",    ):        self.n_neighbors = int(n_neighbors)        self.n_components = int(n_components)        self.distance = distance        self.nlvdis = int(nlvdis)        self.h = float(h) if not np.isinf(h) else np.inf        self.cri = float(cri)        self.weighting = weighting        if self.distance not in ("euclidean", "mahalanobis"):            raise ValueError("distance must be 'euclidean' or 'mahalanobis'.")        if self.weighting not in ("rchemo", "uniform"):            raise ValueError("weighting must be 'rchemo' or 'uniform'.")        if self.nlvdis < 0:            raise ValueError("nlvdis must be a non-negative integer.")        if self.cri <= 0:            raise ValueError("cri must be positive.")     def fit(self, X, y):        X, y = _validate_xy(X, y)        if self.n_neighbors < 1 or self.n_neighbors > X.shape[0]:            raise ValueError("n_neighbors must be between 1 and n inclusive.")        self.X_ = X        self.y_ = y        self.uinv_ = None        self.global_pls_ = None        if self.nlvdis > 0:            ones = np.ones(X.shape[0])            self.global_pls_ = plskern_wpls(X, y, ones, self.nlvdis)        elif self.distance == "mahalanobis":            _whitened, self.uinv_ = mahalanobis_whiten(X)        return self     def _reference_space(self, X_query):        if self.global_pls_ is None:            if self.distance == "mahalanobis":                return self.X_ @ self.uinv_, X_query @ self.uinv_, "euclidean", None            return self.X_, X_query, self.distance, self.uinv_        T_ref = (self.X_ - self.global_pls_["xmeans"]) @ self.global_pls_["R"]        T_q = (X_query - self.global_pls_["xmeans"]) @ self.global_pls_["R"]        uinv = None        if self.distance == "mahalanobis":            T_ref, uinv = mahalanobis_whiten(T_ref)            T_q = T_q @ uinv            return T_ref, T_q, "euclidean", None        return T_ref, T_q, "euclidean", None     def query_local(self, z):        z = np.asarray(z, dtype=float).reshape(1, -1)        X_ref, Z_q, distance, uinv = self._reference_space(z)        idx, d = knn_select(X_ref, Z_q[0], self.n_neighbors, distance=distance, uinv=uinv)        if self.weighting == "uniform":            w = np.ones(idx.shape[0])        else:            w = apply_weight_floor(wdist_rchemo(d, h=self.h, cri=self.cri))        X_loc = self.X_[idx]        y_loc = self.y_[idx]        local = plskern_wpls(X_loc, y_loc, w, self.n_components)        yhat = float(predict_plskern(local, z)[0])        return {            "indices": idx,            "distances": d,            "observation_weights": w,            "yhat": yhat,            "local_model": local,        }     def predict(self, X_new):        X_new = np.asarray(X_new, dtype=float)        if X_new.ndim != 2:            raise ValueError("X_new must be a 2D array of samples by variables.")        if X_new.shape[1] != self.X_.shape[1]:            raise ValueError("X_new must have the same number of variables as X.")        if X_new.shape[0] > 0 and not np.all(np.isfinite(X_new)):            raise ValueError("X_new must contain only finite values.")        return np.array([self.query_local(row)["yhat"] for row in X_new], dtype=float)  def loo_candidate_mask(n, i):    mask = np.ones(n, dtype=bool)    mask[i] = False    return mask 
21

MATLAB

MATLAB has no built-in LWPLSR. The listing implements the same WPLSR and rchemo-style weights. Do not pass observation weights to plsregress and call the result KNN-LWPLSR.

function model = fit_knnlwplsr(X, y, nNeighbors, nComponents, distance, nlvdis, h, cri, weighting)    if nargin < 5, distance = "euclidean"; end    if nargin < 6, nlvdis = 0; end    if nargin < 7, h = 2; end    if nargin < 8, cri = 4; end    if nargin < 9, weighting = "rchemo"; end    [X, y] = validateXY_lwpls(X, y);    n = size(X, 1);    nNeighbors = double(nNeighbors);    nComponents = double(nComponents);    nlvdis = double(nlvdis);    if nNeighbors < 1 || nNeighbors ~= floor(nNeighbors) || nNeighbors > n        error('nNeighbors must be an integer between 1 and n inclusive.');    end    if nComponents < 0 || nComponents ~= floor(nComponents)        error('nComponents must be a non-negative integer.');    end    model.X = X;    model.y = y;    model.nNeighbors = nNeighbors;    model.nComponents = nComponents;    model.distance = char(distance);    model.nlvdis = nlvdis;    model.h = h;    model.cri = cri;    model.weighting = char(weighting);    model.uinv = [];    model.globalPls = [];    if nlvdis > 0        model.globalPls = plskern_wpls(X, y, ones(n, 1), nlvdis);    elseif strcmp(model.distance, "mahalanobis")        [~, model.uinv] = mahalanobisWhiten_lwpls(X);    endend function yhat = predict_knnlwplsr(model, Xnew)    Xnew = validateXnew_lwpls(Xnew, size(model.X, 2));    m = size(Xnew, 1);    yhat = zeros(m, 1);    for i = 1:m        q = query_knnlwplsr(model, Xnew(i, :));        yhat(i) = q.yhat;    endend function q = query_knnlwplsr(model, z)    z = z(:)';    [Xref, zq, distance, uinv] = referenceSpace_lwpls(model, z);    [idx, d] = knnSelect_lwpls(Xref, zq, model.nNeighbors, distance, uinv);    if strcmp(model.weighting, "uniform")        w = ones(numel(idx), 1);    else        w = applyWeightFloor_lwpls(wdist_rchemo(d, model.h, model.cri));    end    local = plskern_wpls(model.X(idx, :), model.y(idx), w, model.nComponents);    yhat = predict_plskern(local, z);    q.indices = idx;    q.distances = d;    q.observation_weights = w;    q.yhat = yhat;    q.local_model = local;end function w = wdist_rchemo(d, h, cri)    d = d(:);    if nargin < 3, cri = 4; end    zmed = median(d);    zmad = 1.4826 * median(abs(d - zmed));    inside = d < (zmed + cri * zmad);    w = zeros(size(d));    w(inside) = exp(-d(inside) ./ (h * zmad));    wmax = max(w);    w = w / wmax;    w(isnan(w)) = 1;end function w = applyWeightFloor_lwpls(w)    w = w(:);    w(w < 1e-5) = 1e-5;end function delta = mweights_lwpls(weights)    weights = weights(:);    total = sum(weights);    if ~isfinite(total) || total <= 0        error('Observation weights must have a positive finite sum.');    end    delta = weights / total;end function fm = plskern_wpls(X, y, weights, nComponents)    n = size(X, 1);    p = size(X, 2);    nComponents = double(nComponents);    if nComponents > min(n, p)        error('nComponents exceeds min(n_local, n_features).');    end    delta = mweights_lwpls(weights);    xmeans = delta' * X;    ymeans = delta' * y(:);    Xc = X - xmeans;    Yc = y(:) - ymeans;    if nComponents == 0        fm.R = zeros(p, 0);        fm.C = zeros(1, 0);        fm.W = zeros(p, 0);        fm.P = zeros(p, 0);        fm.T = zeros(n, 0);        fm.xmeans = xmeans;        fm.ymeans = ymeans;        fm.weights = delta;        return    end    Xd = delta .* Xc;    tXY = Xd' * Yc;    T = zeros(n, nComponents);    R = zeros(p, nComponents);    Wmat = zeros(p, nComponents);    P = zeros(p, nComponents);    C = zeros(1, nComponents);    for a = 1:nComponents        w = tXY;        nrm = sqrt(sum(w .* w));        if ~isfinite(nrm) || nrm == 0            error('Weighted PLS extracted a zero X-weight vector.');        end        w = w / nrm;        r = w;        if a > 1            for j = 1:(a - 1)                r = r - sum(P(:, j) .* w) * R(:, j);            end        end        t = Xc * r;        tt = sum(delta .* t .* t);        if ~isfinite(tt) || tt == 0            error('Weighted PLS score variance is zero for a requested component.');        end        c = (tXY' * r) / tt;        zp = (Xd' * t) / tt;        tXY = tXY - (zp * c) * tt;        T(:, a) = t;        P(:, a) = zp;        Wmat(:, a) = w;        R(:, a) = r;        C(:, a) = c;    end    fm.R = R;    fm.C = C;    fm.W = Wmat;    fm.P = P;    fm.T = T;    fm.xmeans = xmeans;    fm.ymeans = ymeans;    fm.weights = delta;end function yhat = predict_plskern(fm, Xnew)    Xnew = Xnew(:, :);    if size(fm.R, 2) == 0        yhat = fm.ymeans * ones(size(Xnew, 1), 1);        return    end    B = fm.R * fm.C';    intercept = fm.ymeans - fm.xmeans * B;    yhat = intercept + Xnew * B;end function [idx, dsel] = knnSelect_lwpls(Xref, z, k, distance, uinv)    n = size(Xref, 1);    z = z(:)';    if strcmp(distance, "mahalanobis")        z = z * uinv;        d = sqrt(sum((Xref - z).^2, 2));    else        d = sqrt(sum((Xref - z).^2, 2));    end    [~, order] = sortrows([d, (1:n)']);    idx = order(1:k);    dsel = d(idx);end function [Xw, uinv] = mahalanobisWhiten_lwpls(Xtrain)    n = size(Xtrain, 1);    p = size(Xtrain, 2);    sigma = cov(Xtrain, 1);    if p == 1        sigma = var(Xtrain, 1);    end    [U, flag] = chol(sigma);    if flag ~= 0        U = chol(sigma + 1e-5 * eye(p));    end    uinv = inv(U);    Xw = Xtrain * uinv;end function [Xref, zq, distance, uinv] = referenceSpace_lwpls(model, z)    if isempty(model.globalPls)        if strcmp(model.distance, "mahalanobis")            Xref = model.X * model.uinv;            zq = z * model.uinv;            distance = "euclidean";            uinv = [];        else            Xref = model.X;            zq = z;            distance = model.distance;            uinv = model.uinv;        end        return    end    Tref = (model.X - model.globalPls.xmeans) * model.globalPls.R;    Tq = (z - model.globalPls.xmeans) * model.globalPls.R;    if strcmp(model.distance, "mahalanobis")        [Tref, uinv] = mahalanobisWhiten_lwpls(Tref);        Xref = Tref;        zq = Tq * uinv;        distance = "euclidean";        uinv = [];    else        Xref = Tref;        zq = Tq;        distance = "euclidean";        uinv = [];    endend function [X, y] = validateXY_lwpls(X, y)    y = y(:);    if size(X, 1) ~= numel(y)        error('X and y must have the same number of observations.');    end    if ~all(isfinite(X(:))) || ~all(isfinite(y))        error('X and y must contain only finite values.');    endend function Xnew = validateXnew_lwpls(Xnew, p)    if size(Xnew, 2) ~= p        error('Xnew must have the same number of variables as X.');    end    if ~isempty(Xnew) && ~all(isfinite(Xnew(:)))        error('Xnew must contain only finite values.');    endend 
22

Practical notes

  • Local PLS is a family. The public engine is KNN-LWPLSR, not every local PLS variant.
  • Neighbor selection and within-neighborhood observation weighting are different operations.
  • The distance metric and any scaling define local geometry. They are model choices, not universal requirements.
  • SNV, derivatives, or autoscaling used in the 2026 tutorial example are not part of the method definition.
  • k, h, and A have no universal defaults. Tune them without external-test responses.
  • Small neighborhoods can be unstable or rank-limited.
  • A query outside calibration coverage remains poorly supported.
  • Prediction rebuilds a local PLS model and is heavier than applying one global PLSR.
  • Duplicate calibration spectra can dominate a neighborhood. Duplicate curation in the 2026 example is database quality, not a mandatory algorithm step.
  • Use the Open Lab Regression Metrics card. This page does not redefine RMSE, MAE, R2, or bias.
23

Extensions

Kim et al. 2013: adaptive similarity for JIT LW-PLS, not in V1. Hazama and Kano 2015: covariance-based similarity that uses the X–Y relationship, not Euclidean/Mahalanobis X geometry alone; not implemented. Shen et al. 2019: locality in a global PLS score space; available here only as optional . Lesnoff and collaborators: averaging local PLSR predictions over several dimensionalities; not core V1. LW-PLS-DA applies the local idea to classification and is a separate Open Lab card when published. Sicard and Sabatier 2006 develop a local PLS1 framework with multivariate local polynomials. It is related literature, not this KNN-LWPLSR engine.

24

References

  1. 1.

    Lesnoff, M., Metz, M., & Roger, J.-M. (2020). Comparison of locally weighted PLS strategies for regression and discrimination on agronomic NIR data. Journal of Chemometrics, 34(5), e3209.

    doi:10.1002/cem.3209
  2. 2.

    Deryck, A., Fernandez Pierna, J. A., Baeten, V., & Lesnoff, M. (2026). A step-by-step tutorial to local partial least squares analyses for near-infrared spectroscopic data analysis. Analytica Chimica Acta, 1401, 345242.

    doi:10.1016/j.aca.2026.345242
  3. 3.

    Centner, V., & Massart, D. L. (1998). Optimization in locally weighted regression. Analytical Chemistry, 70(19), 4206-4211.

    doi:10.1021/ac980208r
  4. 4.

    Sicard, E., & Sabatier, R. (2006). Theoretical framework for local PLS1 regression, and application to a rainfall data set. Computational Statistics & Data Analysis, 51(2), 1393-1410.

    doi:10.1016/j.csda.2006.05.002
  5. 5.

    Shen, G., Lesnoff, M., Baeten, V., Dardenne, P., Davrieux, F., Ceballos, H., Belalcazar, J., Dufour, D., Yang, Z., Han, L., & Fernandez Pierna, J. A. (2019). Local partial least squares based on global PLS scores. Journal of Chemometrics, 33(5), e3117.

    doi:10.1002/cem.3117
  6. 6.

    Kim, S., Okajima, R., Kano, M., & Hasebe, S. (2013). Development of soft-sensor using locally weighted PLS with adaptive similarity measure. Chemometrics and Intelligent Laboratory Systems, 124, 43-49.

    doi:10.1016/j.chemolab.2013.03.008
  7. 7.

    Hazama, K., & Kano, M. (2015). Covariance-based locally weighted partial least squares for high-performance adaptive modeling. Chemometrics and Intelligent Laboratory Systems, 146, 55-62.

    doi:10.1016/j.chemolab.2015.05.007
  8. 8.

    Dayal, B. S., & MacGregor, J. F. (1997). Improved PLS algorithms. Journal of Chemometrics, 11(1), 73-85.

    doi:10.1002/(SICI)1099-128X(199701)11:1<73::AID-CEM435>3.0.CO;2-#
  9. 9.

    Kim, S., Kano, M., Nakagawa, H., & Hasebe, S. (2011). Estimation of active pharmaceutical ingredients content using locally weighted partial least squares and statistical wavelength selection. International Journal of Pharmaceutics, 421(2), 269-274.

    doi:10.1016/j.ijpharm.2011.10.007
  10. 10.

    Shenk, J., Westerhaus, M., & Berzaghi, P. (1997). Investigation of a LOCAL calibration procedure for near infrared instruments. Journal of Near Infrared Spectroscopy, 5(1), 223-232.

    doi:10.1255/jnirs.115
  11. 11.

    ChemHouse-group (n.d.). rchemo: lwplsr, locw, wdist, plskern, getknn. R package source inspected from ChemHouse-group/rchemo.

  12. 12.

    Lesnoff, M. (n.d.). Jchemo.jl: lwplsr, winvs, getknn. Julia package source inspected from mlesnoff/Jchemo.jl.