Open Lab/Local Modeling · Method
Local Partial Least Squares Regression
Local PLS / LW-PLS
Query-specific KNN-LWPLSR: select nearest calibration neighbors, assign rchemo-style observation weights, and fit weighted PLS regression for each new sample.
LW-PLS
What is Local PLS?
Local PLS is a family of query-specific PLS methods. Instead of fitting one global latent-variable regression on the whole calibration database, a local method constructs a PLS model from calibration objects that are treated as relevant to the sample being predicted.
This card implements one modern member of that family: KNN locally weighted PLS regression (KNN-LWPLSR), as described by Lesnoff, Metz and Roger 2020 and as taught for NIR work by Deryck, Fernandez Pierna, Baeten and Lesnoff 2026. Current rchemo::lwplsr documents itself as that KNN-LWPLSR pipeline. The public Python and MATLAB listings reproduce that pipeline, not every algorithm that has ever been called local PLS.
Why local models?
Lesnoff et al. 2020 present LWPLSR as useful when heterogeneity produces nonlinear relations, including curvature and clustering, between the response and the predictors. That is common in agronomic NIR collections that mix materials or origins. Deryck et al. 2026 give the same motivation for large, clustered spectroscopic databases: a single linear global PLS can degrade when the global – structure is no longer approximately linear.
Local modeling is a strategy. It does not prove that every nonlinear relationship is locally linear, and it does not guarantee improved prediction.
Query-specific regression
For each new observation , neighbors, observation weights, local latent directions, and regression coefficients can all change. The calibration database is part of the prediction machinery. Fitting one PLS model in fit and reusing it unchanged for every query is not KNN-LWPLSR.
Chemometric local-PLS papers often describe this query-specific construction as just-in-time modeling. This card keeps the Lesnoff / Deryck names LWPLSR and KNN-LWPLSR as the primary identity.
A family of methods
Deryck et al. 2026 state that local PLS is a family, not one algorithm. Variants differ in preliminary dimension reduction, neighborhood definition, and within-neighborhood weighting. Names in the literature are not automatic synonyms.
| Name on this card | What it does | Not the same as |
|---|---|---|
| Global PLSR | One latent regression on all calibration samples | Any query-specific local model |
| Uniform local PLS / KNN-PLS | Select k neighbors, equal observation influence, local PLS | k-NN regression |
| LWPLSR | WPLSR on the full calibration set with query-dependent observation weights | KNN-LWPLSR (no KNN preselection) |
| KNN-LWPLSR | Select k neighbors, then WPLSR on that neighborhood | LWPR; Shenk LOCAL; Sicard local PLS1 |
Measuring similarity
There is no scientific notion of nearby without a geometry. Locality depends on preprocessing, the dissimilarity, any dimension reduction used before neighbor search, the neighbor rule, and the weight function. Kim et al. 2013 and Hazama and Kano 2015 emphasize that the similarity definition can dominate local-model performance. Those adaptive or covariance-based similarities are extensions, not part of the public V1 engine.
If Euclidean distance is used in a stated representation,
Euclidean distance is scale sensitive. Current rchemo lwplsr accepts Euclidean or Mahalanobis dissimilarities. Direct Mahalanobis distance in collinear spectral X is often undefined; Deryck et al. note that a reduced score space is usually required. This page does not invent a new regularizer. The educational Mahalanobis option follows rchemo getknn: covariance with divisor n, Cholesky, and rchemo's 1e-5 ridge only if that factorization fails. That ridge is software behavior, not LWPLSR theory.
Current Jchemo also offers cosine, spectral-angle, and correlation dissimilarities. Its correlation distance is , not a unique historical definition of correlation distance. Current rchemo getknn inspected for this card exposes Euclidean and Mahalanobis only. Correlation is therefore documented as a Jchemo option, not as the implemented default.
Selecting neighbors
For KNN-LWPLSR the local calibration set is the k nearest eligible calibration samples to the query in the chosen distance space:
k is a hyperparameter. This card does not use a universal rule such as k = 20 or k = sqrt(n). Equal distances are broken by smaller original training index. That tie rule is an implementation detail, not a result in Lesnoff 2020.
Uniform local PLS
If the selected neighbors all receive the same observation influence, the method is Deryck et al.'s KNN-PLS: Weighting 1 only. A local PLSR is still fitted. This is not k-NN regression, which predicts from neighbor responses without a latent-variable model.
Locally weighted PLS
Lesnoff et al. 2020 define LWPLSR as a particular case of weighted PLS regression (WPLSR): each calibration observation receives a statistical weight different from the ordinary 1/n, and those weights depend on dissimilarity to the query. WPLSR uses those observation weights when calculating PLS scores, loadings, and predictions. It is not ordinary PLS followed by a weighted average of neighbor responses.
Full-database LWPLSR applies WPLSR to all n calibration samples. KNN-LWPLSR first restricts the set. Deryck et al. note that KNN-LWPLS is numerically the same idea as LWPLS with zeros outside the neighborhood, while the KNN step reduces cost.
KNN-LWPLSR
Current rchemo locw separates two operations. Weighting 1 is binary neighborhood selection. Weighting 2 is the optional continuous observation weights on selected neighbors. If Weighting 2 is omitted, neighbors are equally weighted. The public implementation is KNN-LWPLSR: both stages, with Weighting 2 given by current rchemo wdist.
Weighted PLS mathematics
Observation / locality weight controls the influence of calibration sample i for query q. PLS X-weight is a latent predictor direction. They are different objects.
The displayed weight function is the current rchemo / Jchemo convention used by the 2026 tutorial (Deryck et al., Equation 1), adapted there from Kim et al. 2011. It is not the universal definition of LWPLSR. Centner and Massart 1998 study uniform and cubic weights for locally weighted regression. Those are historical LWR options, not this implementation's kernel.
holds the k neighbor distances. MAD follows R stats::mad: 1.4826 times the median absolute deviation from the median. Current rchemo sets a weight to zero when (implemented as a strict less-than inclusion test; default cri = 4). Deryck et al. describe a 4-MAD cutoff of the same form. The public code follows the rchemo inequality, not a separately invented threshold. Remaining weights are divided by their maximum so the largest equals 1. If that arithmetic is NaN, rchemo replaces the vector by ones. Current rchemo predict.Lwplsr then floors weights below 1e-5. Large h flattens the decay toward equal neighbor weights when the cutoff does not intervene. Deryck et al. identify h = infinity with KNN-PLS inside the neighborhood.
Inside WPLSR, rchemo first maps observation weights to that sum to 1, then uses weighted means
The educational engine is rchemo plskern: Dayal and MacGregor improved kernel #1 on the weighted-centered local matrices, with local X scaling off. The weighted cross-product is after that centering. Prediction uses the local intercept and coefficient vector so that
Different queries can produce different . There is not necessarily one global coefficient or VIP vector for the complete predictor. This card does not average local VIPs.
Algorithm
KNN-LWPLSR
InputCalibration X, y; query x_q; k; A; distance space; h; cri
OutputPredicted response y-hat_q
- 01Require calibration , query , , , distance space, ,
- 02Fit any training-only preprocessing and optional global PLS score map ().
- 03Represent calibration samples and in the selected distance space.
- 04Compute dissimilarities to eligible calibration samples.
- 05Select (Weighting 1: neighbors included, others excluded).
- 06Compute rchemo observation weights among those neighbors (Weighting 2), or use equal weights.
- 07Fit WPLSR ( components) on the weighted local .
- 08Predict with the local intercept and coefficients.
- 09Repeat independently for every new query.
Special cases: uniform weights give KNN-PLS; k = n with uniform weights recovers global PLSR under the same centering and plskern conventions; nlv = 0 predicts the weighted neighbor mean of y (Deryck et al. 2026).
Choosing the distance space
Current rchemo can compute dissimilarities in X or after a global PLS with components. If that argument is 0, there is no such reduction. Deryck et al. stress that the subsequent local PLS is still run on the original local (X, y), not on the score map. Using global PLS scores for neighbor search makes Y influence the geometry. That global PLS must be fitted inside each training fold.
Shen et al. 2019 study local PLS based on global PLS scores as a distinct published variant. It is optional here, not the default.
Choosing the number of neighbors
Small k is more local and can be unstable. Large k uses more of the database and moves toward a more global neighborhood. Lesnoff et al. 2020 report that LW and KNN-LW had close prediction performance on their agronomic examples while KNN-LW was faster, and they recommend KNN-LW for large sets. That is not a claim that any k is universally best, and it is not a claim that local methods always beat global PLS.
Choosing the weighting strength
In the rchemo weight function, smaller h is sharper and larger h is flatter. k chooses who enters. h chooses relative influence inside that set. They are not the same control. There is no universal h.
Choosing latent variables
Each local model uses A latent variables. A is not universal. It interacts with k because a small neighborhood may not support a complex latent model. Deryck et al. tune k, h, A, metric, and on a calibration/validation split of the training data. Averaging predictions over several A values is a later Lesnoff pipeline, not core V1.
Prediction of a new sample
Apply training-fitted preprocessing. Map the query into the distance space. Select neighbors among calibration samples. Compute observation weights from those distances. Fit local WPLSR. Predict. The query response is never used. Then repeat for the next query.
A query far from every calibration sample still has poor local support. Local PLS does not solve arbitrary extrapolation. No universal maximum-distance threshold is defined here.
Validation without leakage
Validation must reproduce query-specific prediction. For each held-out query the object is absent from the candidate pool, preprocessing and any global PLS score map are training-only, neighbor search uses training objects only, and local WPLSR uses training responses only. In leave-one-out, sample i cannot be its own neighbor. Hyperparameters k, h, A, metric, and are not selected using external-test responses.
Do not compute one full-data neighbor graph if the distance representation itself was learned from the full data, then reuse it inside folds.
Local PLS versus global PLS
Global PLSR fits one latent regression for all future queries. Local PLS builds a query-dependent model. Lesnoff et al. 2020 report that both LW strategies overpassed standard PLSR on their three agronomic illustrations. Those results are dataset-specific. Neither method is universally better.
Local PLS versus k-NN regression
k-NN regression predicts from neighbor responses, typically by an average. KNN-LWPLSR uses neighbors to fit a local latent-variable regression. Uniform KNN-PLS is therefore not k-NN. Local PLS also differs from ordinary local linear regression by the PLS latent step, which is why it is used for collinear spectra. It is not automatically superior to local linear regression. Kernel PLS builds one kernel latent model. Local PLS builds query-specific local models. Those are different nonlinear strategies.
| Property | Global PLSR | k-NN regression | KNN-LWPLSR |
|---|---|---|---|
| Model scope | One global latent model | Local neighbor responses | Query-specific local PLS |
| Use of neighbors | None | Direct response combination | Local calibration set for WPLSR |
| Latent-variable model | Yes | No | Yes, local |
| Query-specific fitting | No | Yes (neighborhood) | Yes (neighborhood, weights, PLS) |
| Distance dependence | Not for prediction | Defines neighbors | Defines neighbors and optional weights |
| Prediction cost | One stored model | Neighbor search | Neighbor search plus a local PLS per query |
LWPR, Locally Weighted Projection Regression (Vijayakumar, Schaal, and related work), is an incremental receptive-field algorithm. It is not the chemometric LWPLSR / KNN-LWPLSR pipeline. This card never uses LWPR as an abbreviation for locally weighted PLS regression. The NIR LOCAL algorithm of Shenk, Westerhaus and Berzaghi is a specific local calibration strategy, not a generic name for every local PLS method.
Python
The listing is an educational KNN-LWPLSR matching current rchemo: store the calibration database in fit; for each query select k neighbors, compute wdist weights, fit weighted plskern, and predict. The constructor default h=2 is an educational convenience matching current Jchemo winvs, not a universal scientific value. sklearn PLSRegression is not LWPLSR and is not assumed to implement observation weights.
import numpy as np R_MAD_CONSTANT = 1.4826WEIGHT_FLOOR = 1e-5 def _validate_xy(X, y): X = np.asarray(X, dtype=float) y = np.asarray(y, dtype=float) if X.ndim != 2: raise ValueError("X must be a 2D array of samples by variables.") if y.ndim == 2: if y.shape[1] != 1: raise ValueError("This educational implementation supports a univariate response.") y = y[:, 0] else: y = y.reshape(-1) if X.shape[0] != y.shape[0]: raise ValueError("X and y must have the same number of observations.") if min(X.shape) == 0: raise ValueError("X must have at least one row and one column.") if not np.all(np.isfinite(X)) or not np.all(np.isfinite(y)): raise ValueError("X and y must contain only finite values.") return X, y def r_mad(d): d = np.asarray(d, dtype=float).reshape(-1) center = np.median(d) return R_MAD_CONSTANT * np.median(np.abs(d - center)) def mweights(weights): weights = np.asarray(weights, dtype=float).reshape(-1) total = np.sum(weights) if not np.isfinite(total) or total <= 0: raise ValueError("Observation weights must have a positive finite sum.") return weights / total def weighted_colmeans(X, weights): X = np.asarray(X, dtype=float) delta = mweights(weights) return delta @ X def euclidean_distances(X, z): X = np.asarray(X, dtype=float) z = np.asarray(z, dtype=float).reshape(-1) if z.shape[0] != X.shape[1]: raise ValueError("Query z must have one value per predictor.") if not np.all(np.isfinite(z)): raise ValueError("Query z must contain only finite values.") return np.sqrt(np.sum((X - z) ** 2, axis=1)) def mahalanobis_whiten(X_train): X_train = np.asarray(X_train, dtype=float) n, p = X_train.shape sigma = np.cov(X_train, rowvar=False, ddof=0) if p == 1: sigma = np.array([[float(np.var(X_train[:, 0], ddof=0))]]) ridge = np.zeros((p, p)) try: upper = np.linalg.cholesky(sigma).T except np.linalg.LinAlgError: ridge = 1e-5 * np.eye(p) upper = np.linalg.cholesky(sigma + ridge).T try: uinv = np.linalg.inv(upper) except np.linalg.LinAlgError: uinv = np.linalg.inv(upper + 1e-5 * np.eye(p)) return X_train @ uinv, uinv def knn_select(X_ref, z, k, distance="euclidean", uinv=None): X_ref = np.asarray(X_ref, dtype=float) n = X_ref.shape[0] k = int(k) if k < 1 or k > n: raise ValueError("k must be an integer between 1 and n inclusive.") if distance == "euclidean": d = euclidean_distances(X_ref, z) elif distance == "mahalanobis": if uinv is None: raise ValueError("Mahalanobis neighbor search requires a training whitening matrix.") z = np.asarray(z, dtype=float).reshape(-1) d = euclidean_distances(X_ref, z @ uinv) else: raise ValueError("distance must be 'euclidean' or 'mahalanobis'.") order = np.lexsort((np.arange(n), d)) idx = order[:k] return idx, d[idx] def wdist_rchemo(d, h, cri=4.0, squared=False): d = np.asarray(d, dtype=float).reshape(-1) if d.size == 0: raise ValueError("Distance vector must be non-empty.") h = float(h) if not np.isfinite(h) and not np.isinf(h): raise ValueError("h must be a positive scalar or infinity.") if h <= 0: raise ValueError("h must be positive.") if squared: d = d ** 2 zmed = np.median(d) zmad = r_mad(d) inside = d < (zmed + cri * zmad) with np.errstate(divide="ignore", invalid="ignore"): w = np.where(inside, np.exp(-d / (h * zmad)), 0.0) wmax = np.max(w) w = w / wmax w[np.isnan(w)] = 1.0 return w def apply_weight_floor(w, floor=WEIGHT_FLOOR): w = np.asarray(w, dtype=float).reshape(-1).copy() w[w < floor] = floor return w def plskern_wpls(X, y, weights, n_components): X = np.asarray(X, dtype=float) y = np.asarray(y, dtype=float).reshape(-1, 1) n, p = X.shape n_components = int(n_components) if n_components < 0: raise ValueError("n_components must be a non-negative integer.") if n_components > min(n, p): raise ValueError("n_components exceeds min(n_local, n_features).") delta = mweights(weights) xmeans = delta @ X ymeans = float(delta @ y[:, 0]) Xc = X - xmeans Yc = y - ymeans if n_components == 0: return { "R": np.zeros((p, 0)), "C": np.zeros((1, 0)), "W": np.zeros((p, 0)), "P": np.zeros((p, 0)), "T": np.zeros((n, 0)), "xmeans": xmeans, "ymeans": ymeans, "weights": delta, } Xd = delta[:, None] * Xc tXY = Xd.T @ Yc T = np.zeros((n, n_components)) R = np.zeros((p, n_components)) W = np.zeros((p, n_components)) P = np.zeros((p, n_components)) C = np.zeros((1, n_components)) for a in range(n_components): w = tXY.copy() nrm = np.sqrt(np.sum(w * w)) if not np.isfinite(nrm) or nrm == 0: raise ValueError("Weighted PLS extracted a zero X-weight vector.") w = w / nrm r = w.copy() if a > 0: for j in range(a): r = r - float(np.sum(P[:, j] * w[:, 0])) * R[:, [j]] t = Xc @ r tt = float(np.sum(delta * t[:, 0] * t[:, 0])) if not np.isfinite(tt) or tt == 0: raise ValueError("Weighted PLS score variance is zero for a requested component.") c = (tXY.T @ r) / tt zp = (Xd.T @ t) / tt tXY = tXY - (zp @ c) * tt T[:, a] = t[:, 0] P[:, a] = zp[:, 0] W[:, a] = w[:, 0] R[:, a] = r[:, 0] C[:, a] = c[0, 0] return { "R": R, "C": C, "W": W, "P": P, "T": T, "xmeans": xmeans, "ymeans": ymeans, "weights": delta, } def predict_plskern(model, X_new): X_new = np.asarray(X_new, dtype=float) if model["R"].shape[1] == 0: return np.full(X_new.shape[0], model["ymeans"]) B = model["R"] @ model["C"].T intercept = model["ymeans"] - model["xmeans"] @ B[:, 0] return intercept + X_new @ B[:, 0] class KNNLWPLSR: def __init__( self, n_neighbors, n_components, distance="euclidean", nlvdis=0, h=2.0, cri=4.0, weighting="rchemo", ): self.n_neighbors = int(n_neighbors) self.n_components = int(n_components) self.distance = distance self.nlvdis = int(nlvdis) self.h = float(h) if not np.isinf(h) else np.inf self.cri = float(cri) self.weighting = weighting if self.distance not in ("euclidean", "mahalanobis"): raise ValueError("distance must be 'euclidean' or 'mahalanobis'.") if self.weighting not in ("rchemo", "uniform"): raise ValueError("weighting must be 'rchemo' or 'uniform'.") if self.nlvdis < 0: raise ValueError("nlvdis must be a non-negative integer.") if self.cri <= 0: raise ValueError("cri must be positive.") def fit(self, X, y): X, y = _validate_xy(X, y) if self.n_neighbors < 1 or self.n_neighbors > X.shape[0]: raise ValueError("n_neighbors must be between 1 and n inclusive.") self.X_ = X self.y_ = y self.uinv_ = None self.global_pls_ = None if self.nlvdis > 0: ones = np.ones(X.shape[0]) self.global_pls_ = plskern_wpls(X, y, ones, self.nlvdis) elif self.distance == "mahalanobis": _whitened, self.uinv_ = mahalanobis_whiten(X) return self def _reference_space(self, X_query): if self.global_pls_ is None: if self.distance == "mahalanobis": return self.X_ @ self.uinv_, X_query @ self.uinv_, "euclidean", None return self.X_, X_query, self.distance, self.uinv_ T_ref = (self.X_ - self.global_pls_["xmeans"]) @ self.global_pls_["R"] T_q = (X_query - self.global_pls_["xmeans"]) @ self.global_pls_["R"] uinv = None if self.distance == "mahalanobis": T_ref, uinv = mahalanobis_whiten(T_ref) T_q = T_q @ uinv return T_ref, T_q, "euclidean", None return T_ref, T_q, "euclidean", None def query_local(self, z): z = np.asarray(z, dtype=float).reshape(1, -1) X_ref, Z_q, distance, uinv = self._reference_space(z) idx, d = knn_select(X_ref, Z_q[0], self.n_neighbors, distance=distance, uinv=uinv) if self.weighting == "uniform": w = np.ones(idx.shape[0]) else: w = apply_weight_floor(wdist_rchemo(d, h=self.h, cri=self.cri)) X_loc = self.X_[idx] y_loc = self.y_[idx] local = plskern_wpls(X_loc, y_loc, w, self.n_components) yhat = float(predict_plskern(local, z)[0]) return { "indices": idx, "distances": d, "observation_weights": w, "yhat": yhat, "local_model": local, } def predict(self, X_new): X_new = np.asarray(X_new, dtype=float) if X_new.ndim != 2: raise ValueError("X_new must be a 2D array of samples by variables.") if X_new.shape[1] != self.X_.shape[1]: raise ValueError("X_new must have the same number of variables as X.") if X_new.shape[0] > 0 and not np.all(np.isfinite(X_new)): raise ValueError("X_new must contain only finite values.") return np.array([self.query_local(row)["yhat"] for row in X_new], dtype=float) def loo_candidate_mask(n, i): mask = np.ones(n, dtype=bool) mask[i] = False return mask MATLAB
MATLAB has no built-in LWPLSR. The listing implements the same WPLSR and rchemo-style weights. Do not pass observation weights to plsregress and call the result KNN-LWPLSR.
function model = fit_knnlwplsr(X, y, nNeighbors, nComponents, distance, nlvdis, h, cri, weighting) if nargin < 5, distance = "euclidean"; end if nargin < 6, nlvdis = 0; end if nargin < 7, h = 2; end if nargin < 8, cri = 4; end if nargin < 9, weighting = "rchemo"; end [X, y] = validateXY_lwpls(X, y); n = size(X, 1); nNeighbors = double(nNeighbors); nComponents = double(nComponents); nlvdis = double(nlvdis); if nNeighbors < 1 || nNeighbors ~= floor(nNeighbors) || nNeighbors > n error('nNeighbors must be an integer between 1 and n inclusive.'); end if nComponents < 0 || nComponents ~= floor(nComponents) error('nComponents must be a non-negative integer.'); end model.X = X; model.y = y; model.nNeighbors = nNeighbors; model.nComponents = nComponents; model.distance = char(distance); model.nlvdis = nlvdis; model.h = h; model.cri = cri; model.weighting = char(weighting); model.uinv = []; model.globalPls = []; if nlvdis > 0 model.globalPls = plskern_wpls(X, y, ones(n, 1), nlvdis); elseif strcmp(model.distance, "mahalanobis") [~, model.uinv] = mahalanobisWhiten_lwpls(X); endend function yhat = predict_knnlwplsr(model, Xnew) Xnew = validateXnew_lwpls(Xnew, size(model.X, 2)); m = size(Xnew, 1); yhat = zeros(m, 1); for i = 1:m q = query_knnlwplsr(model, Xnew(i, :)); yhat(i) = q.yhat; endend function q = query_knnlwplsr(model, z) z = z(:)'; [Xref, zq, distance, uinv] = referenceSpace_lwpls(model, z); [idx, d] = knnSelect_lwpls(Xref, zq, model.nNeighbors, distance, uinv); if strcmp(model.weighting, "uniform") w = ones(numel(idx), 1); else w = applyWeightFloor_lwpls(wdist_rchemo(d, model.h, model.cri)); end local = plskern_wpls(model.X(idx, :), model.y(idx), w, model.nComponents); yhat = predict_plskern(local, z); q.indices = idx; q.distances = d; q.observation_weights = w; q.yhat = yhat; q.local_model = local;end function w = wdist_rchemo(d, h, cri) d = d(:); if nargin < 3, cri = 4; end zmed = median(d); zmad = 1.4826 * median(abs(d - zmed)); inside = d < (zmed + cri * zmad); w = zeros(size(d)); w(inside) = exp(-d(inside) ./ (h * zmad)); wmax = max(w); w = w / wmax; w(isnan(w)) = 1;end function w = applyWeightFloor_lwpls(w) w = w(:); w(w < 1e-5) = 1e-5;end function delta = mweights_lwpls(weights) weights = weights(:); total = sum(weights); if ~isfinite(total) || total <= 0 error('Observation weights must have a positive finite sum.'); end delta = weights / total;end function fm = plskern_wpls(X, y, weights, nComponents) n = size(X, 1); p = size(X, 2); nComponents = double(nComponents); if nComponents > min(n, p) error('nComponents exceeds min(n_local, n_features).'); end delta = mweights_lwpls(weights); xmeans = delta' * X; ymeans = delta' * y(:); Xc = X - xmeans; Yc = y(:) - ymeans; if nComponents == 0 fm.R = zeros(p, 0); fm.C = zeros(1, 0); fm.W = zeros(p, 0); fm.P = zeros(p, 0); fm.T = zeros(n, 0); fm.xmeans = xmeans; fm.ymeans = ymeans; fm.weights = delta; return end Xd = delta .* Xc; tXY = Xd' * Yc; T = zeros(n, nComponents); R = zeros(p, nComponents); Wmat = zeros(p, nComponents); P = zeros(p, nComponents); C = zeros(1, nComponents); for a = 1:nComponents w = tXY; nrm = sqrt(sum(w .* w)); if ~isfinite(nrm) || nrm == 0 error('Weighted PLS extracted a zero X-weight vector.'); end w = w / nrm; r = w; if a > 1 for j = 1:(a - 1) r = r - sum(P(:, j) .* w) * R(:, j); end end t = Xc * r; tt = sum(delta .* t .* t); if ~isfinite(tt) || tt == 0 error('Weighted PLS score variance is zero for a requested component.'); end c = (tXY' * r) / tt; zp = (Xd' * t) / tt; tXY = tXY - (zp * c) * tt; T(:, a) = t; P(:, a) = zp; Wmat(:, a) = w; R(:, a) = r; C(:, a) = c; end fm.R = R; fm.C = C; fm.W = Wmat; fm.P = P; fm.T = T; fm.xmeans = xmeans; fm.ymeans = ymeans; fm.weights = delta;end function yhat = predict_plskern(fm, Xnew) Xnew = Xnew(:, :); if size(fm.R, 2) == 0 yhat = fm.ymeans * ones(size(Xnew, 1), 1); return end B = fm.R * fm.C'; intercept = fm.ymeans - fm.xmeans * B; yhat = intercept + Xnew * B;end function [idx, dsel] = knnSelect_lwpls(Xref, z, k, distance, uinv) n = size(Xref, 1); z = z(:)'; if strcmp(distance, "mahalanobis") z = z * uinv; d = sqrt(sum((Xref - z).^2, 2)); else d = sqrt(sum((Xref - z).^2, 2)); end [~, order] = sortrows([d, (1:n)']); idx = order(1:k); dsel = d(idx);end function [Xw, uinv] = mahalanobisWhiten_lwpls(Xtrain) n = size(Xtrain, 1); p = size(Xtrain, 2); sigma = cov(Xtrain, 1); if p == 1 sigma = var(Xtrain, 1); end [U, flag] = chol(sigma); if flag ~= 0 U = chol(sigma + 1e-5 * eye(p)); end uinv = inv(U); Xw = Xtrain * uinv;end function [Xref, zq, distance, uinv] = referenceSpace_lwpls(model, z) if isempty(model.globalPls) if strcmp(model.distance, "mahalanobis") Xref = model.X * model.uinv; zq = z * model.uinv; distance = "euclidean"; uinv = []; else Xref = model.X; zq = z; distance = model.distance; uinv = model.uinv; end return end Tref = (model.X - model.globalPls.xmeans) * model.globalPls.R; Tq = (z - model.globalPls.xmeans) * model.globalPls.R; if strcmp(model.distance, "mahalanobis") [Tref, uinv] = mahalanobisWhiten_lwpls(Tref); Xref = Tref; zq = Tq * uinv; distance = "euclidean"; uinv = []; else Xref = Tref; zq = Tq; distance = "euclidean"; uinv = []; endend function [X, y] = validateXY_lwpls(X, y) y = y(:); if size(X, 1) ~= numel(y) error('X and y must have the same number of observations.'); end if ~all(isfinite(X(:))) || ~all(isfinite(y)) error('X and y must contain only finite values.'); endend function Xnew = validateXnew_lwpls(Xnew, p) if size(Xnew, 2) ~= p error('Xnew must have the same number of variables as X.'); end if ~isempty(Xnew) && ~all(isfinite(Xnew(:))) error('Xnew must contain only finite values.'); endend Practical notes
- Local PLS is a family. The public engine is KNN-LWPLSR, not every local PLS variant.
- Neighbor selection and within-neighborhood observation weighting are different operations.
- The distance metric and any scaling define local geometry. They are model choices, not universal requirements.
- SNV, derivatives, or autoscaling used in the 2026 tutorial example are not part of the method definition.
- k, h, and A have no universal defaults. Tune them without external-test responses.
- Small neighborhoods can be unstable or rank-limited.
- A query outside calibration coverage remains poorly supported.
- Prediction rebuilds a local PLS model and is heavier than applying one global PLSR.
- Duplicate calibration spectra can dominate a neighborhood. Duplicate curation in the 2026 example is database quality, not a mandatory algorithm step.
- Use the Open Lab Regression Metrics card. This page does not redefine RMSE, MAE, R2, or bias.
Extensions
Kim et al. 2013: adaptive similarity for JIT LW-PLS, not in V1. Hazama and Kano 2015: covariance-based similarity that uses the X–Y relationship, not Euclidean/Mahalanobis X geometry alone; not implemented. Shen et al. 2019: locality in a global PLS score space; available here only as optional . Lesnoff and collaborators: averaging local PLSR predictions over several dimensionalities; not core V1. LW-PLS-DA applies the local idea to classification and is a separate Open Lab card when published. Sicard and Sabatier 2006 develop a local PLS1 framework with multivariate local polynomials. It is related literature, not this KNN-LWPLSR engine.
References
- 1.
Lesnoff, M., Metz, M., & Roger, J.-M. (2020). Comparison of locally weighted PLS strategies for regression and discrimination on agronomic NIR data. Journal of Chemometrics, 34(5), e3209.
doi:10.1002/cem.3209 - 2.
Deryck, A., Fernandez Pierna, J. A., Baeten, V., & Lesnoff, M. (2026). A step-by-step tutorial to local partial least squares analyses for near-infrared spectroscopic data analysis. Analytica Chimica Acta, 1401, 345242.
doi:10.1016/j.aca.2026.345242 - 3.
Centner, V., & Massart, D. L. (1998). Optimization in locally weighted regression. Analytical Chemistry, 70(19), 4206-4211.
doi:10.1021/ac980208r - 4.
Sicard, E., & Sabatier, R. (2006). Theoretical framework for local PLS1 regression, and application to a rainfall data set. Computational Statistics & Data Analysis, 51(2), 1393-1410.
doi:10.1016/j.csda.2006.05.002 - 5.
Shen, G., Lesnoff, M., Baeten, V., Dardenne, P., Davrieux, F., Ceballos, H., Belalcazar, J., Dufour, D., Yang, Z., Han, L., & Fernandez Pierna, J. A. (2019). Local partial least squares based on global PLS scores. Journal of Chemometrics, 33(5), e3117.
doi:10.1002/cem.3117 - 6.
Kim, S., Okajima, R., Kano, M., & Hasebe, S. (2013). Development of soft-sensor using locally weighted PLS with adaptive similarity measure. Chemometrics and Intelligent Laboratory Systems, 124, 43-49.
doi:10.1016/j.chemolab.2013.03.008 - 7.
Hazama, K., & Kano, M. (2015). Covariance-based locally weighted partial least squares for high-performance adaptive modeling. Chemometrics and Intelligent Laboratory Systems, 146, 55-62.
doi:10.1016/j.chemolab.2015.05.007 - 8.
Dayal, B. S., & MacGregor, J. F. (1997). Improved PLS algorithms. Journal of Chemometrics, 11(1), 73-85.
doi:10.1002/(SICI)1099-128X(199701)11:1<73::AID-CEM435>3.0.CO;2-# - 9.
Kim, S., Kano, M., Nakagawa, H., & Hasebe, S. (2011). Estimation of active pharmaceutical ingredients content using locally weighted partial least squares and statistical wavelength selection. International Journal of Pharmaceutics, 421(2), 269-274.
doi:10.1016/j.ijpharm.2011.10.007 - 10.
Shenk, J., Westerhaus, M., & Berzaghi, P. (1997). Investigation of a LOCAL calibration procedure for near infrared instruments. Journal of Near Infrared Spectroscopy, 5(1), 223-232.
doi:10.1255/jnirs.115 - 11.
ChemHouse-group (n.d.). rchemo: lwplsr, locw, wdist, plskern, getknn. R package source inspected from ChemHouse-group/rchemo.
- 12.
Lesnoff, M. (n.d.). Jchemo.jl: lwplsr, winvs, getknn. Julia package source inspected from mlesnoff/Jchemo.jl.
