Open Lab/Multiblock Data Analysis · Concept
Multiblock Data Analysis & Data Fusion
MB
A conceptual map of linked data blocks, fusion levels, common and distinct subspaces, and the algebraic choices that distinguish chemometric multiblock methods.
MULTIBLOCK
What is multiblock data analysis?
A data block is one structured table that belongs to a larger collection of related tables. Smilde and co-authors describe data fusion as the simultaneous analysis of several such blocks in order to obtain a global view of one system. The same scientific activity is also called multiblock analysis, multiset analysis, or data integration. Those names are not independent algorithms.
Mishra and co-authors treat multi-block chemometric data as multi-source: each block is generated by a distinct source of variability, such as a second spectroscopic platform, a sensory panel, or a process-condition set, while remaining multivariate inside the block. Ordinary single-block PCA or PLS applied to one source extracts only part of that joint information.
Multiblock analysis is not the mere presence of several matrices in a folder. The scientific problem requires an explicit linking structure: shared objects, shared variables, a predictive relationship, or another documented experimental connection.
Why multiple blocks?
Linked blocks arise because one system is probed in more than one way. Smilde et al. 2017 give food-science examples in which formulation, sensory modalities, and liking are collected on the same products, and biological examples in which metabolomics, clinical measurements, and lifestyle variables are collected on the same subjects. Mishra et al. give the chemometric case of the same samples measured by several spectroscopic techniques, such as mid-infrared and Raman.
The scientific motive is not automatic improvement. Additional blocks may contribute complementary structure, overlapping structure, or variation that is irrelevant to the question. Fusion does not guarantee better prediction, clearer interpretation, or greater robustness.
Anatomy of a multiblock data set
This page uses the shared-sample convention of Smilde et al. 2017. For blocks measured on the same objects,
where is the number of variables in block . The first mode is the shared sample or object mode. The variable modes need not match across blocks and often do not.
A different skeleton shares the variable mode rather than the sample mode. Smilde, Westerhuis and de Jong describe that class as several data sets with different objects but the same variables. Under that convention a block may be written
Shared sample mode is the usual chemometric multi-platform layout. It is not the only multiblock configuration.
How blocks can be linked
Smilde, Westerhuis and de Jong classify multiblock problems by which modes are in common. Blocks may share the object mode, the variable mode, both, or neither in an obvious structural sense. Complex arrangements also exist, including L-shaped collections and path structures in which some blocks act as predictors of others. Those geometries are named here only as advanced topologies. They are not derived on this card.
The 2022 monograph of Smilde, Næs and Liland lists, in Chapter 3, the skeleton of a multiblock data set and the topology of that set, with separate headings for unsupervised and supervised analysis. Topology in that table of contents is the scientific arrangement of block roles, not a loose graph drawing. Detailed wording of those sections is not reproduced here from unopened book body text.
Homogeneous and heterogeneous fusion
Heterogeneity has more than one meaning. Smilde et al. 2020 distinguish the type of biochemical measurement, such as metabolomics versus RNAseq, from the statistical measurement scale: binary, ordinal, interval, or ratio. The second distinction is the one used here for homogeneous versus heterogeneous fusion.
Homogeneous fusion, in that sense, is the situation in which all blocks are of the same scale type. The usual chemometric case is several quantitative blocks. Heterogeneous fusion is the situation in which the blocks have different scale types, for example a quantitative block together with a binary block. Traditional multiblock component methods were developed primarily for quantitative, homogeneous tables.
Instrumental modality is not that definition. Two quantitative spectroscopic blocks, such as mid-infrared and Raman on the same samples, are multi-source and may be complementary. They are not heterogeneous merely because the instruments differ. Likewise, numerical magnitude is not scale type. One quantitative block with values near 0.01 and another near 1000 remain homogeneous quantitative data. Their relative influence is a scaling problem, not a change of measurement scale.
Low, mid and high level fusion
Smilde et al. 2020, following the classical fusion-level distinction, separate where information is combined. They write low-, medium-, and high-level fusion. This page uses mid-level as a synonym of that medium-level stage. In machine learning the same three stages are often called early, intermediate, and late integration.
Low-level fusion combines the blocks at the level of the (preprocessed) measurements. For shared-sample quantitative blocks that can be concatenation
Concatenation preserves the original variables in the combined table. It is mathematically simple and scientifically not neutral: scales, numbers of variables, and variance structure still affect the relative contribution of each block. Low-level fusion is not defined as untouched instrument output. The source wording is preprocessed measurements.
Mid-level fusion first summarizes each block, for example by dimension reduction or variable selection, and then fuses those reduced representations. Concatenating PCA scores is one possible implementation, not the definition.
High-level fusion, in the supervised setting of that paper, fits a predictor or classifier on each input block and then combines the prediction or classification outputs, for example by majority voting. Voting is one classification strategy. Combined regression predictions are also high-level fusion. High-level fusion does not combine raw measurements, and it does not always improve prediction.
Fusion level is not an algorithm family. Low, mid, and high describe where information is integrated. They do not determine the objective, the supervision structure, whether components are extracted by deflation, whether predictor blocks enter in a scientific order, or whether common and distinct subspaces are modelled.
Unsupervised and supervised multiblock analysis
Unsupervised multiblock analysis looks for structure within blocks and across blocks without declaring one block to be an external response. Smilde et al. 2017 restrict their common and distinct framework to interchangeable blocks that share the sample mode. In that paradigm the numerical roles of the tables are exchangeable. Exchangeable roles do not imply that every algorithm is invariant to implementation details such as initialization or block scaling.
Supervised multiblock analysis relates a response block to one or more predictor blocks . The predictor and response roles are then not exchangeable. Mishra et al. treat visualization, regression, classification, and variable selection as different tasks on the same multi-source layout. Supervised multiblock analysis is not synonymous with MB-PLS.
What information are we trying to model?
A single quantitative table is often summarized by a component model
with scores , loadings , and residual . That bilinear form is a generic component representation. It is not the model of every multiblock method.
When several blocks are modelled, some methods introduce block scores and block loadings, and some also introduce super-level scores that combine the blocks. Those quantities are method-specific. This page does not assign one global-score equation to every algorithm.
Within-block and between-block variation
Smilde, Westerhuis and de Jong state the generic problem as finding relationships between several possibly related data sets, using components to summarize relevant information between and within the blocks. Those two directions are not the same objective.
Within-block variation is the structure that can be described inside one table. Between-block variation is the association, covariance, correlation, or shared subspace that relates tables to one another. Covariance, correlation, and common-subspace criteria are not mathematically equivalent. Smilde, Westerhuis and de Jong treat the balance between describing within-block variation and describing relationships between blocks as one of two important modelling choices, the other being fairness of block contribution. Different algorithms exist because they compromise these aims differently.
Common, local and distinct variation
Smilde et al. 2017 define common and distinct structure in the column spaces of blocks that share samples. For two column-centred blocks and ,
The common subspace is the nontrivial intersection of the two column spaces, when it exists. Distinct subspaces complete each block as a direct sum. Common variation is therefore a subspace concept, not a list of variables that happen to appear in both tables. A single variable can carry both common-related and distinct-related variation.
Distinct structure is systematic variation associated with a block that is not part of the common subspace under the chosen decomposition. It is not residual noise. Smilde et al. keep irrelevant variation and noise separate from the common and distinct sources they wish to quantify.
Other methods use related but not identical words. Lock et al. decompose each matrix into joint structure, individual structure, and residual error, and label that decomposition JIVE. Löfstedt and Trygg describe OnPLS in terms of globally predictive components and orthogonal variation that may be local to combinations of matrices or unique to one matrix. In that paper, predictive refers to covariance and correlation among the modelled matrices in a symmetric O2PLS extension. It does not mean supervised prediction of an external . Those vocabularies are not treated as proven synonyms of the Smilde 2017 subspaces.
Simultaneous and sequential strategies
The word sequential is used for more than one operation in this literature. The two uses on this page are not synonymous.
Smilde, Westerhuis and de Jong call a multiblock component method sequential when it calculates one component, removes the influence of that component from the matrices, typically by deflation, and then calculates the next component from the corrected matrices. They contrast that organization with simultaneous methods that obtain all components by solving a global optimization problem. Their paper concerns unsupervised component models, not prediction of an external response.
SO-PLS literature uses sequential for an ordered treatment of predictor blocks: later blocks are modelled after information already extracted from earlier blocks has been accounted for, typically by orthogonalization. Mishra et al. describe that family as combining PLS with a sequential orthogonalization step to extract complementary latent variables. Block order then carries scientific meaning. It is not a deflation sequence of unsupervised components.
Supervision and estimation organization are independent axes. Sequential component extraction is not the definition of supervised analysis. Simultaneous estimation is not the definition of unsupervised analysis.
Block scaling and block dominance
Preprocessing can be applied inside a block and between blocks. Mishra et al., following Campos and Reis, separate artefact correction, equalization of variables within a block, and equalization of inter-block effects such as scale, number of variables, and pseudo-rank. Variable scaling inside a block and scaling of a block as a whole answer different questions. This page does not prescribe a universal block-scaling rule, including unit Frobenius-norm scaling of every block.
A block can dominate a solution through its variance, its width, its scale, and the objective of the method. Variable count alone does not determine dominance. Westerhuis, Kourti and MacGregor note that one CPCA implementation divides block scores by the square root of the number of variables so that each block starts with the same variance irrespective of its size, with optional extra factors when prior block importance is known.
Smilde, Westerhuis and de Jong use fairness for whether every block contributes, as opposed to methods that act as block selectors. In their analysis, one Consensus PCA formulation has a built-in tendency for fairness, whereas hierarchical PCA as they analyse it tends to discriminate among blocks. Those statements belong to those algorithms. They are not a numerical fairness index for the whole field.
The algebra behind the main families
There is no single objective of the form maximize "multiblock information". Methods differ because they answer different algebraic questions.
Superblock analysis concatenates shared-sample blocks and analyses . Westerhuis, Kourti and MacGregor show that, when the same variable-scaling conventions are applied and there are no missing values, results from Consensus PCA can be calculated from standard PCA of the concatenated table, and results from multiblock PLS with super-score deflation can be calculated from standard PLS. If CPCA or MB-PLS is run without that block scaling step, the concatenated matrix used for the standard methods must be scaled as to match. The multiblock value they emphasize is then interpretation: block weights, block loadings, super-block quantities, and the variation explained in each block, recovered after the standard model is fit. That is not the claim that multiblock analysis is identical to single-block analysis, nor that every CPCA or MB-PLS variant collapses in the same way. No such equivalence is given for hierarchical PCA or hierarchical PLS. Those hierarchical algorithms are described as lacking a clear objective and as possibly depending on the initial super-score.
A second family summarizes each block by components and then models relations among those representations. A third family seeks explicit common and distinct subspaces, as in Smilde et al. 2017. A fourth family is supervised: predictor blocks are combined to explain or predict , as in MB-PLS or SO-PLS. A fifth family emphasizes association between blocks, including generalized canonical analysis. PCA, PLS, and canonical correlation do not optimize the same criterion.
Shared-sample notation and labelled models
- block m, I samples by J_m variables
- concatenation of shared-sample blocks along the variable mode
- common column subspace of two blocks (Smilde 2017)
- JIVE joint and individual terms. Lock et al. write each matrix as variables by samples. This is not the generic Smilde common/distinct split.
Interpretation
Equation (4) is a generic single-block component representation, not a universal multiblock model. Equation (5) is the Smilde et al. 2017 two-block column-space split. Equation (6) is the JIVE model of Lock et al. and is labelled as such.
Map of multiblock methods
The table orients future dedicated cards. It is not a ranking. Empty cells are omitted rather than guessed. Fusion level is not used as a column, because it does not uniquely classify these algorithms.
| Method | Task | Organization | What it targets |
|---|---|---|---|
| Consensus PCA (CPCA) | Unsupervised | Sequential components with deflation | Super-scores and block contributions. Formulations differ. Westerhuis et al. relate CPCA results to standard PCA under matching variable scaling and no missing values. |
| Hierarchical PCA (HPCA) | Unsupervised | Sequential components | Not interchangeable with CPCA. Westerhuis et al. report an unclear objective and start dependence. |
| SUM-PCA | Unsupervised | Superblock PCA | PCA on concatenated shared-sample blocks. Smilde et al. 2003 identify one CPCA-W form with SUM-PCA. |
| CCSWA / ComDim | Unsupervised | Sequential common components | Common components with block-specific weights (saliences) for tables on the same samples. |
| JIVE | Unsupervised | Iterative low-rank fits | Joint structure plus individual structure plus residual. Joint is JIVE's term, not a synonym of Smilde common. |
| OnPLS | Symmetric multiblock | O2PLS-style extraction | Globally predictive covariance among matrices, plus orthogonal variation local to combinations or unique to one matrix. Not Y-supervised in the original paper. |
| DISCO-SCA | Unsupervised | Simultaneous components then rotation | Common and distinct components in the Smilde et al. 2017 comparison. |
| Generalized canonical analysis | Unsupervised | Between-block association | Correlational relationships among blocks, including the PCA-GCA combination listed by Mishra et al. Not the PCA variance criterion and not the PLS covariance criterion. |
| MB-PLS | Supervised | Related to PLS of concatenated predictors | Predictor blocks versus . Westerhuis et al. relate MB-PLS with super-score deflation to standard PLS under matching variable scaling and no missing values, with extra block-level interpretation. |
| SO-PLS | Supervised | Sequential predictor-block entry | Incremental information for after orthogonalizing later blocks with respect to earlier modelled information. Not MB-PLS. |
| PO-PLS | Supervised | Parallel orthogonalization | Common and distinct Y-related latent variables across predictor blocks, as characterized by Mishra et al. Not SO-PLS. |
Wangen and Kowalski 1989 presented a multiblock PLS algorithm for complex chemical systems. Westerhuis et al. describe that algorithm as allowing pathway relationships among blocks. That historical origin is not used here as a substitute for the later Westerhuis comparison of CPCA, HPCA, MB-PLS, and hierarchical PLS.
Multiblock versus multiway data
A typical shared-sample multiblock problem is a collection of matrices . A multiway problem is a higher-order array such as . Concatenating matrices does not create a PARAFAC tensor.
Mishra et al. note that when the data order increases, Tucker and PARAFAC are the higher-order counterparts of bilinear PCA, and that multi-block methods for mixed matrix and tensor sources also exist. PARAFAC exploits genuine multilinear structure. It is not a generic name for any fusion of several tables. The 2022 monograph places three-way methods inside a broad multiblock discussion when both sample and variable modes are shared. The data objects remain different.
Mishra et al. also place augmented MCR in the multi-source map: matrices are concatenated along a common mode, then resolved under chemical constraints. That is a use of MCR-ALS on a linked layout. It does not redefine the two-way MCR-ALS bilinear model already documented in Open Lab.
Choosing a method family
Method choice follows the scientific question. It is not a deterministic tree of the form if X then always algorithm Y.
Identify the skeleton: shared samples, shared variables, or a more complex topology. Identify scale type: homogeneous quantitative blocks versus mixed measurement scales. Identify the goal: exploration, prediction or classification of , relationships among blocks, or an explicit common/distinct decomposition. Decide where fusion should occur. Decide whether within-block variance, between-block association, covariance with , or subspace separation is the algebraic priority. Decide whether components should be extracted sequentially by deflation, whether predictor blocks should enter in a scientific order, or whether a joint criterion should be solved simultaneously. Then select a family whose documented objective matches that question.
If the aim is exploratory description of several quantitative blocks that share samples, unsupervised multiblock component methods are the relevant family. If the aim is prediction of from several predictor blocks, supervised multiblock regression methods are the relevant family. If block order encodes incremental information, sequential supervised methods such as SO-PLS may be relevant. If common and block-specific subspaces are the object of interpretation, common/distinct methods may be relevant. Mishra et al. state that no technique is best in an absolute sense.
Practical notes
- A multiblock problem requires a documented linking structure. Several unrelated matrices are not a fusion problem.
- Data fusion is not synonymous with concatenation. Concatenation is one low-level option for shared-sample tables.
- Homogeneous versus heterogeneous refers to measurement scale type, not to whether the instruments differ.
- Multimodal quantitative spectroscopy can still be homogeneous fusion.
- Low, mid, and high level fusion locate the integration step. They do not name the algorithm.
- Supervised versus unsupervised is not the same axis as sequential versus simultaneous.
- Sequential component deflation is not the same operation as sequential SO-PLS block entry.
- Common and distinct refer to subspaces of variation, not to shared variable names.
- Distinct systematic structure is not residual error. JIVE residual error is a third term in that model only.
- Variable scaling inside a block and scaling of whole blocks are different operations.
- CPCA, HPCA, and SUM-PCA are not automatic synonyms. MB-PLS, hierarchical PLS, SO-PLS, and PO-PLS are not the same algorithm.
- Westerhuis et al. relate CPCA, and MB-PLS with super-score deflation, to standard PCA and PLS under matching variable scaling when there are no missing values. Block-level interpretation remains the multiblock contribution they emphasize.
- Adding a block does not automatically improve prediction or interpretation.
- Any learned preprocessing, feature extraction, block weighting, or rank selection used to claim predictive performance belongs inside the training and validation structure of the Cross-Validation card.
- PARAFAC models a tensor. Multiblock methods model linked tables. Do not treat one as a notational rewrite of the other.
References
- 1.
Wangen, L. E., & Kowalski, B. R. (1989). A multiblock partial least squares algorithm for investigating complex chemical systems. Journal of Chemometrics, 3(1), 3-20.
doi:10.1002/cem.1180030104 - 2.
Westerhuis, J. A., Kourti, T., & MacGregor, J. F. (1998). Analysis of multiblock and hierarchical PCA and PLS models. Journal of Chemometrics, 12(5), 301-321.
doi:10.1002/(SICI)1099-128X(199809/10)12:5<301::AID-CEM515>3.0.CO;2-S - 3.
Qannari, E. M., Wakeling, I., Courcoux, P., & MacFie, H. J. H. (2000). Defining the underlying sensory dimensions. Food Quality and Preference, 11(1-2), 151-154.
doi:10.1016/S0950-3293(99)00069-5 - 4.
Smilde, A. K., Westerhuis, J. A., & de Jong, S. (2003). A framework for sequential multiblock component methods. Journal of Chemometrics, 17(6), 323-337.
doi:10.1002/cem.811 - 5.
Löfstedt, T., & Trygg, J. (2011). OnPLS: a novel multiblock method for the modelling of predictive and orthogonal variation. Journal of Chemometrics, 25(8), 441-455.
doi:10.1002/cem.1388 - 6.
Lock, E. F., Hoadley, K. A., Marron, J. S., & Nobel, A. B. (2013). Joint and individual variation explained (JIVE) for integrated analysis of multiple data types. The Annals of Applied Statistics, 7(1), 523-542.
doi:10.1214/12-AOAS597 - 7.
Smilde, A. K., Måge, I., Næs, T., Hankemeier, T., Lips, M. A., Kiers, H. A. L., Acar, E., & Bro, R. (2017). Common and distinct components in data fusion. Journal of Chemometrics, 31, e2900.
doi:10.1002/cem.2900 - 8.
Smilde, A. K., Song, Y., Westerhuis, J. A., Kiers, H. A. L., Aben, N., & Wessels, L. F. A. (2020). Heterofusion: Fusing genomics data of different measurement scales. Journal of Chemometrics, 35(2), e3200.
doi:10.1002/cem.3200 - 9.
Mishra, P., Roger, J. M., Jouan-Rimbaud-Bouveresse, D., Biancolillo, A., Marini, F., Nordon, A., & Rutledge, D. N. (2021). Recent trends in multi-block data analysis in chemometrics for multi-source data integration. TrAC Trends in Analytical Chemistry, 137, 116206.
doi:10.1016/j.trac.2021.116206 - 10.
Smilde, A. K., Næs, T., & Liland, K. H. (2022). Multiblock Data Fusion in Statistics and Machine Learning: Applications in the Natural and Life Sciences. Wiley.
doi:10.1002/9781119600978
PCA
Principal Component Analysis
Open
PLSR
Partial Least Squares Regression
Open
PARAFAC
Parallel Factor Analysis
Open
MCR-ALS
Multivariate Curve Resolution - Alternating Least Squares
Open
NS
Normalization & Scaling
Open
CV
Cross-Validation
Open
MB-PCA
Multiblock PCA / Consensus PCA
Open
MB-PLS
Multiblock Partial Least Squares
Open
SO-PLS
Sequential and Orthogonalized Partial Least Squares
Open
ComDim
Common Components and Specific Weights Analysis
Open
OnPLS
OnPLS
Coming soon
JIVE
JIVE
Coming soon
