2  Datasets

Below are curated chemical/process‑engineering datasets you can use for practicing PCA (and MP‑PC/DPCA for process monitoring). For each dataset I give a short description, what kind of PCA experiments it’s good for, data size / variables, a download link, and quick tips for preprocessing and use. I also include a compact Python example (scikit‑learn + pandas) showing PCA on a typical chemical sensor / air‑quality style dataset.

Recommended datasets

  1. Tennessee Eastman Process (TEP)
  1. UCI Gas Sensor Array Drift Dataset (electronic nose)
  1. UCI Air Quality Data Set
  1. UCI Wine Quality (physicochemical tests)
  1. Combined Cycle Power Plant (UCI)
  1. CSTR / Simple reactor and distillation simulation datasets (public notebooks & repos)
  1. Industrial Benchmark / synthetic datasets
  1. Kaggle & institutional repositories

Preprocessing tips for PCA on chemical/process data - Missing values: impute or drop rows/columns; for time series consider interpolation. - Centering & scaling: Always mean‑center. Use StandardScaler (unit variance) if variables have different units/ranges; use autoscaling in chemometrics. - Time correlation: For time‑series processes, consider Dynamic PCA (DPCA) or time-windowed PCA to capture autocorrelation. - Outliers: Remove or analyze separately; robust PCA variants exist if outliers are expected. - Stationarity: If process has trends, detrend first (e.g., remove slow drift) before PCA aimed at fault detection. - Number of components: Use explained variance ratio (e.g., keep PCs covering 85–95%), scree plot, cross‑validation, or criteria from MSPC literature (T2+SPE residuals).

Common PCA experiments / analyses to try - Scree plot + variance explained to choose k PCs. - Loadings interpretation: which variables contribute to each PC. - Scores scatter plots to inspect clusters / fault separation. - Hotelling T2 and SPE (Q) monitoring charts for fault detection. - Reconstruction error and contribution plots for fault isolation. - DPCA / lagged variables for processes with strong dynamics.

Quick Python example: load UCI Air Quality CSV, preprocess, run PCA, show explained variance - This example assumes you download the AirQualityUCI.csv from the UCI page and place it in your working folder.

# example_pca_airquality.py
import pandas as pd
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
import matplotlib.pyplot as plt

# Load (update path if needed)
df = pd.read_csv('AirQualityUCI.csv', sep=';', decimal=',')
# Drop the final empty column if present
df = df.loc[:, ~df.columns.str.contains('^Unnamed')]

# Select numeric sensor/meteorological columns (example)
cols = ['CO(GT)', 'NMHC(GT)', 'C6H6(GT)', 'NOx(GT)', 'NO2(GT)', 'PT08.S1(CO)',
        'PT08.S2(NMHC)', 'PT08.S3(NOx)', 'PT08.S4(NO2)', 'PT08.S5(O3)',
        'T', 'RH', 'AH']
data = df[cols].replace(-200.0, pd.NA).dropna()  # -200 = missing in this dataset

# Scale
scaler = StandardScaler()
X = scaler.fit_transform(data)

# PCA
pca = PCA()
Xp = pca.fit_transform(X)

# Explained variance plot
plt.figure()
plt.plot(range(1, len(pca.explained_variance_ratio_)+1),
         pca.explained_variance_ratio_.cumsum(), marker='o')
plt.xlabel('Number of components')
plt.ylabel('Cumulative explained variance')
plt.grid(True)
plt.show()

# Scores scatter on first two PCs
plt.figure()
plt.scatter(Xp[:,0], Xp[:,1], s=10, alpha=0.6)
plt.xlabel('PC1'); plt.ylabel('PC2'); plt.title('Scores (PC1 vs PC2)')
plt.show()

# Loadings (variables × components)
loadings = pd.DataFrame(pca.components_.T,
                        index=cols,
                        columns=[f'PC{i+1}' for i in range(len(cols))])
print(loadings.iloc[:, :3])  # first 3 PCs

References and reading - Wold, Sjöström, and Eriksson — “PLS in chemistry”; Jackson — “A User’s Guide to Principal Components” (general PCA foundations); textbooks on Multivariate Statistical Process Control (MSPC). - Look up “Hotelling T2”, “Squared Prediction Error (SPE / Q)”, and “Dynamic PCA (DPCA)” for process monitoring theory.

If you want, I can: - Download one of the datasets (TEP / Air Quality / Gas sensor) and prepare a ready-to-run Jupyter notebook with PCA, plots (scores, loadings, T2/SPE), and interpretation. - Or, if you tell me a specific dataset from the list (or a link you already have), I’ll run PCA on it and show results, interpretations, and suggested next experiments.

Which dataset would you like me to prepare a notebook/report for?

https://www.kaggle.com/code/sparshattri/heat-exchanger-lmtd-model-with-noise-tolerance?select=README.md

https://www.kaggle.com/datasets/fuarresvij/steel-test-data/data

https://www.kaggle.com/datasets/natanaelferran/river-water-parameters

https://github.com/StephenGoldie/indpensim-notebook?tab=readme-ov-file

https://data.mendeley.com/

https://data.mendeley.com/datasets/nwy6zpgdys/1