Data to Dynamics


Associate Professor, Chemical Engineering, Curtin University, Australia
SMILE Lab (Sustainable Manufacturing and Intelligent Process Engineering)
Research interests
Experience

At the end of the course you should be able to:
Form groups of 4-5 people.
Introduce yourself to your group members. Tell them:
\Rightarrow What are your Objectives and Expectations?



\Rightarrow Challenging to design, difficult to predict, hard to control, and complex to optimize















| Criterion | First-Principles | Data-Driven | Hybrid |
|---|---|---|---|
| Process knowledge required | High | Low | Medium |
| Operational data required | Low | High | Medium |
| Development time | Weeks-months | Days-weeks | Weeks |
| Computational cost (training) | Low | Medium-High | High |
| Computational cost (prediction) | High | Low | Medium |
| Extrapolation capability | High | Low | Medium-High |
| Interpretability | High | Low | Medium |
| Expertise required | Process engineering | Statistics/ML | Both |
| Maintenance burden | Low | High | Medium |
| Primary applications | Design, scale-up, regulatory | Control, monitoring | Process optimization, soft sensors |
Can data-driven models completely replace first-principles models in chemical process engineering?




The choice of modeling strategy determines the data requirements
| Model Type | Primary Data Needs | Volume | Quality |
|---|---|---|---|
| First-principles | Fundamental properties, structure | Low | Very High |
| Empirical/Regression | Training/validation sets | Medium | High |
| Data-driven/ML | Historical operating data | Very High | Medium |
| Hybrid | All of the above | High | High |
Modelers rarely receive analysis-ready data.
Raw data is rarely clean. Sources of error include:

Raw data requires conditioning before modeling.
Most modeling algorithms expect equal time intervals (constant \Delta t).
Industrial sources have different rates (seconds vs shifts).
Assumption: Process mean/variance are constant.
Points deviating markedly from expected range.
No pipeline should be trusted without validation.


Which is the greater barrier to effective modeling: The Availability of data or the Quality of data?

Process description:

























If our process data is already limited by noise, delays, or slow sampling, when does making a model more “sophisticated” actually start making our control worse? How do we know when to stop adding more complexity?
Data-Driven Modeling of Chemical Processes