A Plant Can Have More Data Than Knowledge
Modern industrial plants generate enormous amounts of data.
Temperatures, pressures, flow rates, valve positions, laboratory results, energy consumption and production rates may be stored every minute—or every second. After several years, a process historian can contain millions of observations.
This creates an understandable expectation:
If we have enough historical data, the optimum must already be hidden somewhere inside them.
Unfortunately, this is often not true.
A process historian tells us where the plant has operated. It does not automatically tell us where the plant should operate. The difference between those two statements is fundamental.
Historical data are produced while operators are trying to maintain stable production, not while engineers are systematically exploring the process. Standard operating procedures keep the most important factors within narrow ranges. Control systems suppress disturbances. Operators avoid conditions that might threaten quality, safety or throughput.
The resulting dataset may be large, but the part of the process it describes can be surprisingly small.
Historical Data Describe Normal Operation—Not The Full Response Surface
Imagine a plant that has operated for five years at approximately the same temperature, pressure and feed ratio.
There may be thousands of measurements at these conditions. However, the data provide little evidence about what would happen if the temperature were changed by five degrees, the feed ratio were adjusted or two factors were changed together.
The historian has observed repetition, not exploration.
In Full Scale Plant Optimization in Chemical Engineering, I describe several recurring weaknesses of historical process data:
- factors vary only across narrow operating ranges
- input variables are often strongly correlated or confounded
- changes are not randomized
- measurements may be missing, delayed or erroneous
- operating conditions and laboratory results may be poorly synchronized
In many cases, more numbers are available than useful information. Data cleaning alone may consume much of the analytical effort before process modelling can even begin.
This does not make historical data useless. Historical data are highly valuable for monitoring, diagnosis, soft sensing, identifying recurring disturbances and generating hypotheses.
The problem begins when we ask them to answer a question they were never designed to answer:
What will happen outside the operating conditions we have already observed?
Correlation Is Easy To Find—And Easy To Misinterpret
Industrial process variables rarely move independently.
When production rate rises, feed flow, steam demand, cooling-water demand and several temperatures may rise at the same time. When operators respond to a disturbance, they may change multiple setpoints together. Raw-material quality may affect both the operating conditions selected by the operator and the resulting product quality.
A model trained on these data may detect strong statistical relationships. But it may not be clear which factor actually caused the response.
For example, a model may conclude that higher steam use improves yield. In reality, steam use may simply increase whenever production rate increases. The apparent relationship is real in the dataset but misleading as an operating recommendation.
This is one of the central limitations of passive observation:
The process decides which combinations of factors appear in the data—not the engineer.
Without deliberate variation, important effects and interactions can remain hidden.
A Predictive Model Can Be Accurate And Still Be Wrong For Optimization
This distinction is particularly important:
Prediction and optimization are not the same task.
A model may predict normal operation with impressive accuracy because new observations are similar to those used during training.
Optimization asks a more difficult question. It searches for combinations of operating conditions that may not yet have occurred. The optimizer moves toward the boundaries of the dataset—and often beyond them.
That is where model confidence can become dangerous.
In the rayon example described in my book, five years of production history produced more than 15,000 records covering 16 process factors and several fiber properties. A neural-network model could be trained to predict fiber strength and explore possible operating settings.
Yet the model inherited the limitations of the historical dataset. The most important factors had varied only across the narrow ranges used during routine production. Extending the model into a wider experimental space required extrapolation, which carries risk.
A model may interpolate well between known observations while behaving unrealistically outside them. This is not necessarily a software error. It is a limitation of the evidence available to the software.
The Optimizer May Find The Model’S Weakness—Not The Plant’S Optimum
Optimization software is designed to search efficiently.
If a model contains a small error, the optimizer may exploit it. It can recommend an operating point that looks excellent mathematically because it lies in a region where the model is too optimistic.
The result may satisfy every equation in the software while failing in the plant.
This can occur in both main categories of process models.
Data-Driven Models
Machine-learning and empirical models learn relationships from available data. They may perform poorly when:
- the new conditions lie outside the training range
- correlations change
- a previously constant variable begins to move
- an unmeasured disturbance becomes important
- the process changes over time
First-Principles Models
Mechanistic models are based on physical and chemical knowledge, but they also require assumptions.
Heat-transfer coefficients, kinetic parameters, phase equilibria, equipment efficiencies and fouling conditions may not be known exactly. Equipment geometry may be simplified. Side reactions or transport limitations may be omitted.
A physically reasonable model can therefore still differ from the real plant.
The question is not whether a model is useful. The question is:
For which decision, under which conditions and with what uncertainty is it useful?
A Digital Twin Does Not Remove Model Uncertainty
A digital twin improves the connection between a physical process and its digital representation. It may combine real-time plant data, mechanistic models, analytics and continuous parameter updates.
That can make the model more relevant and more responsive.
But the term “digital twin” should not create a false sense of certainty. The digital representation is still based on data, assumptions, equations, sensors and software. Each can introduce uncertainty.
Credible manufacturing digital twins therefore require verification, validation and uncertainty quantification throughout their lifecycle. A twin must be evaluated not only for whether its software was implemented correctly, but also for whether its outputs remain trustworthy for the intended real-world decision.
A digital twin can therefore be synchronized with the plant and still be wrong about an operating condition it has never observed.
It can reflect the current state precisely while predicting the response to a new setpoint inaccurately.
The Optimum Itself Is Moving
Even a model that was accurate when commissioned may become less accurate over time.
Industrial processes change continuously:
- raw-material quality varies
- catalysts age
- heat exchangers foul
- tools and liners wear
- instruments drift
- equipment becomes misaligned
- ambient conditions change
- upstream and downstream constraints shift
The optimum identified last year may not be the optimum today.
This is why optimization cannot be treated as a one-time calculation. The process must continue to generate new evidence.
A digital twin that is not challenged by reality can gradually become a digital memory of how the plant used to behave.
The Answer Is Not To Abandon Models
Models remain among the most valuable tools available to chemical engineers.
They can organize knowledge, screen alternatives, identify promising factors, reduce the number of physical trials and help engineers understand interactions that would otherwise remain invisible.
But a model should be treated as a hypothesis—not as the final authority.
A reliable optimization workflow separates two roles:
The model proposes. The plant confirms.
Model-based software can suggest a promising direction. Small, controlled changes in the real process can then determine whether the predicted improvement actually exists.
The observations from the plant update our knowledge, correct the model and guide the next step.
This turns modelling from a one-time exercise into a learning loop.
Evop Closes The Gap Between The Model And The Plant
Evolutionary Operation was created for this exact environment.
EVOP introduces small, planned changes during normal production. These changes are repeatedly tested so that their effects can be separated from industrial noise.
The objective is not to replace process models. It is to provide the missing experimental evidence.
A model may indicate that temperature should be increased and feed concentration reduced. EVOP can test small versions of these changes within predefined quality and safety limits.
Three outcomes are possible:
1. The plant confirms the predicted improvement. 2. The plant shows that the effect is smaller than expected. 3. The plant contradicts the model.
All three outcomes are useful.
Confirmation creates confidence. A smaller effect improves calibration. A contradiction prevents a potentially expensive operating decision.
From A Static Digital Twin To An Experimental Twin
The most powerful future architecture is not a digital twin working alone.
It is a digital twin connected to a controlled experimental learning system.
Such a system can:
- use historical data and physical knowledge to identify promising directions
- quantify uncertainty around model predictions
- reject recommendations outside the safe operating envelope
- introduce small operating changes
- evaluate their effect against normal process noise
- update the model with newly generated evidence
- continue as the process changes
This transforms the digital twin from a passive representation into an active learning partner.
The difference is significant:
A conventional model predicts what the process might do. An active learning system verifies what the process actually does.
The Role Of Pia
PIA—the Process Improvement Agent—builds on this combination of models and full-scale experimentation.
PIA does not assume that historical data contain the complete answer. It also does not assume that every model recommendation is correct.
Instead, PIA uses available data and models to guide safe exploration around the current operating point. Small changes are performed within defined constraints. Their effects are repeated and statistically evaluated. Only improvements supported by real process evidence are adopted.
This creates a continuous correction mechanism between the digital representation and the physical plant.
The objective is not to compete with digital twins.
It is to make them more trustworthy.
Start With One Unanswered Question
A plant does not need a perfect model to begin learning.
A practical first application can start with:
- one important production response
- two or three adjustable factors
- a clearly defined safe operating region
- reliable measurements
- a model or engineering hypothesis worth testing
Historical data can tell us where to look.
A digital twin can tell us what to expect.
But only the real process can tell us whether the improvement is true.
The model is a map. The plant remains the territory.