OFS Validation and Control-System Resiliency Testing
Context
At Lonza (pharmaceutical manufacturing), the control stack included Modicon M580 PLCs, Schneider OFS OPC connectivity, AVEVA Application Server, AVEVA Batch, virtualized Windows Server, and redundant infrastructure components. The work focused on proving that this stack could take disturbances and recover without turning PLC-to-server failures into lasting manufacturing risk.
Problem
PLC–server failures were not just a blank HMI problem. They could produce bad tags, data loss, HELD/ABORTED batches, and failover gaps — symptoms that look like “display issues” but are really control-path and batch-continuity failures.
Role
I led functional and failure-mode testing for the OFS architecture: reviewing expected behavior, running controlled failover and infrastructure tests, monitoring PLC / OPC / Application Server / batch behavior, assessing tag quality and historian data, distinguishing transients from defects, and documenting results for engineering and validation audiences.
Tests and acceptance
Controlled disturbances included:
- OFS/OPC interrupt
- Application Server failover while active
- DA server failover
- VM migration, snapshots, antivirus scan, and server failover scenarios
Acceptance criteria:
- Tags return to Good quality
- Automatic recovery
- Recipes continue
- No persistent HELD/ABORTED batch state
- Missed reads stay within defined thresholds
Result
Takeaway
Test failure modes, not only happy-path function. Measure recovery. Treat automation and IT as one stack. Use the historian to separate display symptoms from real manufacturing risk.