
Modern industrial operations are marvels of instrumentation and automation that generate large and varying streams of process-related data. This combined data is a rich resource, invaluable for making informed operational and strategic decisions, assuming one can reliably access it.
Stone Three’s clients face numerous challenges in collecting and internally democratizing their data to enable this kind of high value decision making. Some clients have multiple operations scattered across the world, requiring centralisation of data into a common endpoint. Other clients have sprawling plants, the results of decades of upgrades, leaving them with a diverse web of instruments and control suppliers who need to talk to each other and form a cohesive operational picture. Old and new plants are facing the question of how they should achieve the goals of data contextualisation and centralisation in the rapidly changing and advancing fields of machine learning and artificial intelligence.
The cohesive consolidation and contextualisation of data must be achieved timeously, cost-effectively, and most importantly, securely. This falls under the twinned domains of Data Science and Data Engineering. Just as a chemist might develop a new process and the chemical engineer implements it on an industrial scale, so does the data scientist extract new insights from data, and the data engineer implements the architecture and solutions to generate these insights reliably and at scale.
Clients had little choice in the past but to use vendor-specific solutions or build bespoke in-house solutions to these problems before the maturation and commodification of data science and data engineering. Often, this would require migrating all operations and systems to a common data or instrumentation platform and using a vendor-specific solution. While this may solve many of the practical problems related to data consolidation, it results in vendor lock-in, expensive and risky migrations, and often onerous licensing fees.
This leaves clients at risk of becoming dependent on a single vendor whose core business is not data science and data engineering, but instrumentation and solving the challenges of distributed control systems. Any change requests fall low on vendor priority lists, and clients may face significant difficulty in recruiting and training resources who specialise in working with and building on these vendor-specific solutions.
It is not viable for most clients to support these kinds of specialised resources, and even the large corporates who’ve traditionally had the in-house expertise for this are finding it difficult to justify the cost and expense of these scarce skills.
This is not to say the industry has been standing still in trying to solve this problem. Through concerted efforts of multiple vendors and industry partners, a cross-vendor real-time communication standard now known as Open Platform Communications (OPC) was developed. The first iterations of OPC were strongly linked with components of Microsoft Windows, but the latest iterations of OPC-UA are platform agnostic, allowing much greater operational flexibility with the use and deployment of OPC.
OPC provides a solution to the first link of the data chain, namely, plant-level inter-communications. Beyond this level, and for historical data, vendor lock-in remains a reality for most clients. However, with the development of data engineering as a general discipline, the historisation of time series data and consolidation from various sources is no longer a specialised issue but a common problem faced by many industries.
This has resulted in the development of numerous open platforms and solutions which can be used to build the complete data chain, from PLC to CEO. The skills and systems involved in this are used across our modern world of commerce, where the only costs are those of compute, storage, and network transit.
Tools such as Microsoft Azure IoT enables the deployment and management of unlimited virtual devices securely across thousands of operations. Stone Three utilises this technology to deploy and maintain OPC data connectors to client operations, which enables a reliable flow of data to IoT Hub, where it can be easily accessed and processed by authorised systems through open standards such as Apache Kafka and Spark.
Databricks, in conjunction with Spark and other open technologies, enables the orchestration of data transforms and management of data lakes through Unity Catalogue, providing a clear indication of data quality through the medallion architecture. This is agnostic of the data layer. Clients can use their own cloud providers or even their own physical hardware to host all their data in perpetuity in common SQL and Parquet formats.
Stone Three uses the medallion architecture to surpass traditional historians and databases to ensure a unified view of all operational data, regardless of the time scale involved. Bronze data represents the rawest form for unlimited reprocessing at any time in the future without risk of irreversible data loss. Silver data is the first level of enrichment, taking Bronze data and transforming it into standardised, human-readable schemas.
Finally, key insights and performance indicators are extracted and processed from Silver into Gold, where it can be used directly by various reporting platforms and stakeholders through open standard-driven interfaces.
This open architecture provides a solid and unbroken chain of data flow from the plant level to any systems, from the lowest to the highest operational levels. Clients should not be expected to limit their options for data processing and reporting in the ongoing industrial data revolution. The open and standards-driven approach that Stone Three has embraced is the future of scalable and reliable data flows from real-time to historical and everything in between.






