Articles
14 minutes
Copy Link
Build vs. Buy a Manufacturing Data Integration Layer: A Decision Framework for Legacy Plants
TL;DR
Build a custom integration when you have a few stable interfaces and funded staff to test, document, and maintain them.
Buy connectors, edge software, or middleware when equipment varies across sites or connector upkeep exceeds your staff’s capacity. Test the choice in a representative legacy plant before expanding it.
Add a shared ontology when equipment, process, and business data need consistent meaning across systems or facilities. Documented mappings can serve a narrow, single-purpose flow.
Add Humble downstream when trustworthy, contextualized data can support scheduling, root cause analysis, and operational recommendations. Humble does not replace the integration or semantic modeling layers.
Treat manufacturing data integration as a set of layered decisions. A plant can use custom code, bought infrastructure, and a decision layer together.
Why legacy plants get this decision wrong
Legacy plants often start a manufacturing data integration project with a specific operational question, such as which orders a line can complete this week. The data needed to answer it may sit in a PLC, an older SCADA system, and an ERP. Moving those records into one place does not establish whether the machine ID in SCADA matches the work center in ERP or whether a reported stop counts as planned downtime.
Three jobs need separate owners. Connectors collect signals and records. Semantic modeling gives equipment, events, and orders consistent meaning. Decision tools use that context to help people assess schedules or investigate a delay. For example, a scheduling tool cannot reliably account for downtime if one plant records maintenance stops under a different code than another plant.
A decision tool cannot repair missing connections, and a data pipeline does not make inconsistent equipment or event definitions agree. Before choosing software, trace one useful decision back to its source records. Identify who maintains each connection, who owns the mappings, and who approves the resulting action. That exercise shows which layer needs work first without assuming the plant must replace its existing systems.
The decision matrix for build, buy, and decision layer
Compare build and buy options by the work each layer must do, not as substitutes for one another. A legacy plant may need more than one row of this matrix. You might keep a stable custom feed, buy connectors for varied equipment, use a shared model across facilities, and add decision support for scheduling. Each row solves a different problem.
Approach | Primary job | Best fit | Total effort | Connector maintenance | Semantic modeling | Latency and outages | Governance | Cybersecurity | Lock-in | Time to value | Scale across facilities |
|---|---|---|---|---|---|---|---|---|---|---|---|
Custom integration | Exchange data between specific systems | Few stable interfaces with funded support | Initial build plus testing and upkeep | Your engineers maintain each interface | Your engineers document mappings | Test response time and recovery | Assign local owners | Maintain access controls and patches | Check code and interface portability | Can be quick for one narrow flow | Repeat site-specific build and support work |
Bought edge or middleware | Connect devices and transport data | Varied equipment or frequent connector changes | Commissioning, licensing, and site support | Vendor supports covered connectors; you configure sites | You still own business definitions | Test buffering, recovery, and site requirements | Set site access rules | Review patching and network boundaries | Check data export options | Depends on connector coverage | Test connector reuse at a second site |
Bought semantic or industrial data platform | Maintain shared equipment and business meaning | Definitions must work across systems or plants | Model design, mapping, and version upkeep | Confirm whether connectors are included | Your data owners govern definitions | Check access to context during outages | Assign cross-site model owners | Review access and tenancy | Check model export and portability | Requires agreement on definitions | Test local variations against the shared model |
Downstream decision layer | Recommend operational actions | Trusted context supports scheduling or analysis | Use case setup, validation, and review | Upstream owner maintains connections | Upstream owners maintain context | Set freshness limits for each decision | Operators approve actions | Restrict access to needed data | Check export of records and reasoning | Depends on upstream readiness | Revalidate decisions and constraints at each site |
Total effort includes discovery, field mapping, commissioning, exception handling, monitoring, version updates, governance changes, and incident response. License cost alone will not tell you who handles a broken connector or an unmapped production event.
Use a representative plant pilot to test the assumptions behind each row. Ask who owns failures, how data behaves during an outage, and whether another facility can reuse the interfaces and definitions. A decision layer should enter the evaluation only when the upstream data gives its recommendations enough context to be useful.
When to build a custom integration
Build a custom integration when a small number of stable systems need to exchange a narrow set of data and you have funded engineering capacity to maintain the connection. For example, a plant might send machine status from one established controller interface to its MES. If you can document the field mappings, test failures, and assign an owner, a custom connection may require less work than introducing a platform.
Budget for maintenance as well as initial development. An equipment upgrade can change a field name or data format while leaving the connection running. Your engineer needs monitoring that catches missing or incorrect records, a test environment for interface changes, and a recovery plan for outages. Documentation must explain what each field means, which version the connection supports, and who responds when it fails.
Reconsider the build choice when you add equipment types, interfaces, or facilities. Each new connection adds code and support obligations. If local engineers cannot keep mappings, tests, and on-call ownership current, bought connectors or middleware may be the more sustainable choice.
When to buy an integration, edge, or middleware platform
Buy an integration, edge, or middleware platform when maintaining the required connections would exceed your engineering capacity. A custom connector may serve one stable machine, but a mix of devices and software versions creates recurring work after upgrades. If you plan to reuse connections across facilities, compare that maintenance burden with the cost and effort of adopting a platform.
These products handle different parts of the data path. Siemens documents an Industrial Edge connector for Beckhoff ADS, while its Industrial Information Hub can transform tag values into attributes. AWS IoT Greengrass provides an edge runtime and management service, and HiveMQ provides MQTT messaging for transport, as described in its manufacturing data sheet. None of those roles, by itself, establishes shared business meaning for the data or decides what an operator should do.
Before buying, ask for the vendor’s supported connector inventory, including equipment and software versions. Test what happens when a device or network connection goes down, how the platform reports missed data, and which export formats let you use that data elsewhere. Run those checks on representative legacy equipment at more than one facility if you expect to scale. A platform saves maintenance work only when its supported connections and operating behavior fit the equipment you actually run.
When does a plant need an industrial ontology or semantic model?
A plant needs a shared industrial ontology when several systems or facilities must interpret equipment, process, and business data consistently. Collecting a machine signal answers what value arrived. Semantic modeling establishes which machine produced it, what the value represents, and how that machine relates to a production order. A decision application can then use that context to recommend an action.
For one stable connection, documented mappings may be enough. If a single machine sends downtime codes to one reporting application, you can record what each code means, who owns the mapping, and how changes are tested. A shared ontology earns more of its implementation and governance cost when plants use different names for equivalent assets or when scheduling, quality, and maintenance applications need the same definitions.
The OPC Foundation’s ISA 95 common object model offers a standards reference for representing manufacturing objects. Using ISA-95 object models can help you define manufacturing objects consistently, but you still need to decide how site-specific assets, orders, and events relate. You still need to define and maintain relationships among equipment, products, orders, process steps, and events where your applications depend on them. Humble’s ISA-95 and industrial ontologies guide examines where the standard’s object models end and site-specific relationships begin.
Before expanding a model across plants, test one cross facility question. For example, can both plants identify which orders ran on presses with the same type of fault, despite different tag names and order codes? If answering requires someone to reinterpret each plant’s data by hand, a governed shared model may be worth building.
Where Humble fits in an industrial ontology architecture
Humble sits downstream of a plant’s integration and semantic modeling infrastructure. Humble uses contextualized equipment, process, and business data to support scheduling, root cause analysis, and operational recommendations. In this architecture, Humble provides decision intelligence after the plant has established what its data refers to.
For example, a plant’s semantic model can connect a machine reading to an equipment ID, a work order, and a product. For scheduling, ask Humble to demonstrate how users describe constraints in natural language, how a schedule responds when those constraints change, and how operators can inspect the evidence behind a recommendation. Test those capabilities against your own MES or ERP data before relying on them. The existing MES or ERP remains the system of record.
Humble is not an ontology platform. It does not substitute for semantic modeling, connectors, middleware, industrial edge, MES, or ERP. If your plant cannot reliably connect a machine reading to the relevant order, Humble does not remove the need to establish that connection first.
Reference architecture for a legacy plant
A legacy plant should keep machine control separate from the data path used for analysis. PLCs control equipment, while SCADA collects and displays operating signals. Edge software or connectors read approved signals from those systems and older devices, then send the data through transport or middleware. Before routing a feed beyond the plant, decide which functions must keep working during a network outage.
A semantic layer gives the collected signals usable meaning. It maps a tag to a specific machine and measurement, then relates that reading to work orders or materials recorded in the MES or ERP. Those applications remain systems of record. Documented mappings may suffice for one stable workflow, while a shared model helps several plants interpret equipment and process data consistently. Humble’s ISA-95 manufacturing integration architecture guide goes deeper on system ownership and data flow between those layers.
Downstream, Humble can use contextualized data to recommend schedule changes or help investigate production disruptions. Operators review recommendations against current constraints and follow the plant’s approval process before changing operations. Decision support must not bypass control and safety boundaries to actuate equipment directly.
For a stable legacy interface, a plant can build its own connector. It can buy middleware where device variety makes connector maintenance harder, then add semantic modeling and decision support as the use case requires. Assign an owner to each layer's access, mappings, and failures.
Failure modes to plan for
A connector can stop delivering usable data after an equipment upgrade, even if the machine keeps running. You should test supported versions before upgrades and alert on stale readings and missing records. Without those checks, silent data drops can reach scheduling or analysis tools before anyone notices.
Shared tags can lose the context that gives a reading meaning. A normalized temperature value, for example, needs its unit, equipment identity, timestamp, and relevant process step. Across plants, the same tag may describe different operating states. Assign an owner to each mapping and version changes so local definitions do not drift unnoticed.
A single-site pilot cannot establish that the same model will work when another plant uses different equipment names, production steps, or access rules. Test a second, meaningfully different site before treating the model as reusable. Track exceptions by owner and resolution time. Otherwise, an unbounded queue can hide mapping errors that require changes to the integration. The multi-plant manufacturing system integration guide covers how to carry shared definitions into later sites without assuming identical local connectors.
Keep the model proportional to the job. A documented mapping may serve one narrow data flow better than a plantwide ontology that nobody can maintain. Likewise, a decision tool cannot repair missing connectors or supply lost context on its own. Verify that each upstream layer delivers trustworthy, contextualized data before relying on its recommendations.
Governance, cybersecurity, and vendor lock-in across facilities
Multi-site governance depends on who owns shared definitions and who can change them. Assign an owner for equipment mappings and require each facility to review changes before they enter a shared model. Otherwise, two plants may use the same field name for different measurements, and downstream reports or recommendations can combine data that does not mean the same thing.
Cybersecurity reviews must account for plant operations, not just access to data. The NIST guide to operational technology security notes that OT has distinct performance, reliability, and safety requirements. For each integration, ask where it runs, which accounts can change it, who patches it, and how you test changes without bypassing plant approval procedures. Apply the same review to custom code, purchased middleware, and downstream decision tools.
Portability becomes easier to judge when you test it with your own data. Ask vendors whether you can export raw records, mappings, and semantic models in usable formats. Check whether model version history lets you reconstruct how a field was defined at a given time. If multiple facilities share a platform, ask how tenancy and access controls keep local data and permissions separate while allowing approved cross-site analysis.
How to run the evaluation at your plant
Run a short pilot on a production line that reflects the interfaces and constraints you expect to face elsewhere. Include one legacy device, a connection to the existing MES or ERP where relevant, and a data field whose meaning changes across equipment or sites. Define the decision the data must support before connecting anything. A pilot that only moves tags cannot prove that operators will receive usable context.
Test a custom integration and a bought option against the same conditions. Check whether each preserves equipment and process context, detects missing records, and recovers after an interruption without silently losing data. Record the work required to map fields, handle exceptions, monitor the connection, and approve access. If you plan to reuse the integration across facilities, test a second site or document which assumptions remain unproven.
Ask your internal builders who will be on call after commissioning and where they can test upgrades without affecting production. Have them specify compatibility checks and a procedure for restoring service and reconciling missed records. Compare those answers with the vendor's support responsibilities, then choose the option your plant can maintain after the pilot ends.
Book a Call with Humble
If you’re evaluating how your plant’s integration and semantic modeling work should support operational decisions, book a call with Humble. Bring your current architecture and a scheduling or root cause problem. The discussion can focus on whether a downstream decision layer fits your existing systems and what data it would need.
See If Humble Fits Your Floor
If your plant already has trustworthy, contextualized data and you want to use it for scheduling or root cause analysis, take Humble’s fit test. It offers a low-commitment way to check whether a decision layer fits your operations. If you still need connectors or shared data definitions, address those first.
Stay Current on Manufacturing Data Strategy
Get practical guidance on manufacturing data integration and semantic modeling in the Humble newsletter.
Frequently asked questions
Do I need an ontology before I can use an AI decision layer?
No. A narrow use case can rely on documented mappings if they give the decision layer reliable equipment, process, and order context. A shared ontology becomes useful when several systems or plants need to interpret that context consistently.
What is the difference between middleware and an industrial data platform?
Middleware moves and transforms data between systems. An industrial data platform may also organize data around shared equipment and process definitions, but you should verify its modeling and governance capabilities rather than assume it provides them.
When does a custom integration outperform a bought platform?
Custom code can be the better choice when you have a few stable interfaces and funded engineers who can test, document, and maintain them. A bought platform becomes more attractive when connector upkeep or reuse across facilities exceeds that capacity.
Does Humble replace MES or ERP?
No. MES and ERP remain systems of record for production and business data, while Humble supports operational decisions using data from existing systems.
How does Humble work with an existing data integration layer?
Your integration layer collects and moves data, while your mappings or semantic model give that data operational meaning. Humble uses the contextualized data for scheduling, root cause analysis, and recommendations. It does not replace semantic modeling, connectors, middleware, or industrial edge infrastructure.