The Semantic Layer Is the Infrastructure
Scientific data platforms keep getting sold as if storage were the hard part. It is not. The hard part is identity, and identity only behaves when ontologies and master data management are owned, governed, and allowed to do the dull work that a generic warehouse never will.
The week’s familiar misdirection
There is always a fresh round of attention on semantic interoperability, on lakehouse promises, on FAIR language, on ontology governance. This week is no different. The rhetoric says the same thing in a cleaner suit: if we unify the files, the science will unify itself.
It does not. A data lakehouse can concentrate data and compute, and that is useful. But concentration is not meaning. A warehouse can hold assay files, sample records, study metadata, and target annotations in one place and still leave them mutually unintelligible when the naming breaks at the seam between R&D and development.
That seam is where pharma keeps pretending integration is a technical footnote. It is not. It is a question of who owns the identifiers, who governs the vocabulary, and who decides whether a target in discovery is the same entity as the target in translational work or a dangerously similar imposter wearing the same label.
What fails when the semantic layer is implied
Picture a Friday handoff between discovery and translational science. An assay has moved labs, a sample lineage has been reprocessed, a study has been split for regulatory reporting, and three teams are still debating whether the master record is the ELN export, the LIMS row, or the warehouse surrogate key. The meeting is called integration because no one wants to say identity.
This is where generic platforms stop short. They can ingest, curate, and govern data. They can even expose rich metadata and flexible schemas. But if target, indication, assay, sample, and study are only understood by convention, then every downstream use is a local interpretation pretending to be enterprise truth.
That is why semantic interoperability has become such a persistent theme. It is the realization, usually late, that cross system decision making depends on shared meaning, not just shared storage. FAIR language points in the right direction, but FAIR without ownership of the semantics is still a poster, not an operating model.
Ontologies are not decoration
An ontology is not a diagram on a slide. It is the contract that says what a target is, how an indication relates to a study, what an assay measures, and which sample identity survives transfer, reformatting, and analytical convenience.
Master data management does not replace the ontology. It enforces the controlled identities that let the ontology work in production. Together, they make scientific objects stable enough to move across research and development without collapsing into aliases, duplicates, and polite uncertainty.
That stability matters because scientific platforms are increasingly being asked to support the full experimental lifecycle and the full decision lifecycle, not just to archive outputs. Scientific data infrastructures are being framed as the foundation for collecting, storing, distributing, processing, exchanging, and preserving data so that knowledge can actually accumulate.
Accumulation is the point. If the platform cannot preserve meaning across a target that changes annotation, an assay that changes protocol, or a sample that changes custody, then it is not a decision infrastructure. It is a very expensive memory problem.
Governance is the part nobody wants to own
Ontology governance sounds tedious because it is tedious. That is also why it matters. Someone has to own term changes, identifier issuance, equivalence rules, provenance, and the friction between local lab vocabulary and enterprise vocabulary.
The recent emphasis on ontology governance is useful precisely because it makes the hidden labor visible. The promise of federated scientific data collapses the moment each domain owns its own meaning and exports it as if it were universal. Federated systems need federation of semantics, not just federation of access.
I have watched teams buy another dashboard hoping it would resolve a sample identity problem. It never does. A dashboard can only render the confusion more elegantly. Marina will disagree with the decoration, but the architecture always tells the truth eventually.
The real engineering question
The question is not whether a platform can store more. The question is whether the platform can answer, with authority, what a thing is, where it came from, what it was compared against, and whether two systems are talking about the same object or merely using the same word.
That is the decision surface. Target selection. Indication stratification. Assay comparability. Sample lineage. Study reconciliation. If those identities are owned in the semantic layer, then research data becomes navigable across the R&D boundary instead of trapped inside local schemas and export rituals.
If they are not owned, then every downstream AI workflow, every search layer, every supposedly intelligent lakehouse, inherits the same problem and decorates it with more compute.
If this handoff problem is sitting on your desk, write to hello@example.com. We build the systems side of that mess.
References
- Defining Platform Research Infrastructure as a Service (PRIaaS) for Future Scientific Data Infrastructure
- EDS Data Platforms - Linac Coherent Light Source
- Data Terra: a federated research infrastructure transforming Earth ...
- A multi-service data management platform for scientific ...
- Why data platforms are key to scientific progress
- International Journal of Scientific Research in Science, Engineering and Technology,IJSRSET
- Data Science Platforms
- Who Will Keep Research Data Infrastructure Open and ...
- Data Infrastructure | Data Science at NIH
- New data infrastructure initiative will accelerate the ...
- Research collaboration data platform ensuring general ...
- Design and Research of Data-driven Scientific Research
- how Dotmatics Luma and Databricks make AI-ready science a reality
- Jean-Christophe Plantin, Carl Lagoze and Paul N. Edwards
- DataScribe: An AI-Native, Policy-Aligned Web Platform for ...
- A cloud-native approach to scientific data management
- Our approach to data
- Data platforms for open life sciences–A systematic analysis of ... - PMC
- How Modern Big Data Platforms Improve Enterprise ...
- What is Data Infrastructure at large companies and how do ...
