Speaker
Description
Scientific facilities are evolving toward automated and data-intensive environments in which experiments continuously generate heterogeneous and large-scale datasets that must be managed under governance requirements. Although metadata catalogues have improved data management, the transient phase between acquisition and long-term cataloguing often remains fragmented, manually handled, script-driven, and beamline-specific. This gap is critical for achieving high-level automatization and especially for autonomous experiments, where decisions must remain reproducible and auditable.
This work presents SIRFlow, a policy-driven framework for data governance automation currently under development for Sirius at CNPEM. SIRFlow operates as an intermediate governance layer between acquisition systems, workflows, and metadata catalogues, coordinating dataset lifecycle states, auditable job and data services events, and policy-aware orchestration.
The framework comprises two main architectural components: a data control plane responsible for policy validation, dataset lifecycle management, and both event and job orchestration; and a data execution plane composed of scalable workers responsible for data movement, processing, streaming, and integration with external systems such as ICAT, high-performance computing (HPC) infrastructures, and AI models. Automation is modeled through auditable job envelopes associated with dataset lifecycle records, enabling reproducibility, safe retries, provenance tracking, and controlled state transitions.
Our key contribution is the explicit distinction between transient and persistent scientific datasets. In the transient phase, datasets may undergo streaming, processing, validation, and curation before promotion to institutional catalogues. Only datasets satisfying governance and quality policies are promoted to long-term systems such as ICAT, reducing storage and catalogue pollution, and enabling robust autonomous experiments.