Speaker
Description
The Shanghai HIgh repetitioN rate XFEL and Extreme light facility (SHINE) will support experiments such as serial crystallography, coherent diffraction imaging, and X-ray spectroscopy. These experiments will generate massive, heterogeneous and time-correlated datasets with scientific and administrative metadata. FAIR data management is therefore essential for facility operation, long-term preservation, data reuse, and future AI-driven scientific discovery.
This presentation describes recent progress in FAIR-oriented data management at SHINE, from policy and metadata standards to system implementation. In collaboration with major synchrotron radiation facilities in China, we contributed to common scientific data policies and metadata standards for large-scale user facilities. Based on these, SHINE has developed facility-specific metadata specifications for XFEL user experiments, covering key entities across the experimental data lifecycle.
An XFEL data management system has been developed based on DOMAS, a scientific data management framework jointly developed by our team and the IHEP, CAS. The system integrates online storage, offline storage, metadata management, and data services. It supports automated migration of experimental data files, metadata extraction and consolidation, controlled data access and data retrieval.
During the joint commissioning with the Spectrometer for Electronic Structure (SES) and Atomic, Molecular, and Optical Science Endstation (AMO) of SHINE, we has identified two practical challenges. The first is consistent permission management for shared storage directories across Linux, Windows, and HPC environments. The second is how to support diverse beamline and endstation DAQ software while maintaining metadata completeness, machine readability, and data traceability. These lessons are important for moving from FAIR data management toward AI-ready datasets.