Speaker
Description
The recent emergence of agentic workflows in scientific research has highlighted the urgent need for high-quality AI-ready experimental datasets for model training and validation. To be AI-ready, these datasets should be described by rich, machine-readable metadata and provenance, following FAIR data principles. Metadata and provenance perform three functions in agentic workflows: they allow datasets to be auto-discovered, they provide context for interpretation of datasets, and they help establish trust in AI models.
At the Cornell High Energy Synchrotron Source (CHESS), we have developed the FAIR Open-Science Extensible Data Exchange Network (FOXDEN), a suite of lightweight, modular data services that supplements researchersʼ existing experimental workflows, allowing them to easily record metadata, provenance, and other research artifacts in real time. FOXDEN is specifically designed to handle large, unportable datasets and heterogeneous use cases. It helps scientists turn their research artifacts into annotated, AI-ready datasets and publish them with Digital Object Identifiers. FOXDEN is also federated for use at multiple sites and facilities. We describe FOXDENʼs architecture and deployment status at both CHESS and the National High Magnetic Field Laboratory, and we present a blueprint for incorporating it into agentic scientific workflows.