Speaker
Description
The OSCARS project PaN-Finder explores how recent advances in artificial intelligence, particularly large language models (LLMs) and vector embeddings, can be leveraged to enhance FAIRness of open data in the European photon and neutron community. We developed a retrieval-augmented generation (RAG) system that improves data findability and relevance in the search results by enabling both expert-driven, domain-specific queries and more general, non-expert searches.
This work provides new insights into data curation and search methodologies, revealing shared patterns and variety present in open data across facilities and scientific disciplines. It also demonstrates how AI assisted methods can enrich metadata, supporting more accurate relevance ranking and improving scientific data reuse. The PaN-Finder challenges traditional top-down curation models, and promotes a more holistic, data-driven approach better aligned with the current rapidly evolving landscape of data and technology.
In this talk, we present the project’s journey, from conceptual design to implementation, and discuss the key choices that help shaping the current PaN-Finder tool, highlighting lessons learned and future directions in FAIR data.