Improving FAIR Compliance for High-Dimensional Data via Automated Metadata Extraction
At a glance
- Citations
- 0
- References
- 16
- Comments
- 0
Abstract
High-dimensional datasets are becoming an increasingly vital asset in the machine learning and AI domains due to their ability to capture large volumes of structured, self-descriptive information. Research data repositories play a key role in supporting the documentation, discoverability, and reuse of such data. This paper highlights the growing presence of high-dimensional formats—particularly NetCDF and HDF5—underscoring the need for better infrastructure to support them. We use a large language model to assess the quality of metadata associated with these files and find that user-provided metadata often falls short when compared to the richness of embedded metadata within the files themselves. In response, we implement an automated metadata extraction process during file ingestion, offering a practical pathway to FAIR-ify high-dimensional data. Our empirical analysis and technical solution are demonstrated through integration with the Dataverse data repository platform.
Publication details
- DOI
- 10.1109/escience65000.2025.00090
- OpenAlex
- W4414898447
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.