Supporting data citation and reuse: introducing the first INFRA-ART versioned data releases

by Ioana Maria Cortea — Published on August 24, 2026 — Reading time: 7 min


Part of the INFRA-ART FAIR Journey article series

Article sections

With the release of the first reference spectral datasets from the INFRA-ART Spectral Library, we have reached an important milestone in our FAIR journey. The core ATR-FTIR and XRF reference collections are now available through Zenodo as independently versioned datasets, providing bulk access to the spectral data and associated metadata and, finally, persistent identifiers for these collections.

In this article, we look at some of the decisions behind the first INFRA-ART data releases and the practical considerations involved in preparing scientific datasets for FAIR publication and reuse.

From spectral files to reusable datasets

The core INFRA-ART spectral datasets are now released through Zenodo, capturing the reference collections at their current stage of development following several years of data acquisition and processing, alongside successive cycles of data and metadata curation.

Getting to this point, however, involved much more than depositing spectral files in a repository. It required decisions about how data should be organized, processed and documented; how raw and processed data should be preserved; how files and metadata should be connected; and eventually, how a continuously evolving database could also provide stable and citable research datasets.

For the INFRA-ART releases, preparing the collections involved establishing consistent SampleIDs and file-naming conventions, reviewing links between spectral files and reference materials, standardizing metadata, developing metadata dictionaries, documenting dataset organization, defining licensing conditions, and determining best file formats.

The resulting releases comprise:

1085 ATR-FTIR spectral records, with both raw and baseline-corrected data

1202 XRF spectral records, provided as raw spectral data

→ sample-level descriptive metadata and metadata dictionaries accompanying both collections

Together, the two datasets represent 1209 unique reference materials, with paired FTIR and XRF data available for 1078 materials. An accompanying Data in Brief article documenting the datasets and their preparation is currently under review.

Further versioned datasets will be released progressively as other collections undergo curation and reach an appropriate level of maturity for independent deposition. Planned releases include Raman and SWIR reflectance spectral collections for reference materials, as well as datasets from composite samples.

Why data citation matters     

Making research data available is only part of making them reusable. Data also need to be identifiable and citable as research outputs in their own right.

The Joint Declaration of Data Citation Principles recognizes data as legitimate, citable products of research and emphasizes that data citation should support credit and attribution, provide access to the data and associated documentation, and enable the specific data underlying research to be identified and verified. Persistent, globally unique identifiers are central to this process.

Example of a data citation based on the Joint Declaration of Data Citation Principles

Data citation therefore does more than acknowledge the creators of a dataset. It establishes a persistent connection between research and its underlying data, supporting discovery, attribution, reuse and reproducibility. For INFRA-ART, however, this raised an additional practical question: how do we provide persistent and citable access to collections that are continuously evolving?

Citing and versioning a living database

The INFRA-ART Spectral Library is a living repository. New records are added, metadata are enhanced, and collections continue to evolve. A citation pointing only to the online database therefore cannot necessarily identify the exact state of a collection used at a particular point in time.

This is a recognized challenge for dynamic datasets. DataCite describes several possible approaches, including citation of a specific subset, a snapshot taken at a defined point in time, citation of the continuously updated dataset with an access date, and timestamped queries against versioned databases.

The RDA Data Citation Working Group similarly emphasizes the importance of identifying and retrieving data as they existed at a particular point in time. Its recommendations include versioning and timestamping dynamic data, with more advanced query-based approaches for persistently identifying specific subsets.

For the current INFRA-ART reference collections, we adopted a snapshot-based approach through versioned data releases. Each Zenodo deposit preserves a defined version of the corresponding collection at the time of deposition and provides it with a persistent DOI. The online INFRA-ART repository can therefore continue to grow and evolve, while each released dataset remains persistently identifiable and citable.

Preparing data for FAIR publication: a few practical lessons

An important lesson from this process is that openly accessible data are not automatically FAIR data. A collection of files, after all, is not necessarily a reusable dataset.

A downloadable file may be accessible, but without sufficient metadata it can be difficult to understand or reuse. Without consistent identifiers and filenames, relationships between files may be unclear. Without documentation, users may not know how the data were generated or processed. And without an explicit license, the conditions for reuse remain uncertain.

FAIRness develops iteratively over time, emerging from many interconnected decisions made throughout the data lifecycle. Data collections evolve, metadata improve, and new standards develop. In that sense, FAIR is perhaps better understood as FAIR + time.

An iterative data lifecycle supporting FAIR data publication, discovery and reuse

Our experience with the first INFRA-ART data releases highlighted several practical considerations that we believe can help streamline future data publication:

  • Plan for reuse from the beginning (FAIR-by-design). Consider early what data, metadata, documentation, standards, and contextual information will be needed to support future understanding and reuse.
  • Preserve original data. Keep an unaltered copy of the original data separate from working and processed versions, ensuring that the source data remain available for future reference and reprocessing.
  • Document data processing: Record the software, processing steps, relevant parameters and other information needed to understand how processed data were derived from the originals.
  • Choose sustainable file formats. Prefer widely supported and open formats while retaining discipline-specific formats when they preserve important scientific information.
  • Organize and name files consistently. Use unambiguous filenames and a logical folder structure suited to your research. Establish conventions early and document them so that others can understand and navigate the dataset.
  • Invest time in metadata and documentation. README files, metadata dictionaries, and clear descriptions of relationships between files are integral components of a reusable dataset.
  • Keep metadata machine-readable. Structure tabular metadata consistently and avoid formatting that makes automated processing difficult.
  • Use an appropriate repository. Repositories such as Zenodo can provide persistent identifiers, citation, versioning, and long-term preservation infrastructure.
  • Version evolving datasets. For living resources, identifiable snapshots can preserve citable states while allowing the underlying resource to continue evolving.
  • Make reuse conditions explicit. Clearly state licensing conditions and provide a recommended dataset citation.
  • Keep track of changes for future versions (change log). Document additions, corrections, metadata improvements, and other changes as they occur, rather than reconstructing them when preparing the next release.
  • Consider publishing a supporting data article. A dedicated data article can complement the repository record by documenting how the data were generated, processed, structured and curated, describing their potential for reuse, and providing a formal scholarly reference linked to the deposited dataset.

Further reading and resources

Ball, A., and Duke, M. (2015). ‘How to Cite Datasets and Link to Publications’. DCC How-to Guides. Edinburgh: Digital Curation Centre. Available online: /resources/how-guides

Cortea, I. M. (2026). From FAIR Principles to Practice: A Case Study of FAIRification in a Heritage Science Data Service. Heritage 9, 228. https://doi.org/10.3390/heritage9060228

Cortea, I. M. (2026). INFRA-ART FTIR Spectral Data Collection – Reference Materials (Version 1.0) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.21993883

Cortea, I. M., and Ghervase, L. (2026). INFRA-ART XRF Spectral Data Collection – Reference Materials (Version 1.0) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.21995097

Perkel, J. M. (2023). How to make your scientific data accessible, discoverable and useful. Nature 618, 1098–1099. https://doi.org/10.1038/d41586-023-01929-7

Rauber, A., et al. (2021). Precisely and persistently identifying and citing arbitrary subsets of dynamic data. Harvard Data Science Review 3. https://doi.org/10.1162/99608f92.be565013

Wild, S. (2025). Need to update your data? Follow these five tips. Nature 643, 868–869. https://doi.org/10.1038/d41586-025-02210-9

Wilkinson, M., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3, 160018. https://doi.org/10.1038/sdata.2016.18

How to cite this resource

Cortea, I.M. (2026, August 24). Supporting data citation and reuse: introducing the first INFRA-ART versioned data releases. INFRA-ART Blog. https://blog.infraart.inoe.ro/2026/08/24/supporting-data-citation-and-reuse-introducing-the-first-infra-art-versioned-data-releases/

Share