Building the open data infrastructure that biomedical science needs.
The FAIR Data Informatics (FDI) Lab at the University of California, San Diego is a leader in developing and providing novel informatics infrastructure and tools for making data Findable, Accessible, Interoperable, and Reusable (FAIR).
Our work spans the development of technology, services, and policy — thereby improving accuracy, reproducibility, and analysis in biomedical research. We bring together expertise in biology, computer science, library science, and ontology engineering to build infrastructure that works for scientists across disciplines.
FDI Lab originated in 2008 as the Neuroscience Information Framework (NIF), an NIH Blueprint Consortium initiative. NIF began cataloging neuroscience resources in 2006 and created one of the first cross-database search engines for biomedical research, indexing hundreds of databases, tools, and resources.
Building on this foundation, the lab developed SciCrunch — a customizable platform for creating domain-specific data portals. SciCrunch now powers infrastructure for multiple biomedical communities, including the dkNET portal for digestive, diabetes, and kidney disease research.
A particularly impactful initiative has been the Resource Identification Initiative (RRID), which provides unique, persistent identifiers for research resources including antibodies, cell lines, organisms, and biosamples cited in scientific literature — dramatically improving reproducibility across the published record.
FDI Lab actively participates in open-science organizations including FORCE11 and the International Neuroinformatics Coordinating Facility (INCF).
The FAIR Guiding Principles, published in Scientific Data in 2016, provide a framework for making data and digital assets maximally reusable. FDI Lab was an early adopter and active contributor to FAIR practice across the biomedical community.
Data and metadata are assigned globally unique, persistent identifiers and are deposited in searchable resources. Rich metadata helps both humans and machines discover datasets efficiently.
Once found, data and metadata can be retrieved using open, standardized protocols. Authentication and authorization rules are transparent, and metadata remain accessible even when data are not.
Data use a formal, accessible, shared, and broadly applicable language for knowledge representation. Vocabularies follow FAIR principles, and qualified references to other data are included.
Data have rich, accurate metadata; a clear, accessible data usage license; detailed provenance; and meet domain-relevant community standards to support replication and reuse.