Deepak Unni
Profile
12 years of experience developing and maintaining tools that support biological and biomedical research. I currently work as a Scientific Coordinator at the Swiss Personalized Health Network at SIB Swiss Institute of Bioinformatics, with a keen interest in biological data modeling, knowledge representation, open-source software development, ETL pipelines, ontology-driven data integration, and knowledge graphs. I am particularly interested in making datasets and other digital assets FAIR by leveraging rich metadata to improve their discovery, interoperability, and reuse across research communities.
Technical Skills
Languages
Python, Java
Semantic Web Standards
RDF(S), OWL, SPARQL, SHACL, ShEx, ODRL
Vocabularies
Dublin Core, DCAT, DCAT-AP, HealthDCAT-AP, VoID, PROV-O, SKOS
Ontologies
GO, HP, Mondo, OMO, SULO
Modeling frameworks
LinkML
Data Models
Biolink Model, InterMine, SPHN RDF Schema
Graph Databases
GraphDB, QLever, Apache Jena Fuseki, Neo4j
Health Data Standards
SNOMED CT, LOINC, ICD-10-GM, HL7 FHIR
Backend Frameworks
FastAPI, Flask
Experience
Aug 2022 – Present
Basel, Switzerland
Scientific Coordinator
SIB Swiss Institute of BioinformaticsSPHN is a national initiative led by the Swiss Academy of Medical Sciences and the SIB Swiss Institute of Bioinformatics. Broadly speaking, SPHN develops, implements, and validates coordinated data infrastructures that make health-related data available, interoperable, and shareable for research in Switzerland. As a Scientific Coordinator at the SPHN Data Coordination Center, I focus on the semantic and metadata foundations of these FAIR data infrastructures, where my work includes:
- ·Engaging with clinical, research, and data provider stakeholders to capture domain semantics and formalize them in RDF
- ·Leading the adoption of semantic web technologies for data modeling, knowledge representation, and validation of health-related data
- ·Contributing to the design and evolution of the SPHN RDF Schema, integrating standard terminologies and ontologies to ensure semantic interoperability across Swiss hospitals
- ·Supporting the development and maintenance of open-source tools and ETL pipelines that help data providers produce, validate, and deliver FAIR-compliant data
- ·Building and deploying the SPHN Metadata Catalog, an end-to-end platform that makes SPHN data assets discoverable through rich, standards-based metadata
- ·Interfacing with local and international initiatives to ensure that SPHN Metadata Catalog is interoperable with other FAIR data infrastructures and aligned with global standards
- ·Promoting FAIR principles through documentation, training, and engagement with data providers, researchers and the broader scientific community
Apr 2021 – Jul 2022
Heidelberg, Germany
Software Developer
European Molecular Biology LaboratoryThe German Human Genome-Phenome Archive (GHGA) is a national initiative within Germany's National Research Data Infrastructure (NFDI) that provides a secure, FAIR-compliant platform for storing, sharing, and analyzing human omics data. In this role, I:
- ·Collaborated with the Architecture Working Group to design and establish the GHGA platform, building on existing frameworks and well-established community standards
- ·Led the Metadata Task Force to design, curate, and maintain the GHGA Metadata Schema, enabling consistent description and discovery of human omics datasets
Feb 2018 – Mar 2021
Berkeley, CA, USA
Staff Software Developer
Lawrence Berkeley National Laboratory- ·Contributed to the design and development of the Monarch API, improving access to integrated genotype–phenotype data
- ·Collaborated with the UI development team to enhance the Monarch web interface and user experience
- ·Partnered with the data ingest team to explore new strategies for building and improving the integrated Monarch knowledge graph
- ·Extended the scope and functionality of the GO API in collaboration with the development team
- ·Worked with scientists and developers to apply next-generation GO annotations to model biological systems
- ·Helped define community standards for representing biological and biomedical data as knowledge graphs
- ·Modeled diverse types of biomedical data and coordinated ETL efforts across the project
- ·Developed tools for exchanging data between heterogeneous knowledge graph initiatives
- ·Led the development and evolution of the Biolink Model, a universal schema for knowledge graphs in clinical, biomedical, and translational science
- ·Built a flexible ETL pipeline for generating a knowledge graph focused on COVID-19 datasets
- ·Assisted in setting up RDF and Neo4j endpoints for the COVID-19 knowledge graph
- ·Engaged with researchers to prioritize potential therapeutic targets for COVID-19
- ·Integrated the COVID-19 knowledge graph into the National COVID Cohort Collaborative (N3C) and the COVID-19 International Research Team (COV-IRT)
Feb 2014 – Feb 2018
Columbia, MO, USA
Research Analyst
University of Missouri- ·Development and maintenance of Apollo, a collaborative genome annotation editing tool
- ·Implementation and maintenance of tools for the Bovine Genome Database and Hymenoptera Genome Database
- ·Implementation, deployment and maintenance of a data warehouse for Bovine genomics (BovineMine), Hymenoptera genomics (HymenopteraMine), and Maize genetics (MaizeMine)
Aug 2013 – Dec 2013
Atlanta, GA, USA
Graduate Research Assistant
Georgia Institute of TechnologyJun 2013 – Aug 2013
NY, USA
Bioinformatics Intern
Regeneron PharmaceuticalsJan 2013 – May 2013
Atlanta, GA, USA
Graduate Research Assistant
Georgia Institute of TechnologyEducation
2012 – 2013
MS Bioinformatics
Georgia Institute of Technology, USA
2007 – 2011
B. Tech Bioinformatics
Dr. D. Y. Patil Vidyapeeth, Pune, India
Publications
2026
2025
Moxon S. A. T. et al
LinkML: An Open Data Modeling FrameworkVasilevsky, N. A. et al
Mondo: integrating disease terminology across communities2024
Callahan, T. J. et al
An open source knowledge graph ecosystem for the life sciencesSIB Swiss Institute of Bioinformatics RDF Group Members
The SIB Swiss Institute of Bioinformatics Semantic Web of data2023
Caufield, J. H. et al
KG-Hub-building and exchanging biological knowledge graphsTouré, V. et al
FAIRification of health-related data using semantic web technologies in the Swiss Personalized Health NetworkScientific Data, 10, 127
2022
Hoyt, C. T. et al
Unifying the identification of biomedical entities with the BioregistryScientific Data, 9, 714
Unni D. R., Moxon S. A. T. et al
Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational scienceClinical and Translational Science, 15(8), 1848–1855
2021
The Gene Ontology Consortium
The Gene Ontology resource: enriching a GOld mineNucleic Acids Research, 49(D1), D325–D334
Reese, J. T. et al
KG-COVID-19: A Framework to Produce Customized Knowledge Graphs for COVID-19 ResponsePatterns, 2(1), 100155
2020
Haendel, M. A., Chute, C. G. et al
The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deploymentJournal of the American Medical Informatics Association, 28(3), 427–443
2019
Shefchek, K. A. et al
The Monarch Initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across speciesNucleic Acids Research, 48(D1), D704–D715
2018
Biomedical Data Translator Consortium
Toward A Universal Biomedical Data TranslatorClinical and Translational Science, 12(2), 86–90
Biomedical Data Translator Consortium
The Biomedical Data Translator Program: Conception, Culture, and CommunityClinical and Translational Science, 12(2), 91–94
2015
Elsik, C. G. et al
Hymenoptera Genome Database: integrating genome annotations in HymenopteraMineNucleic Acids Research, 44(D1), D793–D800
Elsik, C. G. et al
Bovine Genome Database: new tools for gleaning function from the Bos taurus genomeNucleic Acids Research, 44(D1), D834–D839