Hi
@Merel et al. from the Kedro team.
We recently published a new repository of an open-source data resource for biomedical science constructed using Kedro. This work was done at Harvard University.
Here is the released links for OptimusKG:
• Website:
optimuskg.ai,
• Repository:
https://github.com/mims-harvard/OptimusKG
• Paper:
https://arxiv.org/pdf/2604.27269
• Client:
https://pypi.org/project/optimuskg
• Data:
https://doi.org/10.7910/DVN/IYNGEV
Highlights
• A modern biomedical knowledge graph with molecular, anatomical, clinical, and environmental modalities.
• Integrates 65 heterogeneous resources grounded with 18 ontologies and controlled vocabularies using the BioCypher framework and the Biolink Model.
• Contains 190,531 nodes across 10 entity types, 21,813,816 edges across 27 relation types, and 67,249,863 property instances encoding 110,276,843 values across 150 distinct property keys.
• Independently validated using PaperQA3, a multimodal agent that retrieves and reasons over scientific literature.
• Reproducible, deterministic and infrastructure-agnostic data pipeline with parallel execution.
• Distributed as Apache Parquet files and downloadable via the optimuskg python client.
Do you think this might be an interesting initiative to showcase the use of Kedro in your website / documentation?