Announcing dataset "CORD-19 Named Entities KG": an RDF dataset of named entities identified in the CORD-19 corpus
Franck Michel <[email protected]> Thu, 2 Apr 2020 12:18:30 +0200
| Newsgroups | gmane.org.w3c.public-lod,gmane.org.w3c.semantic-web |
|---|---|
| Message-ID | <[email protected]> |
Dear colleagues, In order to foster innovative work based on the cross-linking of COVID-19 literature with the Data Web, we (Wimmics team, Inria <https://team.inria.fr/wimmics/>) are in the process of generating an RDF dataset describing the named entities identified in the research papers of the CORD-19 <https://pages.semanticscholar.org/coronavirus-research> corpus. To identify and disambiguate the named entities, we are using NCBO BioPortal annotator <http://bioportal.bioontology.org/annotatorplus>, Entity-fishing <https://github.com/kermitt2/entity-fishing> (links to Wikidata) and DBpedia Spotlight <https://www.dbpedia-spotlight.org/> (links to DBpedia). We are also taking care of linking to other related works such as CORD-19-on-FHIR <https://github.com/fhircat/CORD-19-on-FHIR> and COVID-19 Literature KG <https://www.kaggle.com/group16/covid19-literature-knowledge-graph>. We shall release this dataset soon, as n RDF dump as well as through a dedicated SPARQL endpoint. Stay tuned! Regards, Franck. -- signature Franck MICHEL - CNRS research engineer Université Côte d’Azur, CNRS, Inria I3S laboratory (UMR 7271) [email protected] <mailto:[email protected]> - +33 (0)4 8915 4277