Session 5
Data modelling: Linked Open Data
Beretta Taught by Francesco Beretta.
In this session
- Introduction to RDF and ontologies
- Discussion of (Beretta 2024) and the SDHSS ecosystem
- Hands-on exercise with Wikidata — producing a CSV you will reuse in Session 8
Linked Open Data
- The five stars of open data, and where most humanities data actually stops
- RDF: everything is a triple — subject, predicate, object
- URIs as identifiers: why
Q64is more useful than “Berlin” - Ontologies as shared vocabularies, and the cost of agreeing on one
- SPARQL, briefly — enough to read a query, not yet to write one
Discussion: Beretta and information production
Beretta’s argument is that modelling should capture not only what was the case but who asserted it, when, and on what basis. Points for discussion:
- What changes when the source of a statement becomes part of the model?
- Why is this particularly acute in historical research?
- What does an ontology ecosystem give you that a bespoke model does not?
Exercise: Wikidata
Working in the Wikidata Query Service:
- Person → birthplace. Take a set of people relevant to your topic and retrieve their places of birth.
- Members → organisation. Retrieve the members of an organisation, or the affiliations of a group of people.
- Export. Produce a CSV from the result.
ImportantKeep this CSV
The CSV you produce today is the input for the network analysis in Session 8. Save it somewhere you will find it again, and put it in your project repository.
TipIn parallel in the Lab
The DH Lab goes further into formal ontologies — RiC-O, LRM and CIDOC-CRM.
Reading for Session 6
References
Beretta, Francesco. 2024. ‘Conceptualising Information Production in the Context of the SDHSS Ontology Ecosystem’. Methodos. Savoirs Et Textes, no. 24 (June). https://doi.org/10.4000/12xqn.