Notes on the Analysis

The exemplary analysis serves not only to further research the literary reception of Sappho in the German-speaking world, but also to test a hitherto untried methodological approach to literary reception histories within the Digital Humanities.

Data Basis

The reception testimonies underlying the lists – and thus also the analysis – were compiled over several years of research, drawing on the German, Austrian and Swiss national bibliographies, newspaper and periodical databases such as ANNO and the Deutsches Zeitungsportal, digital text corpora such as TextGrid and Project Gutenberg, fanfiction platforms such as the Archive of Our Own, and scholarly secondary literature; the material was supplemented with tips from personal and professional contacts.

The texts found were first recorded in a spreadsheet with basic bibliographic information – date of composition and publication, title, author(s), genre, place of publication, publisher, and, where available, a link to a digitized copy. Using OpenRefine, the spreadsheet was then semi-automatically linked to Wikidata to identify persons, works, places and publishers via unique identifiers (QIDs). A purpose-built Python script finally converted the spreadsheet into the XML/TEI files openly available on GitHub, which form the basis for the analysis corpus presented here.

Methodological Approach

All analysed reception testimonies were drawn from the same corpus: the volume Sappho. Texte zur literarischen Rezeption im deutschsprachigen Raum (2023), which presents the reception testimonies – apart from the texts by Huber, Schrott and Schwarz – in the form of their first printing.

Also, all Sappho fragments contained in Andreas Bagordo’s bilingual edition, based on the text-critical edition by Eva-Maria Voigt (Patmos 2009), were analysed (cf. the menu »Texts« > »Sappho Fragments«), as well as 99 sample reception testimonies illustrating typical patterns of reception. These include, for example, the recurring treatment of the Sappho–Phaon legend and the portrayal of Sappho as a great poet or as a homoerotic desiring subject. Texts of various genres (poetry, prose, drama, and other texts, e.g. comics) and of varying degrees of renown were considered – not only works by canonized authors, but also those by lesser-known writers, from the 16th to the 21st century.

Motifs, topics, plots, rhetorical topoi, characters, as well as references to persons, places and works, and quotations were taken into account. On the basis of these phenomena, intertextual relationships were established between the reception testimonies and the Sappho fragments, as well as among the reception testimonies themselves.

Before the phenomena could be integrated into this data model, they had to be extracted from the texts. The 99 reception testimonies were already available in full text through the work on the Sappho anthology; the 197 Sappho fragments were made machine-readable using Optical Character Recognition (OCR). The actual interpretive work was carried out in several successive close-reading passes: the texts were read multiple times, with each pass adding further annotations or checking and refining those already made.

Annotation used an individual XML tagset: motiv, thema, stoff and topos marked motifs, topics, biographical plot variants and rhetorical topoi, werk marked references to works (supplemented, for quotations and paraphrases, with the attribute art="passage"), person with the attributes art="real", art="fictional" and art="type" marked real people, fictional characters and character types, and ort with art="real" or art="fictional" marked real and fictional places. Where a phenomenon could not be pinned directly to the wording, it was given a normalized term via the attribute type; where it could not be located in the running text at all – as was often the case for some plot variants and topics – tags were placed outside the text as well.

Two purpose-built Python scripts supported ongoing quality control of the annotation: one listed all tags already assigned along with their values, the other showed which of these had already been used more than once – groundwork that also formed the basis for the SKOS vocabulary (see below). For the final round of corrections, a static HTML view was developed that displayed, for each text, the full text synoptically alongside a checklist of all annotated phenomena.

Based on the annotated XML documents, six successive Python scripts convert the data into RDF triples. A first script creates person instances from the author information in the list of all reception testimonies and, where a Wikidata ID is available, adds basic biographical data (birth, death, gender, image) as well as further identifiers (DBpedia, GND, VIAF). A second script processes the bibliographic information on the reception testimonies and generates work data and, where a publication date is available, manifestation data with titles, places and dates of publication, as well as, where available, cover images and further identifiers (including a Goodreads work ID); a third script converts the 197 Sappho fragments and Sappho’s complete works in the same way. A fourth script reads the annotated XML documents and enters motifs, topics, plots, topoi, references to works, places and persons, and characters into the corresponding classes of the data model, matched against the SKOS vocabulary (see below). A fifth script uses these to construct directed intertextual relations – both between Sappho’s fragments and the reception testimonies and among the reception testimonies themselves – documenting the respective textual evidence along the way. A sixth script finally merges all previously generated files into a single graph (following a “last write wins” rule); the merged graph is then processed with the Java-based reasoner HermiT. This produced over 400,000 triples, extendable to over 600,000 through reasoning.

To represent the data, a dedicated ontology was developed, building on the existing ontologies CIDOC CRM, LRMoo and INTRO, and additionally equipped with numerous alignments to further ontologies.

Recurring reception phenomena are also defined in a SKOS vocabulary. Since productive literary reception phenomena cannot be fully determined in advance, the vocabulary was developed inductively from the material itself: all phenomena occurring in at least two annotated documents – references to works and plot variants regardless of frequency – were manually structured and hierarchized in the web-based LOD editor VocBench. Internally, they were linked via hierarchical (skos:broader/ skos:narrower) and associative (skos:related) relations, externally, wherever possible, to Wikidata (skos:relatedMatch, skos:closeMatch, skos:narrowMatch, skos:broadMatch). In total, the vocabulary comprises 427 concepts and just over 5,000 triples; motifs and persons make up the largest share.

The resulting data is openly accessible in a GitHub repository and can also be queried via a SPARQL interface on this website. Sample statistical analyses can be found here; all data can also be explored via a network visualization.

Analysed Reception Testimonies

Statistical Overview of the Reception Testimonies