Digital Humanities, University of Bern Digital Humanities, University of Bern DH Lab
  • Lab
  • Programme
  • Assignment
  • Toolbox
  • Student Workflows
  • About
  • Intro to DH ↗
  1. Programme
  2. Session 11
  • Lab
  • Programme
    • Session 1
    • Session 2
    • Session 3
    • Session 4
    • Session 5
    • Session 6
    • Session 7
    • Session 8
    • Session 9
    • Session 10
    • Session 11
    • Session 12
    • Session 13
    • Session 14
  • Assignment
  • Toolbox
  • Student Workflows
    • How to add your workflow here
  • About

  • Introduction to DH (companion course)

On this page

  • In this session
  • OCR and HTR
  • Segmentation
  • Transcription and ground truth
  • Evaluation
  • Export
  • Edit this page
  • Report an issue
  1. Programme
  2. Session 11

Session 11

Text recognition: introduction to eScriptorium

Author
Affiliation

Tobias Hodel

Walter Benjamin Kolleg / Digital Humanities, University of Bern

Published

24 November 2026

Modified

2 September 2026

Hodel Taught by Tobias Hodel.

In this session

eScriptorium, hands-on: from a page image to machine-readable text.

OCR and HTR

  • Optical Character Recognition and Handwritten Text Recognition: the same pipeline, different difficulty
  • Why printed early modern text is often harder than neat nineteenth-century handwriting
  • The pipeline: image → binarisation → segmentation → transcription → export
  • Kraken as the engine underneath eScriptorium

Segmentation

  • Regions and lines, and the baseline model
  • Reading order, which is a scholarly decision as much as a technical one
  • Correcting segmentation by hand — and why an hour here saves a day later
  • Layout that breaks models: marginalia, tables, multi-column pages, insertions

Transcription and ground truth

  • Transcribing lines in the editor
  • Transcription guidelines: abbreviations, u/v, i/j, long s, capitalisation. Decide before you start, write it down, and stay consistent — this is an editorial decision, exactly as in the TEI session
  • How much ground truth a model needs
  • Training and fine-tuning an existing model rather than starting from scratch

Evaluation

  • CER (character error rate) and WER (word error rate): what they measure
  • Why a 5% CER can be excellent or useless depending on what you want to do with the text
  • Where the errors concentrate — and they always concentrate somewhere
  • Ground truth is not truth: it is your transcription decisions, made consistent
NoteThis is the Intro’s machine-learning session, from the inside

The Intro asks where bias enters a machine learning pipeline. Here you occupy the position that answer names: you are the person producing the labels.

Export

  • ALTO and PAGE XML, and what they contain besides the text
  • Getting from ALTO to TEI
  • Publishing images and text together with IIIF
TipIn parallel in the Intro

From algorithms to machine learning

Back to top
Session 10
Session 12
  • Edit this page
  • Report an issue