Skip to the content.

Open Source Projects at UB Mannheim

Optical Character Recognition (OCR)

  1. OCR-fileformat validates and transforms various OCR file formats (hOCR, ALTO, PAGE, FineReader): [code], [GUI], [MIT License].
  2. ocr-gt-tools is an ergonomic line-by-line transcription of scanned text: [code], [GNU Affero General Public License v3.0].
  3. Zotero OCR adds the functionality to perform an OCR for the PDFs selected in Zotero: [code], [GNU Affero General Public License v3.0].
  4. ocrd-pagetopdf is an OCR-D wrapper for prima-pagetopdf. It transforms all PAGE-XML+IMG to PDF with text layer and (optionally) polygon outlines: [code], [Apache License 2.0].
  5. TesseractXplore is an easy-to-use graphical interface to tesseract with full control: [code], [MIT License].
  6. PagePlus is a command-line tool for processing and analyzing PAGE XML files: [code], [MIT License].
  7. Churro is an OCR toolkit for historical document transcription (UB fork of Stanford’s Churro): [code], [Apache License 2.0].
  8. blatt is an NLP helper for OCR-ed pages in PAGE XML format: [code], [MIT License].
  9. keyboardBuilder4eScriptorium builds virtual keyboards for eScriptorium: [code], [MIT License].
  10. ocr-models is a registry of models for OCR engines (UB fork of kba/ocr-models): [code], [MIT License].
  11. Tesseract_Dokumentation provides German documentation relating to the text recognition software Tesseract: [code].

Archived: ocromore (tool for processing, enhancing and evaluating multiple OCR outputs, [code]), crass (tool to crop and splice segments of scanned pages, [code]), Mocrin (coordinates multiple OCR engines into a uniform workflow and folder structure, [code]).

Ground truth data for OCR

Transcriptions (mostly PAGE XML, often created with eScriptorium or Transkribus) for training and validating OCR recognition models.

  1. digi-gt ground truth for the digitized historic collections of UB Mannheim: [code], [CC0 1.0 Universal License].
  2. reichsanzeiger-gt ground truth for the digital edition of the Deutscher Reichsanzeiger und Preußischer Staatsanzeiger: [code], [CC0 1.0 Universal License].
  3. hkb-gt ground truth for the digitised newspaper Hakenkreuzbanner (Mannheim region, 1931–1945): [code], [CC0 1.0 Universal License].
  4. AustrianNewspapers the NewsEye/READ OCR training dataset of the Austrian National Library (19th-c. Austrian newspapers; original dataset published under CC BY 4.0): [code].
  5. digitue-gt ground truth for digitized books and journals of the University Library of Tübingen: [code], [CC0 1.0 Universal License].
  6. stabi-berlin-gt ground truth for digitized publications of Staatsbibliothek zu Berlin: [code], [CC0 1.0 Universal License].
  7. tudigi-gt ground truth for digitized publications of ULB TU Darmstadt: [code], [CC0 1.0 Universal License].
  8. MannheimerZeitungen ground truth for historic newspapers, generated with GTMake: [code], [CC0 1.0 Universal License].
  9. charlottenburger-amtsschrifttum ground truth from the collection Charlottenburger Amtsschrifttum (1879–1919, Fraktur): [code], [CC0 1.0 Universal License].
  10. NZZ-black-letter-ground-truth ground truth for 167 front pages of the Neue Zürcher Zeitung (1780–1947). Fork of the original dataset published by impresso: [code], [CC BY-NC 4.0 License].
  11. mkn-kurrent-gt ground truth for Kurrent handwritten periodicals from the Moravian Knowledge Network. Fork of bertsky/mkn-kurrent-gt: [code], [CC BY-SA 4.0 License].
  12. dach-gt ground truth and full text for selected prints of German archives and libraries: [code], [CC0 1.0 Universal License].
  13. gt-fraktur ground truth for Fraktur/Gothic prints of the 19th century. Fork of the original data by UB Tübingen (ubtue/gt-fraktur, CC0): [code].
  14. Weisthuemer transcriptions of Jacob Grimm’s Weisthümer (Middle High German), for training or validating OCR models: [code], [CC0 1.0 Universal License].
  15. Fibeln transcriptions of 19th-century primers (Fibeln), for training or validating OCR models: [code], [CC0 1.0 Universal License].

Related tools: Reichsanzeiger (software and data for the newspaper’s digital edition, [code]), ra-scripts (scripts used during the Reichsanzeiger project, [code]), reichsanzeiger-nlp (NER/NEL corpus for the Deutscher Reichsanzeiger, [code], [CC0 1.0 Universal License)) and VisualAnzeights (analysis of Reichsanzeiger advertisements, [code]).

OCR projects

Software and data from DFG and other digitization projects:

  1. BeTrial is a Bernoulli trial generator for OCR result validation, part of the Aktienführer-Datenarchiv DFG project: [code], [Apache License 2.0].
  2. GTCheck validates modifications of OCR ground truth in a git repository, showing original text, modified version and image: [code], [Apache License 2.0].
  3. DCC is the software and data for the digitalization, OCR and structuring of the books The Descendants of the Colonial Clergy: [code], [MIT License (code)].

Data management & bibliographic tools

  1. Zotkat is an extension of Zotero for cataloguing in a broad sense and contains also some experimental approaches: [Zotkat], [GNU Affero General Public License v3.0].
  2. malibu (Mannheim library utilities) is a collection of lightweight web-based tools to work with bibliographic metadata from various sources on the web, aimed at supporting the workflows of subject librarians and acquisitions librarians: [code], [GUI].
  3. uma_publist a TYPO3 extension to include publication lists from an EPrints repository in TYPO3 websites where the lists are generated and synced automatically: [code], [GNU General Public License v2.0].
  4. MArs is a web application for seat booking in Mannheim University Library: [code], [GNU Affero General Public License v3.0].
  5. ape (ALMA Print Extension) prints custom letters and notifications from Alma: [code], [GNU General Public License v3.0].
  6. ucompanies downloads PDF forms from ucompanies, extracts the texts (including check marks) and structures them into CSV and DTA files: [code].
  7. validate-mets validates METS files: [code], [Apache License 2.0].

Digital libraries

  1. Kitodo.Presentation is a feature-rich TYPO3 extension for building a METS- or IIIF-based digital library. UB Mannheim maintains a fork with development work (a no-Docker demo site, viewer theming and further features); the fork’s documentation, which also covers that non-upstream work, is published separately: [upstream], [code], [docs], [live demo], [GNU General Public License v3.0].
  2. omeka-matomo is a plugin that integrates Matomo Analytics into Omeka Classic installations: [code], [GNU General Public License v3.0].

AI applications

  1. UBi is an agentic AI-powered assistant (chatbot) for UB Mannheim: [code], [MIT License].
  2. FAIRplexica is an open source AI assistant for research data management (RDM), a fork of Vane (formerly Perplexica): [code], [MIT License].
  3. FAIR-GPT is a documentation for FAIR GPT, a virtual RDM consultant: [code], [CC0 1.0 Universal License].
  4. maidisco is the Mannheim Intelligent Discovery System, an experimental web application that adds AI assisted search to discovery systems like Primo or VuFind: [code], [GNU Affero General Public License v3.0].

Research data management

  1. data-journals-dashboard is a web dashboard for searching and filtering a community-curated list of data journals: [code].
  2. awesome-research-software is a curated list of production-ready open-source research software: [code], [CC0 1.0 Universal License].
  3. madabi (Mannheim Data Bibliography) is a registry of metadata of all data created or collected by the university: [code], [MIT License].
  4. madata is a tool for syncing the dataset metadata between MADATA and Wikidata: [code], [MIT License].
  5. theme-madataplan is an RDMO theme for madataplan: [code].
  6. awesome-RDM is a curated list of awesome RDM resources for researchers and organisations: [code], [CC BY 4.0 License].

Knowledge graphs & Natural Language Processing (NLP)

  1. bbw is an automatic semantic annotator for tabular data using a Wikibase instance and metasearch in SearX: [code], [GUI], [tutorial], [MIT License].
  2. spacyopentapioca is a spaCy wrapper of OpenTapioca for named entity linking on Wikidata: [code], [MIT License].
  3. kg-enricher is a library for enriching strings, entities and knowledge graphs using Wikibase knowledge: [code], [MIT License].
  4. MBI-KG is a knowledge graph of structured and linked economic research data extracted from Monatsschrift für Wirtschaft und Konjunktur: [code], [MIT License].
  5. Amtsgericht-KG analyzes and visualizes a knowledge graph of company registrations across German district courts: [code].
  6. check-fake-references is a script to check references for plausibility: [code], [MIT License].
  7. cas2iob converts UIMA CAS XMI files exported from INCEpTION into IOB TSV files, handling nested NER tags, NEL tags and components: [code], [MIT License].

Archived: RaiseWikibase (tool for fast data import and knowledge graph construction with Wikibase, [code], [docs]).

Screen sharing & team work

  1. PalMA enables people to share several contents on one monitor: [code], [GNU General Public License].

Third-party forks

Forks of third-party projects used by UB Mannheim:

  1. eScriptorium is an open source transcription platform developed as part of the Scripta, RESILIENCE and Biblissima+ projects. The UB Mannheim repository is a fork of scripta/eScriptorium with updates from UB Mannheim: [code], [German documentation], [MIT License].
  2. Kitodo — digitization workflow software: kitodo-presentation (upstream, [code], [docs], [GNU General Public License v3.0]), kitodo-production (upstream, [code], [docs], [GNU General Public License v3.0]), kitodo-workflow-editor (upstream, [code], [Camunda License]), plus the UB’s own docker configurations kitodo-presentation-docker ([code], [GNU General Public License v3.0]) and kitodo-production-docker ([code]), and slub_digitalcollections (Kitodo.Presentation-based templates for SLUB digital collections, [code]).
  3. Churro (UB fork of Stanford’s Churro, [code]) — see also the OCR section.
  4. DataCiteDoi (UB fork of eprintsug/DataCiteDoi, DOI registration via DataCite, [code], [GNU General Public License]).
  5. FAIR-farfalle and FAIR-sensei (Perplexity analogues for RDM based on farfalle and sensei, [FAIR-farfalle], [FAIR-sensei], both Apache License 2.0).
  6. ocrd_pagetopdf (UB fork of OCR-D/ocrd-pagetopdf, [code]) — see also the OCR section.
  7. theme-maobjects and theme-maobjects-nikephoros (customizable Omeka Classic themes for MAObjects, [theme-maobjects], [theme-maobjects-nikephoros]).
  8. rdmo-docs (English documentation for rdmo, [code]).
  9. handbuch-it-in-bibliotheken (working version of the Handbuch IT in Bibliotheken, [code]).
  10. eprints-create_sitemap (optional script for EPrints, [code], [GNU General Public License]).
  11. ocr-model-repo-template (UB fork of OCR-D/gt-repo-template, [code]) and ocr-model-metadata (UB fork of OCR-D/gt-metadata, [code]).