Open Source Projects at UB Mannheim
Optical Character Recognition (OCR)
- OCR-fileformat validates and transforms various OCR file formats (hOCR, ALTO, PAGE, FineReader): [code], [GUI], [MIT License].
- ocr-gt-tools is an ergonomic line-by-line transcription of scanned text: [code], [GNU Affero General Public License v3.0].
- Zotero OCR adds the functionality to perform an OCR for the PDFs selected in Zotero: [code], [GNU Affero General Public License v3.0].
- ocrd-pagetopdf is an OCR-D wrapper for prima-pagetopdf. It transforms all PAGE-XML+IMG to PDF with text layer and (optionally) polygon outlines: [code], [Apache License 2.0].
- TesseractXplore is an easy-to-use graphical interface to tesseract with full control: [code], [MIT License].
- PagePlus is a command-line tool for processing and analyzing PAGE XML files: [code], [MIT License].
- Churro is an OCR toolkit for historical document transcription (UB fork of Stanford’s Churro): [code], [Apache License 2.0].
- blatt is an NLP helper for OCR-ed pages in PAGE XML format: [code], [MIT License].
- keyboardBuilder4eScriptorium builds virtual keyboards for eScriptorium: [code], [MIT License].
- ocr-models is a registry of models for OCR engines (UB fork of kba/ocr-models): [code], [MIT License].
- Tesseract_Dokumentation provides German documentation relating to the text recognition software Tesseract: [code].
Archived: ocromore (tool for processing, enhancing and evaluating multiple OCR outputs, [code]), crass (tool to crop and splice segments of scanned pages, [code]), Mocrin (coordinates multiple OCR engines into a uniform workflow and folder structure, [code]).
Ground truth data for OCR
Transcriptions (mostly PAGE XML, often created with eScriptorium or Transkribus) for training and validating OCR recognition models.
- digi-gt ground truth for the digitized historic collections of UB Mannheim: [code], [CC0 1.0 Universal License].
- reichsanzeiger-gt ground truth for the digital edition of the Deutscher Reichsanzeiger und Preußischer Staatsanzeiger: [code], [CC0 1.0 Universal License].
- hkb-gt ground truth for the digitised newspaper Hakenkreuzbanner (Mannheim region, 1931–1945): [code], [CC0 1.0 Universal License].
- AustrianNewspapers the NewsEye/READ OCR training dataset of the Austrian National Library (19th-c. Austrian newspapers; original dataset published under CC BY 4.0): [code].
- digitue-gt ground truth for digitized books and journals of the University Library of Tübingen: [code], [CC0 1.0 Universal License].
- stabi-berlin-gt ground truth for digitized publications of Staatsbibliothek zu Berlin: [code], [CC0 1.0 Universal License].
- tudigi-gt ground truth for digitized publications of ULB TU Darmstadt: [code], [CC0 1.0 Universal License].
- MannheimerZeitungen ground truth for historic newspapers, generated with GTMake: [code], [CC0 1.0 Universal License].
- charlottenburger-amtsschrifttum ground truth from the collection Charlottenburger Amtsschrifttum (1879–1919, Fraktur): [code], [CC0 1.0 Universal License].
- NZZ-black-letter-ground-truth ground truth for 167 front pages of the Neue Zürcher Zeitung (1780–1947). Fork of the original dataset published by impresso: [code], [CC BY-NC 4.0 License].
- mkn-kurrent-gt ground truth for Kurrent handwritten periodicals from the Moravian Knowledge Network. Fork of bertsky/mkn-kurrent-gt: [code], [CC BY-SA 4.0 License].
- dach-gt ground truth and full text for selected prints of German archives and libraries: [code], [CC0 1.0 Universal License].
- gt-fraktur ground truth for Fraktur/Gothic prints of the 19th century. Fork of the original data by UB Tübingen (ubtue/gt-fraktur, CC0): [code].
- Weisthuemer transcriptions of Jacob Grimm’s Weisthümer (Middle High German), for training or validating OCR models: [code], [CC0 1.0 Universal License].
- Fibeln transcriptions of 19th-century primers (Fibeln), for training or validating OCR models: [code], [CC0 1.0 Universal License].
Related tools: Reichsanzeiger (software and data for the newspaper’s digital edition, [code]), ra-scripts (scripts used during the Reichsanzeiger project, [code]), reichsanzeiger-nlp (NER/NEL corpus for the Deutscher Reichsanzeiger, [code], [CC0 1.0 Universal License)) and VisualAnzeights (analysis of Reichsanzeiger advertisements, [code]).
OCR projects
Software and data from DFG and other digitization projects:
- BeTrial is a Bernoulli trial generator for OCR result validation, part of the Aktienführer-Datenarchiv DFG project: [code], [Apache License 2.0].
- GTCheck validates modifications of OCR ground truth in a git repository, showing original text, modified version and image: [code], [Apache License 2.0].
- DCC is the software and data for the digitalization, OCR and structuring of the books The Descendants of the Colonial Clergy: [code], [MIT License (code)].
Data management & bibliographic tools
- Zotkat is an extension of Zotero for cataloguing in a broad sense and contains also some experimental approaches: [Zotkat], [GNU Affero General Public License v3.0].
- malibu (Mannheim library utilities) is a collection of lightweight web-based tools to work with bibliographic metadata from various sources on the web, aimed at supporting the workflows of subject librarians and acquisitions librarians: [code], [GUI].
- uma_publist a TYPO3 extension to include publication lists from an EPrints repository in TYPO3 websites where the lists are generated and synced automatically: [code], [GNU General Public License v2.0].
- MArs is a web application for seat booking in Mannheim University Library: [code], [GNU Affero General Public License v3.0].
- ape (ALMA Print Extension) prints custom letters and notifications from Alma: [code], [GNU General Public License v3.0].
- ucompanies downloads PDF forms from ucompanies, extracts the texts (including check marks) and structures them into CSV and DTA files: [code].
- validate-mets validates METS files: [code], [Apache License 2.0].
Digital libraries
- Kitodo.Presentation is a feature-rich TYPO3 extension for building a METS- or IIIF-based digital library. UB Mannheim maintains a fork with development work (a no-Docker demo site, viewer theming and further features); the fork’s documentation, which also covers that non-upstream work, is published separately: [upstream], [code], [docs], [live demo], [GNU General Public License v3.0].
- omeka-matomo is a plugin that integrates Matomo Analytics into Omeka Classic installations: [code], [GNU General Public License v3.0].
AI applications
- UBi is an agentic AI-powered assistant (chatbot) for UB Mannheim: [code], [MIT License].
- FAIRplexica is an open source AI assistant for research data management (RDM), a fork of Vane (formerly Perplexica): [code], [MIT License].
- FAIR-GPT is a documentation for FAIR GPT, a virtual RDM consultant: [code], [CC0 1.0 Universal License].
- maidisco is the Mannheim Intelligent Discovery System, an experimental web application that adds AI assisted search to discovery systems like Primo or VuFind: [code], [GNU Affero General Public License v3.0].
Research data management
- data-journals-dashboard is a web dashboard for searching and filtering a community-curated list of data journals: [code].
- awesome-research-software is a curated list of production-ready open-source research software: [code], [CC0 1.0 Universal License].
- madabi (Mannheim Data Bibliography) is a registry of metadata of all data created or collected by the university: [code], [MIT License].
- madata is a tool for syncing the dataset metadata between MADATA and Wikidata: [code], [MIT License].
- theme-madataplan is an RDMO theme for madataplan: [code].
- awesome-RDM is a curated list of awesome RDM resources for researchers and organisations: [code], [CC BY 4.0 License].
Knowledge graphs & Natural Language Processing (NLP)
- bbw is an automatic semantic annotator for tabular data using a Wikibase instance and metasearch in SearX: [code], [GUI], [tutorial], [MIT License].
- spacyopentapioca is a spaCy wrapper of OpenTapioca for named entity linking on Wikidata: [code], [MIT License].
- kg-enricher is a library for enriching strings, entities and knowledge graphs using Wikibase knowledge: [code], [MIT License].
- MBI-KG is a knowledge graph of structured and linked economic research data extracted from Monatsschrift für Wirtschaft und Konjunktur: [code], [MIT License].
- Amtsgericht-KG analyzes and visualizes a knowledge graph of company registrations across German district courts: [code].
- check-fake-references is a script to check references for plausibility: [code], [MIT License].
- cas2iob converts UIMA CAS XMI files exported from INCEpTION into IOB TSV files, handling nested NER tags, NEL tags and components: [code], [MIT License].
Archived: RaiseWikibase (tool for fast data import and knowledge graph construction with Wikibase, [code], [docs]).
Screen sharing & team work
- PalMA enables people to share several contents on one monitor: [code], [GNU General Public License].
Third-party forks
Forks of third-party projects used by UB Mannheim:
- eScriptorium is an open source transcription platform developed as part of the Scripta, RESILIENCE and Biblissima+ projects. The UB Mannheim repository is a fork of scripta/eScriptorium with updates from UB Mannheim: [code], [German documentation], [MIT License].
- Kitodo — digitization workflow software: kitodo-presentation (upstream, [code], [docs], [GNU General Public License v3.0]), kitodo-production (upstream, [code], [docs], [GNU General Public License v3.0]), kitodo-workflow-editor (upstream, [code], [Camunda License]), plus the UB’s own docker configurations kitodo-presentation-docker ([code], [GNU General Public License v3.0]) and kitodo-production-docker ([code]), and slub_digitalcollections (Kitodo.Presentation-based templates for SLUB digital collections, [code]).
- Churro (UB fork of Stanford’s Churro, [code]) — see also the OCR section.
- DataCiteDoi (UB fork of eprintsug/DataCiteDoi, DOI registration via DataCite, [code], [GNU General Public License]).
- FAIR-farfalle and FAIR-sensei (Perplexity analogues for RDM based on farfalle and sensei, [FAIR-farfalle], [FAIR-sensei], both Apache License 2.0).
- ocrd_pagetopdf (UB fork of OCR-D/ocrd-pagetopdf, [code]) — see also the OCR section.
- theme-maobjects and theme-maobjects-nikephoros (customizable Omeka Classic themes for MAObjects, [theme-maobjects], [theme-maobjects-nikephoros]).
- rdmo-docs (English documentation for rdmo, [code]).
- handbuch-it-in-bibliotheken (working version of the Handbuch IT in Bibliotheken, [code]).
- eprints-create_sitemap (optional script for EPrints, [code], [GNU General Public License]).
- ocr-model-repo-template (UB fork of OCR-D/gt-repo-template, [code]) and ocr-model-metadata (UB fork of OCR-D/gt-metadata, [code]).