Cost-Effective Machine Learning for Automatically Processing Bibliographic
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Many digital humanities projects involve tedious and repetitive tasks take time away from the higher-level tasks further down the pipeline that require intelligent decision-making. When funding is available, the tedious and repetitive tasks are often assigned to research assistants, but when funding is scarce, those tasks tend to create bottlenecks that either impede progress or halt it altogether. This paper argues that artificial intelligence and machine learning tools and techniques are worth exploring as cost-effective, accessible solutions to these problems. The Digital Latin Library project provides a case study through its experiments with fine-tuning pretrained transformer language models to process noisy bibliographic metadata. The results show that the models have potential for accelerating this tedious task. But the experiments also had an unexpected, yet positive outcome: the models revealed gaps in the catalog's coverage, helping to focus the efforts of the human experts working on the project.