Mass Digitization, Digital Libraries, and Data Retirement
In this week’s readings, we explored the history and challenges of mass digitization, digital libraries, and data circulation. Brewster Kahle argues for universal access to knowledge through the Internet Archive, the African American Periodical Poetry dataset demonstrates the labor and choices involved in making digitized collections usable, and the Lenna Image article asks when should data die.
In groups, you will build on these readings to explore digital libraries and archives. The goal is to think critically about how cultural objects become digital representations, how those representations become datasets, and when preservation or access might conflict with consent, rights, or harm.
If you have questions, please reach out to the instructors on Slack.
Exploring HathiTrust: Digitization in Practice
As a group, work through the following tasks together. Document your findings and observations as you go:
Finding the African American Periodical Poetry: Using the Hennessey dataset from your readings (which is available in the readings here), select 2-3 poems from different magazines (e.g., The Crisis, Opportunity, Black Opals). Try to locate the original magazine issues in HathiTrust that contain these poems.
- What search strategies did you use?
- Were you able to find all the issues you looked for?
- What barriers or challenges did you encounter?
Examining OCR Quality: Once you’ve found at least one magazine issue, examine the OCR (Optical Character Recognition) quality:
- Can you search for specific words or phrases within the document?
- How accurate is the OCR text compared to what you see in the page images?
- What kinds of errors do you notice? (Consider fonts, layouts, damaged pages, etc.)
- How might OCR quality affect research using this material?
Understanding Context: Compare the poem in the dataset to the original magazine page:
- What additional context do you gain from seeing the original page?
- What else appears on the page or in the issue alongside the poem?
- What information was lost in creating just the dataset of poems?
- What information was gained by creating the structured dataset?
Access and Rights: Look at the viewing options and restrictions:
- Can you download the full PDF? Individual page images?
- Are there any access restrictions? (Full view vs. limited preview)
- What copyright or usage information is provided?
- How does HathiTrust balance preservation, access, and copyright?
- How does this compare to Kahle’s vision of universal access?
Metadata and Organization: Examine how HathiTrust has cataloged and organized the material versus the dataset:
- What metadata is provided for the item you’re viewing?
- How does this compare to the MARC records we discussed in class?
- What other data could have the authors collected?
- Would you organize the dataset differently?
Discovering & Digitizing Cultural Objects
Once you have completed the HathiTrust exploration, you should begin working on this part of the assignment, which is designed to help you start discovering relevant materials and discussing potential focuses for your semester-long project. This section is focused on your group’s selected cultural area. If you have yet to finalize this, please do so before you begin this part of the assignment. Details on that are available in the semester long project assignment.
You are welcome to use AI tools to help you with your assignment, but you should include links and screenshots to all materials you find, as well as an ai-chat-log.md file in your group repository. You should also be prepared to discuss how you found these materials and what tools you used to find them.
In our readings this week, we learned about how cultural objects and practices are turned into digital representations, and how these representations are then shared and preserved online. Building on your HathiTrust exploration, you will now investigate what this process looks like for your selected area of focus.
Here are the following prompts you should answer for your area of focus:
Digital Objects & Representations: What might be considered a digital object or digital representation for your area of focus? You can have multiple examples, but you should explain why you consider it relating to your area of focus (this can be short though).
Digitization Processes: How are digital objects and representations created for your area of focus? What are the processes involved in digitizing these objects? In the African American Periodical Poetry dataset and your HathiTrust exploration, you learned about OCR (Optical Character Recognition) for extracting text from digitized print materials. What are some of the digitization processes for your area of focus? How do they compare to what you observed in HathiTrust?
Historical Equivalents: Kahle repeatedly made reference to the Library of Alexandria in his article. What are some historical equivalents to digital objects in your area of focus? How do these historical objects compare to their digital counterparts?
Born-Digital Materials: Are there examples of born-digital materials for your area of focus? How do these materials compare to digitized objects? For those unfamiliar, born-digital materials are those that were created digitally and never existed in analog or physical form (think most social media, for example).
Oldest Digital Library/Archive: What is the oldest digital library or archive you can find that relates to your area of focus? How has this resource been maintained and updated over time? What metadata or standards exist for the object? You may also include examples that are no longer maintained or have been abandoned.
Newest Digital Library/Archive: Conversely, what is the newest digital library or archive you can find that relates to your area of focus? How does this resource compare to older digital libraries or archives, especially around metadata?
Viral Examples: Are there any examples of your digital object that have gone viral? How did this happen and what impact did it have on the object or the digital library/archive that hosted it?
Free vs. Proprietary Access: Can you find any examples of free vs. proprietary digital libraries or archives for your area of focus? How do these resources differ in terms of access? Think back to Kahle’s discussion of public vs. commercial digital libraries.
Retirement, Refusal, or Restricted Access: Are there examples in your area of focus where cultural materials should not circulate freely forever? This might involve privacy, consent, community protocols, copyright, harassment, platform harm, or changing social norms. What would responsible documentation need to say?
You should aim to find at least one example for each prompt, but you are welcome to find more.
Documenting & Synthesizing Your Findings
The final part of this assignment is to document your findings and prepare a brief synthesis for class discussion. You should create a new folder and Markdown file in your group’s GitHub repository that contains your answers to the prompts above. You should also include any relevant links, screenshots, or other documentation of your findings.
We will discuss selected findings in class. You do not need to prepare formal slides. Your group should be ready to briefly share one useful finding, one problem or tension, and one question. Each member should also be prepared to discuss their contribution to the group’s work.
You are welcome to divide labor any way you choose, BUT please do your best to be equitable and be sure to document who is responsible for what. You may have some group members do part 1 or part 2, but ideally everyone should contribute to either investigating the data in HathiTrust or identifying digital objects. I would highly encourage you to consider the git history (your git log) as a way of making your labor visible in these types of assignments.