In the long nineteenth century, German-language newspapers in the United States were flourishing and migrants were well provided with reading material, surpassing what was available even their homeland. Migrants “read more in America” than they had at home “because there is more going on that they need to know” (Park 9). Already in 1840, New York City, with its four German daily newspapers, outpaced Berlin, the capital of Prussia, and Leipzig, a renowned publishing hub in Europe (Kamphoefner, Germans in American 108). By 1872, the New Yorker Staats-Zeitung (1834–1974), boasting a circulation of 55,000, proudly held the distinction of being the world’s largest German-language newspaper, rivaling in circulation the city’s major English-language papers such as the New York Times (1851–present) and the New York Tribune (1841–1966) (Bergquist, “Ottendorfer” 841). In 1872, the German-speaking population in the state of New York “was estimated at between 600,000 and 700,000” producing 56 German-language newspapers in total (Arndt and Olson 313). Scholars have used newspapers such as the New Yorker Staats-Zeitung, and its distribution and circulation figures, to give a clear idea of the influential role that the German-language press played both within migrant communities and in the wider American context (399).[1] They have portrayed the German-language press as a male-dominated enterprise targeting male audiences, perpetuating a narrative that overlooks the contributions and experiences of women (Blaschke, Die Entdeckung 22). As illustrated in the previous chapter, few have acknowledged that newspapers were widely read, consumed, (co-)edited, and (co-)run by women (322). When it turned into the world’s largest German-language newspaper, the New Yorker Staats-Zeitung, was managed by a woman. In 1837, at the age of twenty-two, Anna Behr migrated as a single woman to the U.S., where she met her husband Jacob Uhl, a printer, one year later. In 1845, they bought the weekly New Yorker Staats-Zeitung. When her husband died, she “refused to sell” the newspaper, which had been turned into a daily by then, and “took over control of its publication” while caring for six children (Bergquist, “Ottendorfer” 841). According to the data in the comprehensive bibliography of German-American periodicals compiled by Arndt and Olson, “Mrs. Jacob Uhl” published the paper from 1852–1859 and was listed as editor-in chief between 1852–1858 as “Mrs. Anna Uhl (Mrs. Jacon Uhl)” (399). When Uhl, however, re-married Oswald Ottendörfer, an editor of the Staats, in 1859, the ownership was transferred to her new husband. In this period, married women could not own property or sign contracts.
Anna Uhl Ottendörfer continued to serve as business manager and “[v]irtually every business day until shortly before her death,” one could find “her in the newspaper’s offices” (Bergquist, “Ottendorfer” 841). When she was the official owner, she moved the paper and its staff to a larger building, increased advertising and circulation numbers. In one of the earliest comprehensive accounts of migrant newspapers in the U.S. published in 1922, Robert Park immediately introduces the newspaper in his section of the German dailies, however, he simply states that it was founded by Jakob Uhl in 1873 (which is incorrect) and owned by Oswald “Ottendorfer” after 1859 (266). He does not even mention Anna Uhl Ottendörfer. By the time she transferred ownership, the paper had already established itself as a nationally circulating publication and a political force in New York City. Looking at this period, the story of Anna Uhl Ottendörfer seems to be an exception: migrant, mother, and successful businesswoman. She could choose to labor at a time when others remained at home. But she was not the only one who owned and edited a German-American paper even though the historiography of the nineteenth-century German-American press reflects women’s underrepresentation as both producers and consumers of newspapers. According to Hoffmann this represents larger issues within migration history. As he points out, “[d]iscussions of immigration seldom distinguish between males and females,” predominantly neglecting the differences between the Americanization processes of men and women, “concentrating on ethnic experiences of male immigrants,” even though a high percentage of women migrated to the U.S. as well (383). In 2009, Helbich called attention to misleading and insufficient historical representations that only show men as actors of migration and women, if they appear at all, simply as followers of their husbands bound by family obligations. He concludes: “there is still plenty of work to be done,” as “only the surface has been scratched” in the context of gendered perspectives (399). The patriarchal system and marriage laws of the time prevented Anna Uhl Ottendörfer from officially continuing to run the New Yorker Staats-Zeitung: a fact that is also reflected in the newspaper’s metadata.
As the example of Anna Uhl Ottendörfer illustrates, metadata (from both analogue and digital sources) should not be taken at face value, particularly in the context of gender-focused research. Metadata can be misleading and, in some cases, entirely inaccurate, underscoring the need for critical examination. The indexing of documents in archives is oriented toward male protagonists, which reproduces the invisibility of women in archives. Searching for traces of women using analogue finding aids in archives and libraries is rarely successful. Before delving into how media portrayals played a significant role in gendering the German-American press in the nineteenth century in the following chapters, the objective of this chapter is to reflect on the digitization of German migrant newspapers and the corpus used in this study for conducting such an analysis.
The advent of digitization of historical newspapers has “revolutionized” the way we can access, read, and search the news from the past (see DiCenzo). Besides discussions about revolution and scale, scholars are worried about “the risks of the new biases introduced by digitisation, such as a lack of standards for the process, issues surrounding the representation of the titles being digitised, diverging quality in metadata,” and OCR, as historian Estelle Bunout points out (279). “Historians share concerns about the bias implicit in the archives scholars use, both as bias represents a subject of discourse analysis and as it represents a challenge, insofar as the silences of archives limit the possibility of a complete factual account,” as Jo Guldi in her review article on text mining for historical analyses outlines (522). In The Nineteenth-Century Press in the Digital Age, Mussell urges researchers working with nineteenth-century digital archival materials to consider the transformation of an analogue source into one that can be displayed on a computer screen to be able to understand further processing and contextualize suitable data (192–202). In line with Bunout’s and Mussell’s statements, this chapter seeks to address these concerns and expose the implicit and explicit choices that have influenced the construction of the newspaper corpora used for this study. Subsequently, it will turn to reflections on selection criteria and the representativeness of the dataset compiled for this project, which Beelen et al. describe as an “environmental scan” to evaluate how it is suited for a study on text reuse and gender of the German-American press (3). Working with digitized newspapers also presupposes examining the history of their digitization from selection criteria to OCR recognition and online representation. Digitization projects have to adhere to certain guidelines that have an impact on what one can find and access online (and in Open Access). “[S]uch historicizing does not foreclose the possibility of new OCR editions […], but [instead] it recognizes the realities of our editions at a given moment—and the limits those realities place on re-search” (Cordell, “Q i-jtb” 216). As Melanie Althage briefly puts it, when using datasets that were not created by oneself, it is recommended to critically reflect on and question who created the data, why, and in addition how complete and representative it is (251).
First, as a whole, the corpus does not provide an accurate depiction of the newspapers that were published at that time. Due to the politics of digitization, which I will outline in this chapter, the corpus does not include any German-language newspapers from New York, such as the New Yorker Staats-Zeitung. Second, the metadata compiled by Chronicling America is limited and misleading. Therefore, additional “analogue” metadata taken from Arndt and Olson’s study and the information extracted from the newspaper content needed to be added to provide comprehensive information about the newspaper titles, the people who created them, and the ones who consumed them. Metadata help in a variety of ways. They help search and browse and document technical characteristics. Metadata, as any kind of data, is always already “cooked” or in the words of Johanna Drucker “captured” within a set of epistemological assumptions (6). As shown in the case of the New Yorker Staats-Zeitung, even this resource might not be enough to understand who was behind the production of the news. Metadata compiled by Chronicling America represent more than just information about these titles: they offer insights into the nineteenth-century newspaper landscape and who received recognition. Only two of the newspapers used in this study were, at some point, owned (and edited) by women. This is, however, not necessarily transparent to the user when only using the information from the digital archive. Third, apart from these representation issues, there are further obstacles regarding basic search algorithms that complicate this research. Chronicling America’s interface provides full-text search which allows users to browse through thousands of pages of German-American newspapers. This search method is useful for conducting layout analyses by being able to zoom in and out of the pages, keeping in mind OCR errors that have an impact on search results. However, it is not very helpful for text reuse analyses, in other words to find widely reprinted texts because it only allows for keyword queries and not for n-gram queries that could compare n-grams across documents.[2] In the following pages, I will sketch a critical bibliography of Chronicling America’s German-language press. However, since I am dealing with several newspapers from various states, this chapter cannot describe the histories of every individual newspaper and its digital transformation. Therefore, I will use specific papers, dates, locations, subscription numbers, and other information to address how digitization guidelines have influenced selection and digital preservation. This approach sheds light on the reasons and methods for finding these newspapers online and highlights the usefulness of such a corpus for studying nineteenth-century German-American newspaper reprinting culture while recognizing its limitations.
2.1. The “Environmental Scan” of Chronicling America’s German-Language Newspapers¶
Understandings of materiality and availability seem to dissolve in the digital environment because one cannot immediately trace the spaces between documents, between records and archives, or between front and back-end (Kaminsiki 231). Using a physical or digital archive for research poses different sets of questions regarding access and methodologies. In “The Transnational and the Text-Searchable: Digitized Sources and the Shadows They Cast,” Lara Putnam highlights the impact that the combined shifts toward transnational and digital approaches have had on historical research. For the scholar, free access to millions of newspaper pages from diverse locations around the world and reducing travel expenses of time and cost, are one of the examples that illustrate the promises of using digitized archives for historical research in the twenty-first century. Despite all these opportunities to find novel evidence, Putnam calls attention to digital practices arguing that the use of digital archives has received little critical reflection among mainstream historians (394). She writes:
You may think none of this concerns you. The phrase “digital turn” evokes specialized techniques like text mining and distant reading. Tools for counting, graphing, and mapping we recognize as “digital methods.” But the mass of historians’ research is about finding, and finding out. That so many of us are now finding and finding out via digital search has significant consequences, regardless of whether we count, graph, or map anything at all. (378)
In the words of Guldi, Putnam “advertised that the era of text mining might mean an important moment for transnational history” (537). However, she argues that researchers have been increasingly using digital archives in the last decades without reflecting on the fact that the material available online may not accurately represent the newspaper landscape of that time. The convenience of accessing freely available content can lead to the neglect of other newspapers that are not digitized or are behind paywalls. “Today it is hard to imagine conducting historical scholarship without technologies such as search engines, yet these technologies significantly impact historiography” (Romein et al. 310). Ian Milligan’s work has been groundbreaking in showing how the use of digitized material has changed research on the history of the press leading to overrepresentations of specific newspaper titles. In his study, Milligan examines newspapers cited in Canadian dissertations. He quantitatively reveals that overall citations of newspapers have increased in “the post-database period” and demonstrates that citations increasingly favor papers that have been digitized over those that have not (550). As he writes:
Before digitization, a newspaper like the Ottawa Citizen was roughly equivalent in historical usage to the Toronto Star, as one might expect, given their relative prominence in Canadian history. After the Star was digitized and made available, however, it became far more prominent in dissertations. (567)
Consequently, Putnam, Milligan, and others have advocated for understanding the scope of digitized source material and recommend complementing digital searches with physical archives whenever possible. For the German press landscape, Astrid Blome has called attention to the distorting image that digitized material imposes. In Germany, leading media outlets, long-running and widely circulated newspapers, and historically significant publications have primarily been (or are being) digitized (13). This creates a distorted representation of the structural diversity and media reality of the German press landscape. Leading media, quality newspapers, and representatives of the mass press are significantly overrepresented in digital collections, while the many small, regional, and local newspapers (with more limited circulation but central to the everyday lives of their readership) receive little attention (14). Consequently, the typological spectrum of the press is far from adequately represented. For preservation reasons, as Blome points out, sources that are either in high demand or particularly endangered are often prioritized for digitization: an approach that does not necessarily follow a coherent research logic (15). As these examples illustrate, national digitization programs follow different guidelines and priorities. Researchers working with a specific national archive must therefore familiarize themselves with its particular archival strategies and practices.
In the pre-digitization era, scholars of the German-American press have disproportionally focused on newspapers from metropolitan areas neglecting rural areas, the Trans-Mississippi region, and papers owned and edited by women. Do the digitized German-language newspapers in Chronicling America provide a different avenue to study the German-American press and its network of reprinting? The answer is yes. However, a careful examination of implicit and explicit information extracted from different source material is necessary to get a fuller view of the press, its digitization history, and to evaluate the usefulness of the corpus for approaching a transregional study of the textual ecosystem. As Blome points out, Chronicling America at the Library of Congress, which provides access to more than 20 million digitized newspaper pages, is one of the most research-friendly newspaper databases (41). For each newspaper, “essays” offer additional information about the publication’s history and editorial orientation. In addition, the platform features a variety of thematic entry points, tailored suggestions for further research, and interactive maps and graphics that offer guidance beyond simple full-text searches.
Starting in the 1990s, private, public, and private-public institutions have been digitizing their collections (Terras 3–4). The decision to digitize newspaper collections was largely influenced by the abundance of newspaper issues available in libraries and the absence of copyright restrictions for material older than 100 years, as highlighted by Sarah Oberbichler and Eva Pfanzelter (126). In 2003, building on former newspaper programs, the National Digital Newspaper Program (NDNP), a partnership between the National Endowment for the Humanities (NEH) and the Library of Congress (LC), was established as “a long-term effort to develop an Internet-based, searchable database of U.S. newspapers with descriptive information and select digitization of historic pages” (Beals and Bell 11).[3] The digitized pages were made available to the public through the Chronicling America website.[4] As of 2025, the digital newspaper portal Chronicling America provides free access to more than 20 million pages of historic American newspapers, spanning from 1770 to 1963.[5] The site includes more than 4,000 newspaper titles from across the United States and the number of newspaper titles continues to increase. While in 2018, when I began this project, 58 newspapers were available in the German language, in 2025, the site already offered access to 76. In 2026, Chronicling America lists 108 German-language newspapers. This is just a tiny fraction of all German-language periodicals published in the United States since the eighteenth century.
Providing full-text search over millions of OCRed pages via online portals, digitization has greatly simplified how users can access, visualize, and search historical newspapers. Digitizing historical newspapers can be a time-consuming and costly process, but it provides a valuable means of preserving and sharing historical information. It involves the process of converting physical newspapers into digital format, which can be made accessible through online databases and archives. The process typically involves the following steps: the first is to identify which newspapers will be digitized. This may involve prioritizing newspapers of high historical value, those with fragile or deteriorating physical copies, or those that are frequently requested by researchers. Once the newspapers have been selected, the physical copies are scanned using specialized equipment that captures high-quality images of each page. To make the text searchable and accessible, OCR technology, “the most widespread ML technology in libraries,” is applied to the scanned images to recognize and convert the text into digital format (Cordell, “Machine Learning” 20). Metadata, or information about the digitized newspapers, is created to aid in their organization, discovery, and access. This may include information about the newspaper title, publication date, location, and keywords. The digitized newspapers are then uploaded to an online platform, such as a website or database, to make them available for public access and research.
Scholars typically employ keyword searches and filtering options to construct a corpus for approaching specific research questions, a method known as the bottom-up approach (Bunout 287). In contrast to these corpus-creation techniques, the newspapers for this study were drawn from the Chronicling America database as part of the Oceanic Exchanges project using the entire digitized collection of each given newspaper. Based on the metadata feature “language: German” and “date: 1830–1900,” I received 36 German-language newspapers in 2018 as unstructured and structured data, which were Chronicling America’s available digitized German-language newspapers at that time. Selection criteria are of crucial interest for scholars since institutions decide what is being digitized and how one can access it (online). One of the objectives of Oceanic Exchanges was to gain a clearer understanding of the ways digitized collections were built, both in terms of the institutional decision-making and the metadata structure of newspapers. Chronicling America’s newspapers are a collection of sources that are derived from different digitization programs which has an impact on various factors from selection criteria to OCR quality (as different OCR software has been used over the years). To include titles in the Chronicling America database, the National Digital Newspaper Program stresses four primary considerations for selection guidelines, which Beals and Bell summarize as follows:
First, the title should reflect the political, economic and cultural history of the state or territory, with preference given to titles recognized as “papers of record.” Second, they should provide state, or multi-county, coverage of the majority of the state or territory’s population. Third, titles with longer chronological runs are preferred over those with short or sporadic runs. Finally, particular consideration is given to titles that have ceased publication and therefore are less likely to be digitized by other providers. (12)
These guidelines, however, do not completely conform with the selection of Chronicling America’s digitized German-language newspapers. First, they do not all fit the definition of “papers of record” as the example of newspapers from Pennsylvania illustrates. The dataset covers four newspapers from the total of “8 daily, 36 weekly, 9 monthly” that were published in the state of Pennsylvania between 1733 and 1955 (Arndt and Olson 501): the weekly Scranton Wochenblatt (1865–1918) from Scranton, the weekly Der Friedens-Bote(1812–1932) and the weekly Der Lecha Patriot (1827–1872) from Allentown, and the weekly Der Liberale Beobachter und Berks, Montgomery und Schuylkill Caunties Allgemeine Zeitung (1839–1864) from Reading. Pennsylvania has played an important role in the history of German migration, which is also reflected in its history of the press. Despite being a publication hub since the eighteenth century, the dataset only includes newspapers from towns, which their periodization (weekly) reveals, and does not include any paper “of record” published in Philadelphia. Why was the Scranton Wochenblatt, for instance, prioritized for digitization compared to other newspapers published in this state with much higher subscription numbers such as the daily Philadelphia Freie Presse (1848–1887), which in 1880 had 7,400 subscriptions compared to Scranton Wochenblatt’s 1,200? Additionally, why is the Scranton Wochenblatt available through the Library of Congress’ Chronicling America archive, while the Philadelphia Freie Presse is available through another digital archive (HathiTrust)?[6] One reason is that Chronicling America is based on different state digitization programs. Penn State University Libraries was one of the first NDNP grantees. So far, they have received two rounds of supplemental funding and “are the third largest CA [Chronicling America] contributor” (Cordell, “Q i-jtb” 210). To receive the grant, they had to follow specific guidelines such as focusing “on digitizing material from microfilms,” “expanding the time frame,” and selecting titles from “counties with very little or no digitization” and “that represent the Commonwealth’s German” ethnic heritage (Cordell, “Q i-jtb” 210–211). Thanks to a “generous gift from the German Society of Pennsylvania,” Penn State University Libraries had already received funding for digitizing “nine German-language newspapers published in Philadelphia” (Vilanova).[7] The NDNP did not grant money to the city “papers of record” because they had already received funding. Instead, they focused on weekly papers in towns with a high representation of German Americans. While the HathiTrustcollection focuses on one specific location, Chronicling America brings together German-language papers from three different locations connecting urban and rural areas.
Table 1: German-language newspapers from Chronicling America. All daily newspapers started as weekly editions but after the first years of publication started to publish daily.
| Title | City | State | Start | End | Frequency | Subscription Numbers (1890) |
|---|---|---|---|---|---|---|
| Der Demokrat | Davenport | Iowa | 1851 | 1918 | daily | 2,300 |
| Der Deutsche Beobachter | New Philadelphia | Ohio | 1869 | 1911 | weekly | 1,008 |
| Der Deutsche Correspondent | Baltimore | Maryland | 1841 | 1918 | daily, weekly | 12,250 (daily), 4,118 (weekly) |
| Der Fortschritt | New Ulm | Minnesota | 1891 | 1915 | weekly | 1,000 (1900) |
| Freie Presse für Texas | San Antonio | Texas | 1875 | 1918 | daily | 850 (daily), 9,000 (Sunday) |
| Der Friedensbote | Allentown | Pennsylvania | 1812 | 1932 | weekly | 3,000 (1895) |
| Hermanner Volksblatt | Hermann | Missouri | 1854 | 1928 | weekly | 1,000 |
| Der Lecha Patriot | Allentown | Pennsylvania | 1838 | 1872 | weekly | / |
| Millheim Journal | Millheim | Pennsylvania | 1872 | 1912 | weekly | 960 (1892) |
| Nebraska Staats-Anzeiger | Lincoln | Nebraska | 1881 | 1901 | weekly | 2,700 |
| Der Nordstern | St. Cloud | Minnesota | 1874 | 1931 | weekly | 2,540 |
| Ohio Waisenfreund | Pomeroy | Ohio | 1874 | 1953 | weekly | 40,000 (1880) |
| Richmonder Anzeiger | Richmond | Virginia | 1853 | 1917 | daily | / |
| Scranton Wochenblatt | Scranton | Ohio | 1865 | 1918 | weekly | 835 |
| Süd Dakota Nachrichten | Sioux Falls | South Dakota | 1890 | 1901 | weekly | / |
| Die Südliche Post | Goldsboro | North Carolina | 1858 | 1877 | weekly | / |
| Der Westbote | Columbus | Ohio | 1843 | 1918 | weekly | 4,500 |
| Westliche Blätter | Cincinnati | Ohio | 1865 | 1919 | weekly | 23,000 |
| Der Vaterlandsfreund | Canton | Ohio | 1829 | 1846 | semi-weekly | / |
How do these locations correspond to demographic patterns? Seen from a national perspective, the corpus does not only include papers from towns in Pennsylvania but from 17 locations as table 1 reveals. The dataset represents areas such as the Midwest, Texas, and East Coast, in which German migrants settled with “ambitions like building a new German world in different surroundings” (Häusler 114). The newspaper titles range from Maryland to Texas. The corpus covers states such as South Dakota, Pennsylvania, Missouri, Maryland, and Ohio which have a high German ancestry. The corpus does not include papers from Alabama, Arkansas, California, Delaware, District of Columbia, Illinois, Idaho, Michigan, North Carolina, North Dakota, West Virginia, and Wisconsin, which have only been digitized and added after 2018. Thus, this corpus only functions as a sample of what has been published (and digitized at a specific point in time). “CA is not a single digitization project run by the Library of Congress, but the portal to data generated by the NDNP, which awards grants to groups in individual states seeking to digitize their historical newspapers” (Cordell, “Q i-jtb” 208). Such state-level granting means that some states with a German population are represented, others less so. For instance, New York has not yet participated in the NDNP, meaning that Chronicling America also includes no papers from New York City such as the New Yorker Staats-Zeitung (1834–). However, even if New York would participate, the New Yorker Staats-Zeitung could not be selected for digitization based on the NDPD’s guidelines because, first, the paper has not ceased to exist. It is still published to this date. Second, it is already being digitized by the New York State Library. A complete digital representation of any periodical collection, except for the Finnish press, for instance, which is unique since it covers all published issues up to 1920, will always be out of reach because newspapers continue to be siloed in physical and digital archives (Salmi et al. 14). While not all papers have been digitized, some do not exist anymore, and others have been wrongfully classified or remain hidden in physical archives.
Subscription numbers (if available) are entry points to study the extent of circulation and, simultaneously, the popularity of a newspaper. Studying how these numbers changed throughout time offers insight into migration dynamics, adaptability, and quality. As shown in the previous chapter, scholarship of the German-American press has predominantly focused on papers from metropolitan eras, which had a high subscription flow, to write histories about the circulation of German-American newspapers. James Bergquist, for instance, proposes the hypothesis that the so-called “papers of record,” such as New Yorker Staats-Zeitung and the Philadelphia Alte und Neue Welt, were probably “most frequently copied by others” since “they were known for their editorial talent, were situated at active German cultural centers, and had the most immediate access to fresh news from Europe” (“The German-American Press” 135). While his argument is a compelling one, Bergquist does not provide examples of texts that were reprinted in different locations. His claim elucidates the role of the “influencer,” to use a term popularized by social media users, neglecting other places of active German cultural centers and what texts they copied and modified. His assumptions are not based on textual evidence, but partly on subscription numbers compiled by Arndt and Olson. In contrast to Berquist’s study, the newspapers used in this study range from high to low subscription numbers. These numbers, however, changed throughout the century and were influenced by advances in printing technology and German settlement movements and migration waves. Even though Arndt and Olson’s work provides these numbers, it also shows data gaps: there is, for instance, no entry for the Ohio Waisenfreund. In some cases, they were not able to collect the subscription numbers for every decade.[8] Given the fact that the newspapers were well established by the end of the nineteenth century, most data are available for 1890. The corpus covers small papers such as the Hermanner Volksblatt(1890: 1,000) in Missouri published in predominantly monoethnic environments, middle-size papers such as Der Westbote (1890: 4,500) in Ohio, which saw a large migration of German and Irish people in the last decades of the nineteenth century, and large papers such as Der Deutsche Corrrespondent(1890: 12,250) in multiethnic centers such as Baltimore (Arndt and Olson 240, 476, 188).[9] While subscription numbers varied among these places as the examples illustrate, what they have in common is that the subscription numbers doubled or, in some cases, even tripled in the second half of the nineteenth century. Der Westbote (1843–1918), one of the forty-four weeklies in Ohio (Columbus), tripled from 1,250 in 1880 to 4,000 in 1890 (476). Der Nordstern(1874–1931), a Roman Catholic newspaper published in English and German in St. Paul, Minnesota, did not change much between 1880 (2,200) and 1890 (2,540), but tripled in 1906 (7,320) (227). It was one of five German newspapers in Minnesota in 1870. The number of papers tremendously increased at the turn of the century. Further south, the Freie Presse für Texas(1865–1945) in San Antonio, Texas, first doubled almost from 2,000 (est.) in 1870 to 3,500 in 1880, and finally tripled to 9,000 in 1890 (630). The paper was established directly after the Civil War and had one of the longest time spans even surpassing the challenges faced by the German-language press during World War I. However, its publication ceased simultaneously with the end of World War II. The statistics about these subscription numbers can be used for a variety of purposes. One can analyze them as economic data to understand trends in the market or as social data to study the relation between subscription and migration waves. The data, however, are more than numerical representations, a quantity or an amount, an arithmetical value, used to count or to make calculations. These numbers represent migrants, newspaper readers, men and women, young and old, Jews and Protestants.
To complement insights into publication histories, I heavily relied on newspaper metadata from the non-digitally available database compiled by Arndt and Olson. This step was necessary as metadata created by the Library of Congress is in some cases limited or even misleading. Publication dates in Chronicling America reflect the time frames of the available digitized content, which do not necessarily resonate with the entire publication period of a paper. Since a digital archive such as Chronicling America prioritizes quantity instead of quality, that is making more titles available instead of focusing on fewer titles and providing extensive context on them, it is up to us scholars who are using these materials to make the reader aware of the blank spots in the collection (Cordell, “What has the Digital Meant” 4). Since the NDNP stresses that titles with longer chronological runs should be preferred, one would not expect to find those with short or sporadic runs. While this is the case with most newspapers in the corpus, there are exceptions. One of the “shortest,” but also one of the oldest papers in the dataset with only a one-year span of publication, is Der Vaterlandsfreundfrom Canton, Ohio. According to Arndt and Olson, Der Vaterlandsfreund was one of forty-four weeklies next to the six daily and eight monthly German-language newspapers in Ohio (427). The paper, however, did not only exist for one year. It was first published by Johann and Salomon Sala from 1829 to 1846 as a Democratic weekly (434).[10] The structured data on Chronicling America, however, says that Der Vaterlandsfreund was published from 1845 to 1846 by H.J. Nothnagel & Co.[11] Next to the metadata (e.g. publisher or date of publication) and the data (machine-readable text), Chronicling America also publishes bibliographic information about the titles, which they refer to as “newspaper title essays.” The short text Der Vaterlandsfreund on Chronicling America reveals that Johann Sala bought the paper in 1827 from Eduard Schäffer, who had already founded Canton’s first German-language newspaper in 1820 titled Westliche Beobachter und Stark Caunty. Before Sala bought it, it had undergone several name changes, as it was distributed to neighboring counties, including Columbiana, Tuscarawas, and Wayne. Starting in 1831, Peter Kaufmann began to publish the paper and throughout his period it changed titles several times, before Heinrich Joseph Nothnagel, Kaufmann’s son-in-law, took over in 1842. The information on Chronicling America is correct in the sense that Nothnagel was the publisher from 1845 to 1846 (as well as insufficient as we do not find their family affiliation). However, the metadata alone display a limited view of the paper’s entire existence. The dates of publication as found in the metadata and the available material in the archive in digital or physical form do not necessarily correspond with one another. Metadata from Chronicling America also do not reflect its changing owners and titles. As it turns out, there are three other entries of the paper in the dataset: Der Vaterlandsfreund (1833–1836), Der Vaterlandsfreund und Westliche Beobachter (1836–1837), and Der Vaterlandsfreund und Geist der Zeit (1837–1845). Even though it is the same paper, digitization produced different entries whenever the metadata “publisher” changes. The corpus in this study reflects data and metadata about 36 different publishers, but “only” 18 newspapers. In contrast to physical archival practices that categorize publications such as newspapers according to their dates, digitization differentiates between publications once it finds a new publisher on its masthead. In addition, the weekly editions or Sunday editions of some papers, such as Westliche Blätter, the Sunday edition of Tägliches Cincinnatier Volksblatt(1836–1919), were added without their daily editions as shown in table 1. Der Vaterlandsfreund is not the only newspaper in the dataset with several entries. Digitization here draws unwilling attention to an important question of newspaper publishing: Does a newspaper have to be considered a new paper once it changes its owner? My personal answer is no. However, digitization makes us aware of the break or change that occurred in the publication of the newspaper.
Influenced by previous decisions of prioritizing and eliminating, blind spots and gaps are a given in archival collections whether they are digitized or not. Comparing and linking different sources and sites is necessary, even though it is very time-consuming, to uncover what is represented in the data, how this information relates to other bibliographic works, and, finally, who is represented in the data and how.
2.2. Unveiling Gender Metadata in the Digital Archive¶
When Chronicling America incorporated digitized German-American newspapers into their database, they added the metadata “publisher” to complement the newspaper text as machine-readable data and convert it into a structured database. As Hauswedell et al. have shown, institutions that digitize newspapers and make them available to the public via web interfaces aim for simplicity (154). This implies among other things that even though information about the sources is being added, such as metadata “publisher,” these data often just include one person even if the paper’s publisher changed throughout the century. The digitization of a daily paper of record Der Deutsche Correspondent displays such a case. Der Deutsche Correspondent (which is featured on the cover of this book) was a Democratic newspaper. It held the distinction of being the most widely read and enduring daily newspaper among the 120,000 German-speaking population in Maryland (144). Situated in Baltimore, the newspaper benefited from its proximity to the U.S. capital and Europe, providing it with convenient access to news and information. Despite facing competition from numerous English-language newspapers in its vicinity, Der Deutsche Correspondent maintained its success until World War I. When looking at the metadata for Der Deutsche Correspondent, the queryreturns the features such as “publisher” (F. Raine), “dates of publication” (1841–1918), “frequency” (daily<1847–1918>), “language” (German) as well as “notes,” where one receives information that “Description based on: Jahrg. 6, Nr. 63 (26. Mai 1846).” Apart from digitized issues, especially covering the first years of the paper’s existence, what is missing of Der Deutsche Correspondent in Chronicling America are the names of its other publishers as metadata features. According to Arndt and Olson, the paper had at least two other family members as publishers (and a handful of other editors) at different times (188). The paper was founded by Friedrich Raine, a Prussian emigrant, in 1841. His brother Edward Raine, however, took over the paper in the 1860s while the older brother pursued a civic career (Blaschke, Die Entdeckung 145). When zooming in on the newspaper page, the masthead reveals that “Frederick Raine” was the publisher. When looking at the digitized pages after Edward Raine took over, however, one can detect that information about a specific person as a publisher disappeared. In 1911, Edward Raine passed away, leaving the paper to his daughter Annie V. Raine. This information is not, however, available in the metadata. During this time, it was still not necessarily common to leave one’s business to the daughter. If Chronicling America would (or rather could) include various metadata entries about publishers, readers could also identify other women as publishers in other digitized newspapers.[12] The newspaper title essay on Chronicling America also leaves out Annie V. Raine’s contributions for a female readership, which Blaschke’s research brings into focus. One year after she officially took over the paper, she introduced a women’s page and successfully managed the paper (Die Entdeckung 146). Under her leadership, more space across all pages was officially declared to women, the women’s page in the Sunday edition was removed and subscribers instead received the monthly magazine Die Deutsche Hausfrau. By circulating a national magazine, Annie V. Raine helped in creating a network not only among readers in the local region but also across the entire country. This displays the cooperation between German-American publishers in securing the survival of German-language publications. Blaschke, who examines reprinted texts in women’s pages of five newspapers (1911–1912), argues that similar fictional and non-fictional texts were published in Der Deutsche Correspondent, the Milwaukee Vorwärts, the Milwaukee Seebote, and the New Yorker Staats-Zeitung (154). Such observations already hint at the network connection between papers from Maryland and New York and Milwaukee (which are not part of the corpus used in this study). Even though Blaschke does not provide specific examples of text reuse, she further notes that there was already a noticeable increase in women-related topics and sections in the press prior to this period, indicating Annie V. Raine’s active role in editing the paper before 1911 (146). As Blaschke’s study indicates, in order to study the gender history of the German-American press, we have to move beyond metadata about publishers and instead examine the evolution of the newspaper’s content across time to receive information about the people behind the papers. Therefore, it is useful to dig into the newspaper’s content including its layout to receive further clues about who edited, wrote, and read the paper.
As the publisher of Der Deutsche Correspondent in the beginning of the twentieth century, Annie V. Raine had to navigate the decline of financial opportunities due to a decrease in German readers and advertisers, as well as growing xenophobic sentiments arising from Germany’s participation in World War I (see Kluge, “The First World War”). It was crucial for her to strategize how the paper would position itself as a German-language publication in the United States. Like many other papers in this dataset, such as the Freie Presse für Texas or Westliche Beobachter, Der Deutsche Correspondent ceased publication by the end of the war. The time frame of this study (1830–1914), therefore, provides fertile ground to examine the period of excessive growth and decline of the German-American press. It is particularly interesting since most of the German-language papers were published during this period as a direct consequence of the increasing waves of German migration, the transregional expansion of migratory settlement, and advances of technologies in print and communication. Most German-language newspapers were established after the end of the American Civil War (1861–1865), and by 1874 over 300 periodicals including newspapers and magazines were already in circulation. This number even rose to 796 by 1890 (Blaschke, Die Entdeckung 80–82). For the migrant newspapers to flourish, publishers depended on a growing body of German speakers. Local studies, for example from Missouri, which used census data to analyze the communication of German migrants and their descendants, provide “evidence that ethnic newspapers were an effect rather than a cause of language” retention, and that “the crucial variable was the absolute numbers of Germans” (Kamphoefner, “German Language Persistence” 3). Studies have shown that the German language persisted much more strongly in rural areas and particularly among the farming population, compared to the large cities where “exposure to American culture and the English language were practically unavoidable” (ibid.). There, the conventional areas of language retention, such as family, school, church, community life, and the press, were more effective than in the growing, dispersed, multiethnic society on the West Coast (Häusler 122).
There are only two newspapers in this study that survived during World War I: Der Nordstern and Ohio Waisenfreund. One newspaper that flourished up to the end of the 1920s in a predominantly monoethnic environment was the Hermanner Volksblatt. As of February 2022, if one goes on Chronicling America and generates a query to find out how many German-language papers the site includes, the query returns the number “96.”[13] However, when reading through the list of titles, one realizes that this number is also misleading at first, as some papers are listed several times with different time frames of publication. The Hermanner Volksblatt, for instance, appears three times in this list (1860–1871, 1872–1873, 1875–1928). These numbers are not in line with the actual publication time frame of the paper (1856–1928) and reveal gaps in the digitized collection as the years between 1873 and 1875 are missing. Despite these inconsistencies in the timeframe, there are also anomalies related to its listed publishers. While the first two units list “Jacob Graf” and “Carl Eberhardt” as the publisher, the third one informs that it was published according to the metadata by“Frau Graf & Co.” (Mrs. Graf & Co.). When looking at one example of an issue, such as on June 24, 1875, one can see on the top left in the masthead: “Frau Graf & Comp., Herausgeber” (see figure 5).[14] This indicates that (in some cases) information for the metadata was extracted directly from the physical objects. Instead of treating this newspaper as one publication, it is broken into three different pieces as publishers change. Even though Chronicling America contains the metadata “succeeding titles,” it lists the paper in three different units. When Jacob Graf passed away in 1870, as the essay on Chronicling America reveals, his widow, Christine Graf, took up his work and published the paper until 1873.[15] She sold it to Charles Eberhardt but bought it back in less than a year. In contrast to the bibliographic information, this particular digitization division into three units indicates—and a careful examination of the masthead also provides evidence for this. She did not publish the paper in 1873; instead, it was “Karl Eberhardt” who published it under the title Hermanner Volksblatt u. Gasconade Zeitung. As this shows, it is not only the metadata on Chronicling America that can be misleading, but also the newspaper title essays. After her husbands’ passing, Christine Graf published the paper not on a strict weekly schedule, but intermittently—every couple of weeks—under the name “Die Wittwe Graf” (“the widow Graf”).
When Christine Graf assumed ownership of the Hermanner Volksblatt again, she also became owner of the Hermann Advertiser (also The Advertiser-Courier [1875-1964]), which was the first English-language paper in the area. This paper has been digitized but not integrated into the Chronicling America database. However, it can be found online in the “Directory of U.S. Newspapers in American Libraries” by the Library of Congress.[16] The directory is a compilation of metadata detailing the number of titles that the Library of Congress believes have existed in the territory now known as the United States. Instead of “published,” the directory features the metadata entry “Created/Published,” which in the case of the Hermann Advertiser (1873–1877) lists “Hermann, Mo.: C. Graf & Co., -1877.” Next to Arndt and Olson’s metadata compilation, this directory is also useful for a study on the American press system. While it is large, this collection is also far from complete. One the one hand, these metadata indicate that Christine Graf had prior involvement in the newspaper industry before her husband’s passing, as she not only took over the paper but also owned multiple businesses. On the other hand, Christine Graf not only provided German but also English-language material to the reading community of Gasconade County shortly after the Civil War and in the decades before the World War I, emphasizing her contribution in establishing a growing multilingual environment in a rural area in Missouri. The paper, which was later owned and edited by her sons, ceased publication after 71 years in 1928. Were Anna Uhl Ottendörfer and Christine Graf exceptions? According to Blaschke, cases of (German-)American history of the press reveal that it was not unusual that widows took on editorial work after their husbands’ passing (Die Entdeckung 321). As these examples show, even if women took over the papers, we do not necessarily find adequate information to study their involvement based solely on metadata of both analogue and digital sources. Inadequate information baked into metadata contributes to women’s invisibility in the archive.
Figure 5: Masthead of Hermanner Volksblatt from June 24, 1875, listing as Frau Graf & Comp. as publisher.

Newspapers are widely recognized as valuable historical resources, especially for smaller communities where they may be the sole surviving record of local events (Arndt and Olson 8). As one can see in figure 5, apart from metadata from Chronicling America, the Directory of U.S. Newspapers in American Libraries, and Arndt and Olson, the layout of a newspaper can further reveal information about the title, the town, its German population, and its owner. The Thursday edition of the Hermanner Volksblatt, which could be obtained through a subscription of 2 dollars per year, printed poetry and fiction on its front page. Christine Graf did not change the price of the paper. A decade earlier, in 1860, when Jacob Graf was still officially the publisher, the yearly subscription cost was already 2 dollars. The paper remained four pages long, adhering to the standard American practice of the time. The outer inner side was printed and then folded to create a four-page newspaper. Additionally, even in 1860, the front page primarily featured texts that could be categorized as feuilleton. However, by the 1870s, these texts were framed by various advertisements for German(-American) products and services. I am adding “American” here to emphasize that ads about beer products and drinking venues (“Hermann Brauerei,” “Wein u. Bier Saloon von John Pfautsch,” “St. Charles Hall Winde & Beer Saloon”), verses (“Warnung”), and tales (“Der Friedensrichter: Ein amerikanisches Lebenselixier”) do not only show the transcultural and multilingual dimension of the migrant press, but also to illustrate the high impact of advertising (and entertainment) even in the case of a newspaper from a small-town monoethnic environment. However, I am putting it into brackets to stress that this hybrid nature of the press has received little attention from scholars so far next to the examination of the role of women in these developments. Christine Graf’s traces in the data highlight that even though she was a successful entrepreneur, in both the physical and the digital object, she simply remains “Frau Graf” (Mrs. Graf). By omitting the metadata, that is her surname, it is implied that she is defined by her husband. During that era, such labeling was prevalent, but it symbolizes the unseen contributions of German migrant women in the newspaper landscape. Officially, most publishers, editors, and journalists were men. However, the close professional entanglement between couples was not uncommon even in these days (Blaschke, Die Entdeckung 322). According to Blaschke, at least 22 of the male partners of the 70 Germans-speaking female authors that she lists in her last chapter were similarly engaged in the editorial or publishing business. Nevertheless, it is usually difficult to find references to women’s names, particularly their maiden names and surnames. As Blaschke further points out, we can assume that an even larger number of women worked alongside their husbands without being duly acknowledged in the sources (323).
A digital source criticism is necessary for every study that uses digitized material not only to evaluate the representativeness of the corpus for a particular research question, but also to evaluate the implementation of guidelines and to formulate changes, if necessary. Databases such as Chronicling America provide access to millions of newspaper pages ranging from urban newspapers and cover more than twenty different languages. But like most archives, they can never be complete. “[W]hether housing physical materials or seeding virtual information in ones and zeroes in the cloud,” as Abel points out, “all archives are selective. They are selective, of course, because hardly everything of the past survives—there are always absences, gaps—and, equally important, because someone has selected from what does survive, according to certain principles” (6). Only few newspapers have yet been digitized. “While the past decades have seen significant efforts toward digitization and digital curation, digital libraries represent only a fraction of the materials held in physical libraries and archives” (Cordell, “Machine Learning” 34). Many sources cannot or will not be digitized because some do not exist anymore, or some cannot be read by OCR software due to damaged pages. Some sources remain undisclosed either due to incorrect classification by archivists or because they have not yet been classified. Additionally, certain sources may be inaccessible as they are behind a paywall. Criteria that address OCR quality, selection of publications, or data formats must be examined by scholars to evaluate the use of digital archives and tools for historical practice (Priewe 406–407). Digitization will continue and digitization guidelines will change during this process because scholars are increasingly using digital sources and uncovering biases or archival gaps. For the English-language newspapers in Chronicling America, for instance, Ben Fagan found out, when he analyzed the collection, that it privileged newspapers that served “the dominant white, middle-class readers of the nineteenth century” (11). This has led to the near exclusion of Black and other minority-run papers such as the ones published by Native Americans, migrants, or women. Pointing out these inequalities has already led to changes in selection criteria (Cordell, “What has the Digital Meant” 3).
The examples above illustrate that publishing a newspaper was more than just the achievement of a single person. First, it depended on a German-speaking readership in the urban and rural areas. Second, publishing a migrant newspaper was often a family business and a collective effort and its success depended on the next generation’s participation and willingness to read German-language publications. Like a few other exceptional women of her time who owned newspapers in the U.S.—think of Susan B. Anthony, Elizabeth Cady Stanton, Ida B. Wells-Barnett, Matthilda Franziska Anneke, or Amalie Struve, to name just a few—Christine Graf contributed significantly to the formation of transnational imagined communities. And yet her having once been at the epicenter of a migrant newspaper network during the peak of the German-American press is a little-known fact. Being a woman and a migrant, meant, in her case as in that of most other migrant women, being judged and labeled based on gender all the time. But it would be far-fetched to reduce Christine Graf to being a victim incapable of acting, as it was precisely her conditions, being a migrant and widow in a rural area, that provided privileges, agency, far-reaching opportunities, and networks that others lacked. While the examples of Christine Graf and similarly Annie V. Raine illustrate the biases and blank spots that have been directly baked into archives and digital technologies in the form of missing or abbreviated names impacting methods and historical narratives, they also emphasize women’s publishing endeavors beyond famous feminist figures. As these examples show, digitization to this date further strengthens the image of the German migrant newspaper landscape as a male-run business that focuses on the individual instead of on the collective. This above all impoverishes the understanding of women’s participation in the past as long as they remain invisible. “As several generations of feminist labor studies scholars have observed,” Klein stresses, “it is both a cause and an effect of this invisibility that these forms of labor are undervalued and undercredited (or uncredited altogether) in the end result” (“Dimensions of Scale” 24).[17]
As this study includes only two newspapers owned by women, I aim to demonstrate that research on the role of women in the German-American press can be conducted by examining newspapers owned by men. Implementing this approach necessitates a close examination of the newspapers’ textual content. A layout analysis alone, without the need to shift through thousands of pages for widely reprinted texts, already reveals the high visibility of women in these publications. The following example of a Texan newspaper seeks to illustrate not only the different layouts, periodization, and costs of German-American newspapers, but also highlights that gender was a highly debated topic, extending beyond discussions of the suffrage movement, as mentioned in the introduction of this book. As one can see in figure 6, the Freie Presse für Texas (May 12, 1899) edited by Robert Hanschke had a different layout compared to the Hermanner Volksblatt.[18] In contrast to the small-town paper, this title was published daily and, therefore, resulted in a higher subscription cost (10 dollars per year or 2.50 for three months). While the weekly paper had a more structured layout, with the front page dedicated to advertising and entertainment followed by news on subsequent pages, the front page of the daily Texan newspaper exhibited a high degree of content mixing. Almost the entire right side was covered in ads, framed with ornamental designs that indicated their commercial nature. Adjacent to the ads, readers found information about the train schedule, as well as local, national, and international news. In these texts, no particular gender was addressed except for the largest textbox on the top right, which was an ad for a men’s clothing store. Interestingly, the ad addressed women, starting with the phrase “Eine fashionable Dame” (a fashionable woman), suggesting women as potential purchasers of men’s clothing. While Blaschke argues that women’s pages and magazines primarily targeted women as key consumers, similar marketing strategies can also be observed in non-gendered sections of the German-American press (“German Immigrant Women” 327).
Figure 6: Interface of Chronicling America displaying the front page of Freie Presse für Texas, in San Antonio, May 12, 1899.

When examining the left side of the front page in contrast, gender discussions and differences between men and women become visible. Since its inception, the front page of a newspaper has been considered the most important, designed to immediately capture the reader’s interest and entice them to purchase the paper. In other words, it matters whether a text appears on the front page or the last. Front page texts were the first encountered by readers, granting them not only greater visibility but also heightened perceived importance. The front page of the Freie Presse für Texas did not start with an editorial but instead directly addressed readers’ demands for fiction by featuring an excerpt of “Die Wildkatze” (“The Wildcat”), a serialized novel by German writer Ida Peisker. Unlike other examples discussed in this study, Peisker’s name was not abbreviated, which demonstrated her popularity, preceding her story on gender roles, stereotypes, and women’s rights, which featured an independent female protagonist, often referred to as “the wildcat.” In between the text units of the serialized novel, Hanschke also offered more fruit for discussion about gender and health, particularly issues that exclusively concerned masculinity and men’s issues, by integrating an advertising text about a booklet titled “Männer! Eine Warnungsstimme” (“Men! A Warning Voice!”). This advertisement offered factual information for men experiencing a loss of sexual strength (whose “geschlechtliche Stärke verloren geht”) and guidance on how to regain it. The booklet was offered for free by a doctor from Chicago, with recipients only needing to pay postage costs. The advertisement assured readers that their personal information would be handled with the highest discretion. This ad highlights that men were also facing issues in the context of gender and health. However, in contrast to gender issues that related to women that were represented and debated in numerous ads (see chapter 5), detailed discussions about men’s problems were predominantly redirected into the private sphere, where they could be handled in a discrete manner, rather than openly discussed in the press. Although the layout (and significance) of the front page has evolved over time, this exemplary page provides insights into the German-American newspaper and its audience indicating that it catered to both men and women. It reveals that gender was a topic of discussion in both novels and advertisements—even on the front page of a newspaper from South Texas. The question arises of how to “mine” thousands of newspaper pages to examine women’s participation as readers and writers in the news beyond ownerships and editorships in the long nineteenth century.
2.3. Chronicling America’s Interface: Interacting with Software through Visual Elements¶
The analysis of a single page demonstrates that women were visible as readers and writers, and it shows that gendered discussions in the press encompassed more than just the role of women. Unlike the limitations of physical archives, digital searches enable extended time frame analyses and the integration of newspapers from various locations. Instead of having to travel to different archives, we have newspapers from rural areas in Pennsylvania to urban centers in Texas in one location. This capability facilitates comprehensive quantitative and qualitative approaches to study transregional information flow and investigate how gender was constructed and debated in news, advertising, and entertainment content. Before the widespread use of computers, historical research on newspapers had been a challenging task (Oberbichler et al. 226). Scholars had to shift through massive collections of news items, often needing to travel to other institutions to access local publications (Priewe 404). This laborious and time-consuming process did not guarantee the discovery of relevant artifacts at all. Modern computing and digitization, however, have transformed the preservation and accessibility of these resources. Digitized historical newspapers are now globally available via online databases, which has eliminated barriers of access and enhanced the speed and efficiency of search methods. “No longer do I have to take frequent, lengthy trips to public libraries or state historical society libraries to view microfilm copies of local and regional papers, sometimes on less-than-reliable readers, and depend on photocopies, often of poor quality” (Abel 2). The availability of digitized historical newspapers from a diverse array of national, cultural, and linguistic sources offers scholars new opportunities to explore the conceptual history of news network systems and how certain concepts, topics, and news items traveled across time and space. However, each digital archive comes with its own set of possibilities and limitations ranging from diverse (and often insufficient) metadata standards to the quality of OCR-derived text. While structured data such as metadata provides information to learn more about the history of the newspapers as well as of their preservation, unstructured data of the newspaper content generated through OCR software allows for full-text search and to explore the pages of Chronicling America’s German-language press archive by means of keyword search (KWS). KWS are information retrieval applications that simplify the process of finding relevant data. However, the effectiveness of these tools depends on the accuracy of OCR transcriptions. “One recurrent concern [when performing queries in these interfaces] is the impact of OCR quality on the resulting list” (Bunout 288). Since Chronicling America is based on different state digitization efforts and rounds of digitization that have worked with different software over time, “[t]he quality of the OCR data” in their database “is so variable that any summary would be largely inaccurate, even for a specific title” (Beals and Bell 12).[19] The quality of the data, therefore, differs tremendously among the newspapers that were chosen in the beginning to the ones that have been recently added. Evaluating OCR quality, however, is essential since it is this machine-readable text that is being used for the full-text search. The data quality as regards OCR varies considerably owing first to being digitized by different actors with different software applications, but second to variations in the source material and its changes throughout the course of the nineteenth century. In 2023, researchers of the Living with Machines project, which studied how people experienced the Industrial Revolution through the British press, revealed that searching the British Library’s digitized newspaper collection would return politically biased results.[20] According to the scholars, OCR was better at reading the fonts favored by more expensive and conservative papers than those used by less expensive, liberal ones.
“[I]n spite of missing documentation, researchers need to develop strategies to assess the quality of OCR in different titles or periods, by using a few test keywords, of different lengths or frequency and focus on segments of time or titles in the collection they are using” (Bunout 288). Before showing how Chronicling America’s interface is limited in finding reprinted texts, I want to show how to manually detect and evaluate OCR errors by examining different representations of the material. To find particular texts, users can enter keywords, select a language and a specific time frame on Chronicling America to receive a list of documents in which a word or words appear. The interface allows users not only to access the scanned image of a newspaper page as shown in figure 6 (“PDF” or “JP2”) but also the machine-readable text (“Text”) that has been generated through OCR.[21] The user has two options to close read the text of such an ad as “Für $ 3.00 nach Deutschland” (For $ 3 to Germany), which was printed on the second page of this edition, online by examining the entire digitized newspaper page as a scan or the machine-readable text of the article.[22]
Fiir 83.00 nach Deutschland.
Deutsch – amerikanische Zeitungen
werden ini alten Vaterlande mit gro
ßem Interesse gelesen, weil sie sich frei
aussprechen nnd häufig von driiben
mehr Neuigkeiten nnd Nachrichten brin
gen, wie die deutschen Zeitungen selbst.
(Extract of OCR-ed text of “Für $ 3.00 nach Deutschland.”)
Both representations can be useful for an analysis of the newspaper. A scan allows us to examine how the text was embedded in the page or the entire edition to address questions about visual characteristics: from a distance, for instance, the text looks like a regular news report compared to the ads with large titles to the right or compared to the ads on the front page with the ornamental framing. Information including location on the page, size, number, and character proportion can turn into valuable hints about a text’s genre, but at the same time close readings are necessary to fully capture the text type. Switching between different formats and representation options is useful because it allows you to interrogate the source from different angles. Such a digital search makes it possible to quickly zoom in and out of the page to compare and contrast features, to read the page from a distance, or to zoom into specific parts of the text for close readings. While a scan also adds a flood of additional information to investigate, the OCR extract, in contrast, only shows the machine-readable text. In this way, a newspaper page is being reduced to its simplest unit: words. This view provides insight into the OCR quality and shows how some letters have been misidentified or not identified at all. In this ad, which was reprinted in several editions of the Texan paper, the umlaut “ü” was identified as “ii” and the dollar sign “$”as the numeral “8,” which means that the title and the text claim “for 83 to Germany” instead of “for three dollars.” Such errors vary even among editions. In contrast, in the one published on September 23, 1899, the umlaut was correctly identified. As the OCR extract further exposes, in some cases, the letter “m” was incorrectly identified as “ni,” “u” as “n,” or “e” as “c.” The quality of OCR output depends on a range of factors, including the quality of the input image, the type and complexity of the text, the language and font used, and the quality of the OCR software. The issue with OCR on Fraktur text lies in the distinctive and complex nature of its typography (Resch 93). Fraktur is a form of blackletter typeface, which was commonly used in German-language printing until the mid-twentieth century. Its unique letterforms, characterized by ornate, Gothic-style script with intricate serifs and ligatures, differ significantly from the Latin-based alphabets such as Antiqua (English-language texts in Chronicling America). An understanding of these and other technical processes is important as they influence the displayed results when a word or a series of words is entered into the search interface of an archive that may contain millions of inaccurately scanned and converted objects (Priewe 407). “Technologies such as OCR and search engines are often not directly visible in a historical argument, especially since historians tend to cite the physical archival sources” (Romein et al. 312). Interfaces do not enable us to quickly see and evaluate the hierarchical structure of code, or the relationship between what is being presented and the entire dataset. Additionally, it is difficult to find out what is missing or what one is missing through such searches and filtering. When conducting a search query, that is inserting a keyword or an entire phrase in Chronicling America’s search interface, the list of search results shows figure 6 instead of the OCR extract shown above, which is the source that has been used for the query. In comparison to the German Newspaper Portal, for instance, which displays parts of the OCR next to the scanned image, Chronicling America displays only the scanned images first in the search list. “This prioritization can be seen as an act of manipulation because the image recalls memories of the user reading the paper in the physical archive” as print or on microfilm (Keck, “Let’s Talk About Data” 69). Reflecting on these properties is relevant to gain a deeper insight into the data that is being extracted to create a corpus for analysis (even if one is simply downloading a page as a PDF file).
The use of machine reading on historical documents holds significant transformative potential, and the primary task ahead is to modify and advance suitable technologies that enable efficient searching, retrieving, and exploring of information from this vast “big data of the past” (Kaplan and Di Lenardo 1). Typing a keyword into a collection interface’s search box sets off a chain of “algorithmic actions and relies on a series of preceding decisions,” all of which are not visible to the user but are critical to any research based on the search results (Bunout 281). The search algorithm used to select from an unknown pool of elements is not visible, long lists of query results may hide “archival gaps” giving a false impression of completeness and the large number of results obtained can create an illusion of the importance of a phenomenon or the correlation between terms (Cordell, “What has the Digital Meant” 3–5). In addition, as datasets increase in size, traditional methods such as close reading become less effective, as Andrew Torget writes: “If, for example, a search for a particular term yields 4,000,000 results, even those search results produce a dataset too large for any single scholar to analyze in meaningful ways using traditional methods” (47–48). With such an abundance of information available, scholars may struggle to extract meaningful and novel patterns from massive datasets using basic keyword searches. The sheer volume of digitized historical newspapers available begins to pose a challenge for scholars. When conducting a keyword query, for instance, for “Weiber” (women) using the basic search bar (1830–1914), a total of 87,344 results are generated. Employing a wildcard search for “Weib*” to encompass related terms such as “Weiber” or “weiblich” (female) yields 90,950 results. However, this initial discrepancy is quickly resolved by consulting the help files, which clarify that wildcards, case sensitivity, and basic Boolean operators are not supported. The search engine employs language-specific dictionaries with stemming functionality to include word variants (see Ehrmann et al. 152). To refine or expand the search results, it is recommended to use word combinations or use the features provided in the “Advanced Search” option, which includes filters for states, titles, years, front pages, language, word combinations, phrase search, and distance search. When searching for the phrase “Mann und Weib” (man and woman), which was one of the titles in the reprinted story of this study’s introduction, a total of 1,989 results were obtained in July 2023. However, upon a quick scan of the results, it becomes apparent that the two terms often did not occur together in the same news item, rendering this keyword search ineffective.[23] Keyword searches possess the seemingly contradictory weaknesses of finding too few documents (under-inclusion) and finding too many documents (over-inclusion) (Keck, “Let’s Talk About Data” 71). If selecting, for instance, a shorter time frame or a specific date, a digital search with keywords can be promising when using specific terms related to persons or events. It is less useful, however, for the analysis of concepts (gender) or entire groups (women). The interface is additionally not advantageous for detecting reprinted texts like “Mann und Weib” or “Für $ 3 nach Deutschland” in thousands of newspaper pages. A keyword search (to find reprinted texts) in the digital archive would also presuppose that we already know which (reprinted) texts we are looking for and insert specific keywords or phrases accordingly (see Michel et al.).
The use of a digital archive of immigrant newspapers presupposes the examination of data and simultaneously the reflection on algorithmic performance in the case of filtering the metadata about language and ethnicity. To work with other non-English language newspapers, different filtering options to examine the newspapers from the perspective of language and ethnicity are possible in Chronicling America. On the “All Digitized Newspapers 1777–1963” page, users have the option to select “language” and “ethnicity” to explore the newspapers that have been digitized and made available on the site. If one selected “location: all states,” “ethnicity: German,” and “languages: all languages,” in February 2021, the search revealed that there were 59 newspapers available to view on the site. If one then selected “all states,” “all ethnicities,” and “languages: German,” the search returned a result of 74 newspapers. When combining the two approaches by selecting “location: all states,” “ethnicity: German,” and “languages: German,” the search again generated the response that 59 newspapers were available. Although the first and last selection options seemed to produce the same result, despite the differences in each step, this was due to the hierarchical organization of procedures in the programming. The algorithm followed a set sequence of steps based on certain decisions, which are dependent on the order of the selections made. In this case, the priority was (and still is) given to the “ethnicity” feature when multiple features were selected. Many newspapers in Chronicling America contain multilingual texts. However, the search used the metadata, which is not based on the individual texts, but on the newspaper as a whole. German-American newspapers also published English-language content. When one conducts a keyword search in Chronicling America and selects “English” as the language, the query, however, does not return texts that would fit this keyword in a German-American newspaper, when their metadata only lists “German” as its language.
Many pragmatic decisions implemented in digital archives are based on the principle of simplicity. In a summarized overview of interviews conducted with institutions that are digitizing sources and making them available through web-based text-search as part of the Oceanic Exchangesproject, Hauswedell et al. conclude that creators aim for simplicity. Since these platforms are created predominantly for the greater public and not necessarily for scholars, simplicity means that the traditional method of keyword searching “is here to stay for the foreseeable future, since the majority of users have been habituated to this mode of search” (155). Exploring the interface and trying out different filtering options for search queries are simple ways to get familiar with the available data formats and structures of algorithms.[24] Hiltmann argues that the current process of digitization represents more than just a regular shift. Instead, it signifies an essential transformation in how people communicate and work, impacting both scholars and society as a whole (“Vom Medienwandel zum Methodenwandel” 16). Since keyword search may not be the best solution for a specific research question, scholars have started to experiment with different computational approaches to understand how to implement humanistic ideas into computational language. However, as outlined in this chapter, historical research poses several challenges for keyword searches due to the intricate nature of language. Oberbichler and Pfanzelter’s study, for instance, also criticizes keyword search for the analysis of concepts. Their work, which focuses on finding articles on return migration in Austrian newspapers (1820–1920), similarly uncovers these challenges, as concepts “do not equate to single words” (136). These challenges are further exacerbated by alternative spellings, abbreviations, polysemy, shifting word usage, idiomatic expressions, as well as misspellings or the exclusion of parts of words and sentences as the authors further argue.
This chapter has shown that “[t]he efficacy of our scholarship” in periodical studies and beyond “depends upon a largely missing source history of these digital collections” (Fyfe 548). Large collections can create a sense of awe, a “digital sublime” (ibid.). At the same time, this awe can vanish when peeking behind the scenes of the digital archive to examine the datasets thoroughly and uncover the decision-making processes within digitization and preservation, which is not necessarily transparent to the user. This chapter has reflected on both the potential of digitization and limitations of digitized newspapers compiled by Chronicling America to study text reuse and gender quantitatively and qualitatively. By examining their database and discussing the potential and pitfalls of the interface, the chapter has provided insights into the representativeness and selection criteria of the dataset compiled for this study and uncovered some of the blind spots in the digitized newspaper collection of Chronicling America’s German-language press, such as limited and misleading metadata. As with almost any dataset, however, the corpus used in this study only represents a sample (of the five thousand German-language periodicals published between 1732–1955). While the selection from the Library of Congress does not offer a comprehensive portrayal of the German migrant press during that time, it still presents a valuable case for studying text reuse practices across different newspapers. To move beyond basic search algorithms, the forthcoming chapters will demonstrate the practicality of employing the computational study of text reuse detection. The next chapter will introduce the text reuse software and explain text reuse detection as a method to analyze reprinting at scale. It will reveal the usefulness of the approach to computationally study text reuse, and also its limitations, particularly by exploring forms that go beyond the mere repetition of words. It will show that reprinted texts and their recreations such as translations or parodies provide textual evidence to study gender roles and stereotypes relating to both men and women in the press, even though they do not necessarily contain words denoting women. The digitization and computational analysis undertaken in this study have been instrumental in identifying and examining these texts, offering insights into the experiences of migrants, particularly women. These findings also seek to contribute to contemporary discussions on digital literacy and to provide a historical context for media-related topics such as virality, plagiarism, circulation, and representation bias.
The paper was founded in 1834 as a weekly and was then turned into a daily in 1843, by Jacob and Anna Uhl. The circulation increased from 15,000 in 1857 to 41,500 (est.) in 1870 (Arndt and Olson 399). When Jacob Uhl was the owner of the New Yorker Staats-Zeitung, his wife had already been pursuing an active role in the business management of the paper while raising the couple’s six children (Bergquist, “Ottendorfer” 841).
An n-gram is a contiguous sequence of “n” items from a given document. In natural language processing, an n-gram can be a word, a phoneme, a letter, or any other smaller or larger unit of language.
Since the 2010s, it has become a common practice to conform to the standards established by the Library of Congress, specifically the METS and ALTO format. To put it simply, the METS standard can be viewed as a table of contents for all the ALTO files, which includes information about the textual and spatial arrangement of the text on each page. These files are subsequently stored in a database by the organization that requested the digitization, and an index of the textual content is generated (Beals and Bell 12).
For a detailed explanation of the digitization of newspapers from scanning the pages to how the data is being stored, see Bunout.
HathiTrust is a collaborative digital repository that provides access to millions of digitized texts from libraries and research institutions around the world.
These include Philadelphia Demokrat, Neue Welt, Philadelphia Tageblatt, Philadelphia Freie Presse, Philadelphia Morgen-Gazette, Philadelphia Schwäbischer Merkur, Schwäbischer Merkur, Philadelphia Sonntags Journal, and Volks-Stimme: das Socialistische Wochenblatt für die Ost-Staaten.
The subscription numbers and info on frequency were taken from the bibliographic information on the Ohio Waisenfreund from Chronicling America.
1890 serves as an example to illustrate publication size because information about this year was mostly available in Arndt and Olson for the papers that are used in the present study.
According to the 1870 Census, Ohio had a visible “German” presence “making the numerical strength of the German-Americans in this State second only to New York” (426).
The structured data on the site provides information among other things about title, geographic coverage, publisher, dates of publication, frequency, language, notes, library identification numbers, and holdings. Here, one can also access a text (unstructured data) that provides more information about the paper’s founding and development.
I am using the verb “could” to account for the possibility of adding comprehensive bibliographic information that comes with the creation of digital databases and, at the same time, I do not want to express too much criticism on institutions that are providing digital archives such as the Library of Congress. Digital archives such as Chronicling America have done an important job in helping to preserve fragile objects for the future and making them available to a large audience. However, very often, like physical archives, they do not have the financial resources to provide such extensive bibliographic information and, due to guidelines for selection and preservation, they very often aim for quantity instead of quality, as the interviews by Hauswedell et al. reveal.
In June 2023, Chronicling America provided access to 96 newspapers based on the metadata “language: German.” Some of them, however, were published after Das PW-Echo, which was a semi-monthly prisoner of war (POW) newspaper, circulated at Camp Rucker in Pike, Alabama (1945–1946).
Hermanner Volksblatt, June 24, 1875, p. 1.
Christiane Graf was born in Württemberg in 1820. She and her husband emigrated in the 1840s. Christiane Graf had eight children, of whom only four were alive when she passed away in 1902.
The newspaper was digitized by the State Historical Society of Missouri (SHSMO). It can be accessed through the SHSMO Digital Newspaper Project.
Klein, who focuses on abolitionist papers published by American women, shows that even though women voiced their concerns about larger social and political goals, “they were not always able to claim the same credit as men for their forward-thinking work” (“Dimensions of Scale” 34). Her study analyzes the newspaper texts to fill these gaps and reconstruct women’s collective past.
Freie Presse für Texas, May 12, 1899, p. 1.
OCR quality can be measured in terms of different factors such as the accuracy of the character recognition, the level of formatting retained (such as font styles, bold/italic text, and text alignment), and the presence or absence of errors such as incorrect line breaks, misspelled words, or missing characters.
The Alan Turing Institute, “Living with machines: The environmental scan” (video recording, 2024). Accessed July 20, 2025.
OCR is a field of research in pattern recognition, artificial intelligence, and computer vision. When records are digitized, scanning is only the first step. The software creates an image of the document, but that image, and the data that composes it, is neither editable nor searchable. OCR is a type of software that converts those scanned images into structured data that is extractable, editable, and searchable. Thus, OCR involves two steps: converting print text into machine-encoded text as well as creating information about this data: i.e. metadata. In short, metadata is data about data. Many distinct types of metadata exist, including descriptive metadata, structural metadata, administrative metadata, reference metadata, and statistical metadata. See Schöch, “Big?”
For three dollars, one could send Die Freie Presse für Texas to friends and family directly in Europe. To convince the reader of the popularity of the paper, the text claimed that it was not only widely read by German Americans, but also by people in the “Vaterland” (“fatherland”), who were reading German-American newspapers with great interest because these sources expressed their opinions freely and provided more authentic information on the U.S. than the German papers.
Bulk downloads are not available at this level of access.
For a variety of digitized newspaper collections such as Chronicling America, Beals and Bell created The Atlas, which is an open access guide to information about the histories and data of digital archives around the world. The Atlas, which is an online publication, provides information about digitization choices, the evolution of newspaper terminology and the variety of metadata available. Additionally, it shows “how machine-readable information about an issue, volume, page, and author is stored in the digital file alongside the raw content or text, and provides a controlled vocabulary designed to be used across disciplines, within academia and beyond” (Beals and Bell 1–5).