Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

While women in the German-American press were represented in news, information, advertising, and entertainment, one can observe varying degrees of their agency and visibility in different textual forms. In some of the texts, violence against women including reference to physical and emotional abuse was not only a topic in news reports about murder and rape: ideas about harassment and brutality were also evident in poetry, as a story printed on the front page of Die Freie Presse für Texas on June 17, 1884, demonstrates: “Das Mädel wird ’mal Braut / Ich haue, du haust, er haut, / Aus Bräuten werden Frauen / Wir hauen, ihr haut, Sie hauen.”[1] This short poem featured in a humorous sketch, which was found by keyword searching “Gattung” (genre) in the dataset of reprinted texts. The sketch illustrated a report by well-known German novelist Paul Lindau, which he told to a group of friends to express his view on women who write. The text informed that this short poem was his reply to a letter that he had received by a woman, “Helene v. S.,” asking for his opinion on a new type of poetry she invented, illustrated in the following lines: “Des Lebens kurze Frist / Ich bin, du bist, er ist, / In kurzer Zeit verrinnt / Wir sind, ihr seid, sie sind.”[2] Lindau explained that he simply intended to respond with a poem in the same style and praised her for her originality. This story can be interpreted from different angles: on the one hand, if one examines it through the lens of text category, one can read it as a humorous sketch. In the nineteenth century, such brief entertaining texts circulated in newspapers across the globe. Sketches enabled authors to “write about ordinary life” as it “did not require great events or disasters. Small, private events would do,” as Norman Sims observes (46). This humorous sketch is an example of what Fitzgerald and Cordell computationally classify as the vignette. They “are short prose accounts that blend fact and fiction, anecdote and sketch, information and entertainment,” which thrived in the nineteenth-century newspaper.

“Whether classified by human readers or computationally,” as Fitzgerald and Cordell argue, “vignettes are a bit like fiction and a bit like news.” Whether the incident—described above in this “untitled” short prose piece about sentiments towards women who wrote—that declared itself neither as news nor fiction, or more accurately suggested that it was both by using a “real” author such as Lindau as its narrator gossiping with his friends, occurred or was entirely fabricated, does not lie at the heart of the investigation here.[3] The objective is to examine this vignette from the perspective of gender and to analyze what information it spread about women. Lindau’s report shows how derogatory remarks about women were disguised as humor and embedded in a narrative of a well-known German author, in which a private conversation is made public through print (see Horn).[4] By adding imagery about physical violence through the repetition of the verb to “hit” (six times), the text represented not only a humorous remark about a woman’s intellectual qualifications. It also spread ideas about the reproduction or normalization of gender-based violence and criticized women who wrote poetry, claiming that first they lack the ability to create true art and second that it did not make any sense that women write since they become wives anyways. Masquerading as humor, this brief piece conveyed notions about distinct gender roles, highlighting the perceived differences in abilities between men and women.

This gendered message about women being intellectually inferior to men and whose proper place was the domestic sphere traveled further than Texas. Through the textual ecosystem of reuse, it came into the hands of editors in the South, the East, and the Midwest, who printed it in their papers, such as on June 11, 1884, in Der Nordstern, St. Cloud, Minnesota, and on June 5, 1884, in Der Westbote, Columbus, Ohio. Through scalable reading, one can trace the travel of such stories to identify similarities and differences between texts, editors, and publications.[5] When exploring the dataset of reprinted texts, one can always get redirected to the Chronicling America interface to examine how these texts were embedded into the page layout of different newspapers. This reading aids in kindling a dialogue between both the reused text(s) in different newspapers and their neighboring items on the page to identify localized differences. Gunther Kress and Theo van Leeuwen argue that the layout, i.e. the placement of elements in the visual space, provide additional information for the analysis of texts (181). In Minnesota, for instance, the text was printed on the penultimate page under “Miscellen” (“Miscellaneous Section”), where it was surrounded by the time schedule of the railroad and various ads about saloons, beer breweries, and products such as tobacco, whiskey, or “St. Jakob’s Öl,” a medical cure-all. In Chronicling America, one can see that in Texas, however, the text was printed not on the last, but on the front page on top of an ad about baking powder and under the news section “Ausland” (“Abroad”) that was preceded by an extract of a novel Gebannt und Erlöst (Banned and Redeemed) by E. Werner, which was the pseudonym of German writer Elisabeth Bürstenbinder. Just a year earlier, this novel that discusses societal conventions and restrictions for women was published in Germany’s family magazine Die Gartenlaube, the precursor to all modern magazines (Palatschek 41).[6] A decade earlier in 1876, the popular magazine had printed a profile about her that was titled Eine Heldin der Feder (“A Heroine of the Pen”). It started with an introduction about the rising antagonistic positions “welche den schriftstellernden Evastöchern den Krieg erklären” (“which declared war on Eve’s daughters who write”) (464). Even though these voices have been piling up, the author of the profile, E. Z., argued that works by women such as E. Werner contributed significantly to literary productions in this century of emancipation: “Zu den erfreulichsten Errungenschaften dieses Jahrhunderts der Emancipation gehört ohne Frage die freiere Stellungnahme der deutschen Frauen zur Literatur” (464).[7] The statement highlights the significant progress made in women’s rights. These progressive advances have resulted in the increased involvement of women in the literary world, which in turn have influenced and triggered innovative styles and movements.

When one looks at how the text was embedded in the Texas newspaper, it seems rather ironic that Lindau’s report, which represented one of the “vielfach laut gewordener Stimmen,” was preceded by a text of a woman who wrote quite successfully for women readers as contemporaries stated (Maack 84). Bürstenbinder contributed a variety of novels and novellas to the literary record, many of which were first published in magazines such as Die Gartenlaube. According to Claudia Stockinger, who studies Die Gartenlaube, the popularity of a periodical cannot be measured by individual writers, works, or genres, but always in relationship to one another (54). Popular magazines not only served as role model for other magazines, but also for German-American newspapers, which by the end of the nineteenth century were flooded with advertising and entertainment. As the example above shows, her works were not only published in a popular magazine but also traveled to the U.S. where they were printed in the daily press of German migrants. As stated in chapter 2, it is nothing new that German-American newspaper editors reused material from Germany (Hofacker 2). To this day, however, scholars mainly highlight texts written by men, such as Goethe or Gerstäcker, thereby neglecting not only the works by American women as in the previous chapter, but particularly the numerous works by German women that were reused freely and printed in German-American newspapers, which shows the strong interest of German migrants to read these works. Such portrayals limit understandings of German-American print culture as they neglect the manifold numbers of texts for, by, and about women that were published before the invention of women’s pages.[8] Recent scholarship in American book and press history emphasizes that the literary canon established in the twentieth century does not reflect the reading preferences of nineteenth-century audiences, who often favored familiar and traditional styles over novel or original content (see Cordell and Stokes). Before diving into women’s literary contributions in the German-American press in the following chapters, the objective of this chapter is to sketch the workflow of computationally classifying genre to emphasize both the opportunities and challenges of operationalizing concepts such as genre and gender.

Text reuse detection, when combined with digitized archives, makes it possible to trace how texts were repurposed by different editors across the country: as news, novels, poetry, or advertisements. Examining not only the reused texts themselves but also their textual surroundings reveals the heterogeneity of content found even on a single newspaper page and highlights the blurred boundaries between fact and fiction, including the embedding of poetry within prose. While text reuse detection proved to be a useful way to extract reprinted texts from a digitized corpus, the generated output (more than 40,000 reuse clusters) was too large to close read all reprinted texts and spot similarities and anomalies between different newspapers issues. To help researchers analyze and make sense of large amounts of cultural heritage data, machine learning (ML) is increasingly being used for digital archives to identify patterns, to study themes, and analyze language use over time (see Jaillant). The objective of this study is to use machine learning to divide the larger dataset into a binary system of two subsets—“fiction and poetry” and “non-fiction”—to systematically examine texts for, by, and about women. By merging different methods, that is unsupervised and supervised machine learning, I was able to expand the workflow by classifying these texts not only into two, but ten different categories: hard news, advertisements, lists, factual texts, announcements, poems, jokes, short stories, novels, and religious texts. This allowed for a more nuanced analysis as other textual forms turned out to be relevant for the present study since information about women was highly present in non-fiction as well.[9] The computational text classification and its output, which will be discussed in this chapter, does not provide a binary judgement but instead discloses a range of possibilities and probabilities to study the hybridity of textual forms by locating multi-classification. This implies that a text can have the likelihood of belonging to several types (Gledhill 65). The goal is not to have a perfectly classified corpus of different genres. Such an endeavor would be impossible given the relationship between different text types published in the German-American press that developed and mutually inspired one another in the nineteenth century. “Genre hybridization and intertextuality are not an exception but a current practice that has actually contributed to the emergence of news genres,” as Langlais points out (204). The goal is instead to use computational methods to understand the similarities and differences of text types by merging qualitative and quantitative analyses. This chapter first provides a brief overview of machine learning applications in computational periodical research before it shows how unsupervised approaches (topic modeling and clustering) and supervised approaches (manual annotations in a topic model browser) were combined in this study. The last section introduces the predictions made by the classifier before turning to frequency analyses which suggest that information about women was disproportionally featured in jokes, ads, and short stories. Together with qualitative analyses, these results lead me to argue that women were not only already addressed in newspapers before 1890 but also played a more significant role in news production than previously assumed. Therefore, the focus of the last two chapters is on the roles of women as editors, reporters, writers, and consumers in the German-American newspaper industry, with particular emphasis on reused content in the areas of advertising and short fiction.

4.1. Current Research on Machine Learning in Computational Periodical Studies

Machine learning applications can be found in almost all institutions that host “big data” including libraries and archives. Since many of these institutions started to digitize material in the 1990s, they have explored diverse approaches to improve discovery in their large-scale collections (see Terras).[10] Users of digital archives such as Chronicling America are often already confronted with machine learning output when they use interfaces since the digitized content, which they examine with keyword search or filtering, was generated into data by means of applications such as OCR.[11] Machine learning can also be used to identify common topics across documents, to tag the content of digital collections, and to group similar materials. Even though libraries experimented with machine learning to make use of such techniques to connect similar items across digitized collections, they have been slower to integrate the results “into their interfaces and catalog systems,” as Cordell highlights (“Machine Learning” 20).[12] “Machine learning,” a term coined by Arthur Lee Samuel in 1959, is a branch of artificial intelligence focused on creating algorithms that enable computer systems to learn and improve from “experience” autonomously, without the need for explicit programming.[13] Learning from experience means that these models learn to recognize patterns and make predictions based on data inputs, which are usually large amounts of training data extracted from the internet.[14] Machine learning can be broadly categorized into three types: unsupervised, supervised, and reinforcement learning. “Supervised methods rely on a typical artificial intelligence process: the model is trained on an annotated corpus (or ‘ground truth’) and evaluated on a test corpus. Unsupervised classification methods (or ‘clustering methods’) produce automated classifications,” as Pierre Langlais pinpoints (201).[15]

Several projects in computational periodical studies have started to experiment with machine learning to analyze the abundance of digitized material that is impossible for the human mind to process. “Automated classification of news is an old practice” (Langlais 200). Since Reuters Ltd., the largest international text and television news agency, started to release over 800,000 manually categorized newswire stories in 2003 for research purposes, scholars have been eager to use this data for machine learning (Lewis 362). The following project examples sketch a brief overview to show how machine learning is used “as an aide to discoverability and serendipity amidst informational abundance” (Cordell, “Machine Learning” 1). Verheul et al. use deep learning, which is a subset of machine learning, to examine how concepts such as national and illness have changed in meaning over time and migrated across countries using a digitized multilingual newspaper collection.[16] Klein examines English-language US-abolitionist newspapers with unsupervised machine learning (topic modeling)and statistical analysis to identify and describe the invisible editorial labor of women and to outline their contributions in discussions about the abolitionist cause and the women’s suffrage movement (“Dimensions of Scale” 24). While these studies predominantly use machine learning and statistical analyses to detect patterns in the data, other studies apply techniques such as text classification to automatically allocate texts into predefined categories or classes based on its content. Langlais develops a supervised classification model to automatically categorize newspaper genres of French-speaking newspapers from 1800 to 1945 into, for instance, rubrics such as columns, economics, sports, or theatre and to study how newspaper genres evolved over time (198). In contrast to Langlais, Fitzgerald and Cordell combine unsupervised and supervised machine learning (using English-language texts from Chronicling America) to classify newspaper texts into the four categories: literary, news, informational, and ads.[17] While these studies follow different research objectives, they show the usefulness of machine learning approaches to cope with OCR “noise” and to study large digitized historical newspapers at scale.

Scholars can use supervised classification approaches to sort a corpus of texts by theme or structure, or unsupervised machine learning to group texts by shared topics or linguistic structures. Studies like the ones described above are not only relevant to gain a deeper insight into the historic newspaper landscape, but they constitute useful resources to make machine learning more accessible for a non-technical audience. “To prepare students for a world where information is filtered by computers,” as Ted Underwood writes, “we will need a stronger alliance between the humanities and math” by combining “cultural criticism of the mathematical models shaping our world, and mathematical inquiry about culture” (“Machine Learning”). In line with Underwood’s vision, the goal of integrating computational approaches into the Humanities, in other words the productive integration of the algorithmic to the humanist, is to learn from the process of classification, to use it as a tool, and to regard the generated output not as a final argument but a productive ground for exploration (see Klein et al. 132). The objective of machine learning is to explore the data. We can automatically uncover previously unknown or hidden patterns and relationships within the sources, drawing attention to documents or objects that might otherwise have been overlooked in research (Althage 258). Such a perspective also helps to examine black boxes (see Schwandt), to demystify narratives about AI (see Meinecke and Voss), and to reflect on who gets counted and classified in technical systems (see D’Ignazio and Klein, or Noble). The objective of this chapter is to provide a deeper understanding of machine learning for historical research by guiding through the process of working with language models to classify newspaper texts into genres and to analyze gender in German migrant newspapers by systematically identifying texts for, by, and about women. The classification workflow combines unsupervised and supervised approaches for computational text classification. It first uses topics, derived from a topic model of the reprinted texts, which were subsequently enriched with hand-coded data by tagging a sample of texts according to textual forms to train a classifier. By describing the development of this model in the following sections, the aim is to not only discuss the generated output, that is the types predicted by the classifier in the corpus, but to formulate the human-computer interaction that was designed to analyze which textual forms were printed and reused in the German-American press.

4.2. An Unsupervised Approach: Texts and Topics

“The concept of genre is as old as literary theory itself, but centuries of debate haven’t produced much consensus on the topic,” as Underwood concludes (“Life Cycles” 1).[18] Contemporary practices of statistical modeling put “different groups of texts into dialogue with each other” (ibid.). Underwood is one of various scholars who started to experiment with machine learning for genre analysis seeing the potential of working with probabilities to examine categories such as genres because they are not fixed or immutable, but rather dynamic and constantly evolving (24). Moreover, machine learning can handle multiple categories simultaneously. The present study incorporated the works of various researchers such as Underwood, Hettinger et al., and Schöch (“Topic Modeling”) who use computational methods for text classification by using digitized books as well as Langlais and Fitzgerald and Cordell who work with digitized newspapers.[19] In contrast to scholars working with digitized books, which can make use of classifications assigned by previous scholars and library cataloguing systems, the ones working with digitized newspapers have to start from scratch as the majority of texts in newspapers have not been tagged according to text type in the process of digitization. To enrich these texts with metadata (genre), the classification process starts with unsupervised machine learning by employing one set of such methods, topic modeling, to examine the usefulness of topics and to identify more specific indicators of thematic content and how they correlate to other textual forms.

Topic modeling is a statistical machine learning technique used to identify and extract the underlying topics from a large corpus of text documents. The goal of topic modeling is to uncover the hidden semantic structure in the text data, which can be used to gain insights into the underlying themes and patterns of the text (Klein, “Dimensions of Scale” 27). “As an information-retrieval tool, topic models excel at organizing documents into clusters of texts based on shared clusters of terms that themselves occur together with high probability” (Algee-Hewitt). The most used algorithm for topic models to capture the semantic relatedness between documents and terms is called Latent Dirichlet Allocation (LDA). Introduced by Blei et al. in 2003, it was the first complete generative topic model. This approach assumes that each document is a mixture of a small number of topics. “A topic is considered as a list of words with probabilities indicating their belonging to a certain” group (996). Every word is associated with each topic at varying probabilities, meaning that each term follows its own distribution across topics. The algorithm iteratively estimates the topic distribution of each document and the word distribution of each topic until it converges on a stable solution.[20] For this study, a topic model was created by using MALLET to divide 44,750 clusters of reprinted texts into 300 “topics.”[21] Once the algorithm converged, the topics and their associated words were extracted from the model and interpreted using an interactive visualization as shown in figure 10.[22]

Figure 10: Interactive visualization with an intertopic distance map, representing topic similarity in a two-dimensional bubble plot derived from word distributions.

Interactive interface with a bubble graph on the left showing topic distances by similarity, and a right panel listing the top 30 terms for topic 2.

“[F]requently co-occurring words are presumed to form tighter or more coherent semantic groups than less frequently co-occurring words,” as Anna Shadrova explains (10). The concept of words working together to construct a topic is represented by spatial proximity, as one can see on the left side in figure 10.[23] Each bubble stands for a distinct topic. Larger bubbles generally indicate topics that are more prevalent or significant in the analyzed dataset. “This fact somewhat qualifies the general assumption that the topics with the highest scores in a given” dataset represent the major themes, as Schöch suggests (“Topic Modeling” 13).

The objective of the present study is to examine which dominant themes can be identified through topic modeling and if topics correspond to core vocabularies of text types that can then be used to build a classifier. The topics produced by topic modeling techniques are clusters of similar words. To examine the degree to which topics reflect different textual forms, Underwood uses “words as clues,” but this does not imply that text categories are “a linguistic phenomenon” (“Life Cycles” 6). This suggests that words, or rather, cluster of words, are a convenient tool for tracing the implicit similarities and differences between distinct practices of text classification. Even through different textual forms respond to “social phenomenon; words just happen to be convenient predictive clues, allowing us to trace the implicit similarities and dissimilarities between different practices of selection” (“Life Cycles” 6). A text typically concerns multiple topics in different proportions. A topic model mathematically formalizes this intuition, enabling the analysis of a collection of texts to identify the underlying topics based on the statistical distribution of words within them. Andrew Piper describes a topic as “a heterogeneity of statements under a semantic field,” in other words a topic can encompass several meanings within its distribution (74). When encountering topic modeling for the first time, one observes that the topics generated are not conventional topics, as one may typically expect. These topics do not directly reveal the subject matter of texts. Klein therefore describes topic modeling rather as “a technique that stirs the archive” (“The Carework”). “When a topic model sorts a corpus into an unexpected arrangement, then, it can defamiliarize our understanding of its organization and surface new connections that are equally as meaningful for that corpus as any of the critical schemas that we may seek to impose on it” (Algee-Hewitt). One assumes that if a text focuses on a specific topic, certain words would appear more or less frequently within it. Topics containing a high frequency of terms such as “Korrespondent” (“correspondent”), “Mittheilung” (“message”), “Bericht” (“report”) as shown in figure 10 are more likely to refer to texts that can be computationally classified as news. Other topics that include a high representation of nation and place names such as “Deutschland” (“Germany”), “Frankreich” (“France”), “Amerika” (“U.S.”), London, or Frankfurt are useful for a classification of news as well. Using the intertopic distance map and exploring different topics revealed that there are many topics that closely correlate with other types of non-fiction. Topics that contain terms such as “Schwäche” (“weakness”), “Energie” (“energy”), “Fieber” (“fever”), and “Geheimnis” (“secret”) point to advertisements, specifically for patent medicines, which dominated the nineteenth-century newspaper ad market (see chapter 5). Topics, however, containing terms such as “Diener” (“servant”), “Herzen” (“heart”), “Liebe” (“love”), “Augenblick” (“moment”), or “treuer” (“loyal” or “faithful”) point toward fiction. Data exploration shows that many terms such as “Schwester” (“sister”), “Mädchen” (“girl”), or “Liebste” (feminine version of beloved) are listed as some of the most frequent words in the topics suggestive of fictional texts. In most topics, these terms have a higher frequency of occurrence than their counterparts such as brother or boy. While some men’s names such as Johann, Jakob, John, can be found, no specific women’s names feature in these topic lists.

In topic modeling, term lists serve as heuristics to help researchers label clusters not as analytical endpoints themselves (Algee-Hewitt). This first data exploration suggests several striking patterns: first, there are many topics that imply a strong correlation to a specific text type. Second, topics that hint at belonging to a certain type also appear in close proximity to each other. Third, there are differences in the representation and frequency of gendered terms in distinct text categories. While suggestive, these patterns of similarity and differences of topics require a more finely grained analysis. Moreover, other scholars may find other patterns useful for text classification simply due to the reason that “we still lack systematic theoretical categories for classifying periodicals,” as Matthew Philpotts points out (404). Klein argues that “[w]hen considering its uses within an expanded conceptual frame,” one “might find yet more meaningful applications of topic modeling” (“Dimensions of Scale” 27). However, despite their thematic suggestions, topic models are just models. Their significance depends on domain experts critically examining the proposed associations and contextualizing them within the archive (Klein, “The Carework”). In the present study, such an expanded framework implies that topics are integrated into the training set for an automated text classification process of newspaper texts. An automated text classification process uses algorithms and machine learning models to categorize or label large amounts of text data into predefined categories or classes. It involves training a model on labeled examples so that it can learn the patterns and characteristics of different types of texts. Once trained, the model can automatically assign new, unseen texts to the appropriate category based on their content. To create training data for the classifier, I manually annotated topics and their corresponding texts, a process described in the next section. This step in the workflow raised a number of new questions: How easy is it for a researcher to label texts as news or advertisements? Can one distinguish between text types without seeing their placement on the newspaper page? Is it more difficult to tell jokes from poems? Which categories are most difficult to identify (for the human and the computer)? This process proved invaluable, offering me an entirely new way of engaging with the digital newspaper archive. It brought together texts that, based solely on their vocabulary, had been clustered in unexpected and revealing ways.

4.3. A Supervised Approach: Scalable Readings of Topics and Texts

The main template for applying machine learning to classification in historical newspapers is supervised machine learning, even though “large pretrained language models and image representations have improved many tasks in natural language processing” (NLP) and computer vision (Langlais 201). Current work on supervised machine learning covers several tasks including layout analysis and page segmentation (see Wevers and Smits), which determine article breaks, figures, and reading order, or entity linking, which seeks to connect mentions of named entities or definite descriptions to entries in a knowledge base (see Ehrmann et al.). Text classification is an additional approach that allows researchers to group text units (see Aysenur et al.). While these tasks have different objectives, they share a canonical workflow that starts with creating a labeled dataset. This dataset is later “partitioned into a training set, validation set, and test set. A machine learning classifier uses the training set, including classification labels, to learn a classification function” (B. Lee 523). For a text classification task, there are several ways of creating a training set. One can create data in the form of a list of terms to describe a category (see Kantner and Overbeck) or hand-coded data by annotating items such as pages, lines, articles, or names and tag them with belonging to a certain category (see Fitzgerald and Cordell; Langlais). Either the list or specific items are then used as training data.

Another way, which is the objective of the present study, is to use topic-based annotations for the training of a supervised classifier. In this case, topics derived from topic modeling are used for training in conjunction with hand-coded data, that is tagging texts into “fiction” (including poetry) or “non-fiction.” In the beginning of this approach, the objective was to automatically determine whether a text can be classified as the one or the other using support vector machines (SVMs).[24] Merging several approaches, however, revealed that a classifier can be developed that not only distinguishes between these two labels but instead allows a more fine-grained analysis to work with several categories from news, advertising, information, and entertainment. The exploration and annotation platform used for this approach by linking different data from the topic models, the dataset of reprinted texts, and *Chronicling America’*s interface, to label data was a key factor for this realization. The following section outlines this iterative classification process that involves annotation, training, and generated output. Any scholar working on supervised machine learning says that most of the time has to be spent on annotating data as a model can only learn from what the program has previously seen. The decisions of the user about what to label and how has a strong impact on the ability of the classifier to learn. “[T]he process of annotating training data is often the most costly element of an ML project” but “researchers need ground truth and training datasets compiled specifically for cultural heritage work, curated by domain experts” (Cordell, “Machine Learning” 36–37). A major challenge in developing tools for intertextual analysis is the absence of gold standards or “ground truth” datasets that aim for comprehensive coverage (Gruber 274). Yet no such standard exists that fully captures the characteristics of a specific nineteenth-century newspaper genre or distinguishes it reliably from others. As a result, I created my own training dataset, shaped by both statistical outputs and human interpretation. This approach shows that computational text classification is as much a matter of cultural and historical context as it is of computation.

Training data is key to the performance and usability of a machine learning application. To interrogate the topics together with a selection of texts from the reprint clusters and create a training dataset through hand-coded data, the dfr-browser developed by Andrew Goldstone was used, but extended with an annotation option.[25] “This browser seeks to reflect the multifaceted nature of a topic model (it is a model of patterns in both words and documents) by offering you multiple views of the model” (Goldstone and Underwood, Quiet Transformations).[26] One can interactively explore the topics by using different views to access both qualitative and quantitative data. There are four different versions of the overview of the topics that are represented by their most frequent words: the grid, scaled, list, and stacked subview. These views provide a variety of ways of exploring the newspaper data from dismantling visualizations reading word lists or texts. Starting with figure 11, the list view shows how the topics are listed in table form. Here, one can examine the keywords (top words) and the distribution of the individual topics across the corpus. The blue bar on the right, for example, represents the percentage of all words in the corpus that are attributed to the specific topic. Just looking at the topics 1–16 in figure 11, there is nothing conspicuous to observe in relation to the proportion distribution except for topic no. 9.

Figure 11: List subview (overview) of the topic model browser, presenting the first 16 topics alongside distribution visualizations and top-word lists.

List subview of a topic model browser showing the first 16 topics, each with a small distribution visualization indicating corpus proportions ranging from 0.1 to 3.9 percent and a list of top-ranked words.

Goldstone, who, however, did not create the browser for text classification purposes, argues that these high-proportion topics “are often the least interesting parts of the model—agglomerations of very common words without a clear thematic content” (Goldstone and Underwood, Quiet Transformations). However, in the classification process for this study, this was not necessarily the case. Topic no. 9, for instance, with a high proportion (3.9%) compared to the other topics including top words such as “herr” (sir), “gut” (good), “na” (an interjection used in expressions of anger, impatience, hesitating approval, disbelief, or astonishment), “Lieber” (masculine version of dear), “fräulein” (reference to an unmarried woman), or “bitte” (please) lacks a clear thematic content but turns out to be a relevant indicator for texts that correlate with fiction, more precisely with jokes. Such features get weighted more heavily in the machine’s language model because they significantly help its ability to distinguish or classify one type of text from another.

Other views and their integrated visualizations in the browser provide the opportunity to analyze topics across the time frame. A visualization that turns out to be useless for one question can be potentially helpful for another. Drucker, who advocates for a critical and creative approach to visualizations, argues that visualizations can be used not only to represent data, but also to explore and communicate complex ideas and relationships (Digital Humanities, 86–110). She suggests that visualizations can be designed to provoke thought and discussion and can serve as a tool for inquiry and exploration.[27] In the list subview, for instance, the column to the left of the topic keywords (top words), one can see a miniature graph of the topic time-series, which provides a general idea of how the topic is distributed across different document years. Topic no. 9 is the only one in figure 11 that shows a steady increase throughout the century, even though there is a high peak at the beginning of the time frame. To examine if topic no. 9 is the only prominent topic or if jokes become more common throughout the century, one can go a step further into the stacked subview (overview) as shown in figure 12. This view provides a streamgraph of the entire corpus to understand how the composition of topics within the corpus changes over time. In figure 12, the height of each topic’s stream gives the proportion of the topic in each time interval, and colors are used to distinguish topics from one another. This view allows “a chance to pick out islands of prominence or topics with unusual spreading over the time space,” as Goldstone explains (Goldstone and Underwood, Quiet Transformations). Given the number of topics and colors, this view may be overwhelming at first. Several observations, however, can be made through examining this visualization that account for the complexity of the following phenomena—generating new questions. First, one can clearly observe a shift between 1862–1865, which correlates with the American Civil War. Starting around the period of the war, the distribution of topics changes and it seems as if suddenly more different topics emerge. In the period before 1865, the spaces of the individual topics seem much bigger.

Figure 12: Streamgraph in the stacked overview, encoding topic proportions over time through stream height and color differentiation.

Streamgraph displayed in a stacked overview, with differently colored streams whose varying heights indicate the relative proportions of topics across successive time intervals.

Around the 1860s, however, the lines and spaces become smaller, which points towards newspaper content becoming more complex. This shift can be explained both by the dataset that includes less newspapers before 1860 (see chapter 2) and by the likeliness of having fewer topics in this period in general because less newspapers were published.[28] In the 1860s, 200 German-language newspapers and magazines were published in the U.S., which increased to 1,000 in 1880 (see Arndt and Olson).[29] Prior to this, there was less variation in both editors and readers. Moreover, papers contained fewer pages, usually around four, which increased during the century from eight to twelve. Starting in the 1860s, the number of newspapers rose due to mass migration movements in the 1850s which becomes visible in this visualization at scale.[30] The diversity of readers thus influenced the heterogeneity of content. To meet the tastes and needs of these German readers with various linguistic, political, economic, religious, and social backgrounds, the editors had to provide information on a range of topics from politics on the local, regional, national, and international level to advice on health and household. This visualization does not necessarily produce new insights but makes it possible to quantitatively display how war, migration, and technological innovation increasingly shaped newspaper content.

4.4. Tagging Topics and Texts: (Machine) Learning from Annotation and Classification

This topic model browser allowed me to examine the German-language newspapers from different perspectives to spot anomalies at scale. In the streamgraph in figure 12, one can explore dominant topics in this period, even though their number is low compared to pre-1865, such as topic no. 9 (red). When inspecting the parameters of the most prominent latent distribution, one can observe in the streamgraph that 10% of the topics (36 of 300) have a higher proportion in the corpus as indicated by their listing. Similar to topic no. 9, the majority of the 36 topics also serve as useful indicators for classification such as, for instance, topic 184 (blood, stomach, liver) that correlates with ads, topic 20 (door, saw, stood) with short stories, topic 30 (heart, God, world) with poems, and topic 48 (German, government, party) and 264 (foot, meter, inch) with news.[31] To read texts that correspond to these topics with varying degrees, one can switch to the topic view as shown in figure 13. Here, on the left side, one can read through an extended version of the top words list, while on the top of the right side, one can see a larger version of the histogram that visualizes the distribution of the topic to identify patterns and trends. Below this graph is a sample list of documents. The topics are derived from documents, i.e. from texts in the text reuse clusters. The tokens in these documents relate to this topic with a high likelihood (100%) and a slightly lower one (93.3%). What one can also see in this document list is that the two top items and the two last ones are accompanied by blue labels titled “jokes.” These documents were tagged by me as “joke.” As stated earlier, this topic model browser was extended with a tagging option to be used as an annotation platform. In this way, data could be represented in different ways and examined from distinct perspectives to iteratively identify text types and enrich texts with annotations.

Figure 13: Topic view for topic 9, showing the top-ranked words, the topic’s distribution across the corpus, and the associated document list.

Topic model interface showing topic 9 with top-ranked words, a distribution graph, a document list, and four documents tagged as jokes.

Tagging texts allows researchers to capture nuanced themes, relationships, and patterns that might not be immediately visible through close reading methods. By systematically applying tags, the corpus can be explored more deeply, enabling sophisticated queries and facilitating the discovery of underlying trends, connections, and variations across different text types or categories. As this subchapter shows, the use of computational methods does not imply neglecting close reading. As Langlais emphasizes: “Model categories should be informed by historical research and by a thorough close reading of primary sources” (203). Textual forms differ with respect to their level of generality. In the case of novels, for instance, one differentiates between different levels such as literature (superordinate), novel (basic), and detective novel (subordinate) (see Underwood, Distant Horizons 48). German-American newspapers brought together texts from many different newspaper and literary genres and levels. When humans classify a text according to a Textsorte (text type), they typically consider higher-level aspects such as the overall feeling or tone of the text. In the present study, to tag texts with a label, one can click on a document that redirects one to the document view as shown in figure 14. Here, one can read the text titled “Gute Freundinnen” (good female friends), a story about gossiping housewives, and access the metadata of the corresponding text reuse cluster. In some cases, this information gained through reading the text was enough to be able to identify the text type.

Computers rely on counting what linguists call “tokens,” which includes individual characters, words, phrases, or sentences. Among these, words are the most frequently used tokens for computational classification experiments because they are easy to isolate within texts and can also be interpreted in meaningful ways by humans. In contrast to a detective novel, for instance, where terms such as “police,” “murder,” “investigation,” and “crime” define the thematic premise of the genre (48), the category of “joke” in newspapers cannot be computationally classified by “puns” and “surprise.” Texts like “Gute Freundinnen” were annotated as jokes because in comparison to other fictional texts, they were extremely short, and they often illustrated a narrative in the form of a dialogue ending in a punch line. From a computational linguistic perspective, these texts were noticeable because they consisted heavily of personal pronouns, interjections such as “Ach” or “Na” as well as quotation, exclamation, and question marks. Reading through other examples in topic no. 9 reveals that what they share is not that they all focus thematically on girls and gossip, but what brings them together as a unit is their text type. Clusters are not predefined but are derived by the algorithm from the data itself during the analysis. In contrast to texts in topic no. 9, most texts are examples of soft clustering, that is they correlate with more than one topic such as, for instance, a text in topic 44, which was annotated as news. Titled “Männlich und Weiblich” (Male and Female), the text was a report that openly declared that it was inspired by a story from the British press that described a study about the differences between the two “Geschlechter” (genders) and commented that the comparison, even though it was true, painted a negative image of women. This text that correlates with seven other topics (including 10.2% with no. 9) was printed in 1894 in Texas, Maryland, and Nebraska (see Introduction). Apart from multi-topic distribution of texts, these examples illustrate that information about gender could be found in all different kinds of texts, which circulated across states and decades.

Figure 14: Document view of the annotation interface, where texts are enriched with labels through manual tagging.

Document view in an annotation interface showing a text tagged as a joke, alongside a label panel and references to two newspapers that reprinted it, with links to Chronicling America.

The objective here is to determine what these topics reveal about the degree to which they refer to different textual categories and if these features seem reasonable for a supervised machine learning approach, in which they are used to train a classifier that automatically classifies unknown texts into distinct categories. When one is not able to see a text’s title or its embeddedness in the layout of the newspaper page, and only focuses on its textual component, one can observe that many texts share characteristics with diverse forms such as the increased dialogue in jokes, short stories, and novels. In some cases, though they remained few in total, the top words lists, and the texts themselves, were not enough material to evaluate a text. Thematically, this was specifically the case when having to distinguish between fictional texts such as short stories or serialized novels, but also in the case of news and information. While figure 14 shows the complete text of the joke, in some instances the text reuse detection did not extract the entire short story or excerpt of a serialized novel (mainly due to OCR errors). If the type could not be identified immediately, the corresponding link to get redirected to Chronicling America was used to examine the page. Here, neighboring articles, the number of the page, or preceding phrases such as “5. Kapitel” (Chapter 5) served as indicators for evaluating the textual form. This allows one to combine “practices with longstanding histories like editing, annotation, analysis, teaching, translation, contextualization, commentary” in a digital setting and to link different datasets to read texts up close and at scale (Flanders 9). This step enables one to identify and name which textual forms were actually reprinted in the German-American press. One of the key features of the data exploration in the topic model browser, but one whose impact was not immediately clear in the beginning, was that texts can even be annotated on the basic level instead of the superordinate level of “fiction” and “news.”

To be usable in a digital environment, as Hiltmann reminds us, “information and its relationships must be formalized always involving selection and abstraction, which in turn shape how we understand and interpret that information” (“Hermeneutik in Zeiten der KI” 207). “The range of genres,” however, “cannot be defined in advance and will generally be partly informed by the actual categorization” (Langlais 204). When scholars refer to newspaper content items, they often distinguish between articles, advertisements, images, obituaries, and similar categories (Düring et al. 2). However, these classifications do not necessarily reflect the textual reality of the nineteenth-century American press especially in the context of reused texts. As a result, I developed a different set of categories that better capture the types of texts most commonly encountered and are also more amenable to computational analysis for identifying similarities and differences. In contrast to Fitzgerald and Cordell who distinguish between the four categories of literary (poetry or prose such as sermons, sketches, vignettes, or essays), news, advertisement, and informational, I created the following annotation guidelines for the present study and to tag texts accordingly. As demonstrated, these descriptions are deliberately not overly prescriptive; they allow for interpretive flexibility in categorizing texts while simultaneously striving to delineate boundaries between distinct genres.

Table 2: Annotation guidelines to classify texts in the German-language press.

Text TypesDescription
NewsUp-to-the-minute news and events that are reported immediately.
Advertisement (ad)A message printed in the newspaper in space paid for by the advertiser.
ListA collection of items or elements such as numbers, letters, words, or objects that are arranged in a particular order or sequence.
NotificationA public or formal statement that provides information about a particular event, situation, or decision. They may include details such as the date, time, and location of an event, changes to policies or procedures, or updates about a particular situation.
Religious textA text that includes teachings, stories, and beliefs that guide the faith and practice of its adherents.
Factual textA text to inform and educate readers about a particular topic similar to a textbook, research paper, technical manual, or instructional guide.
Short StoryA fictional “short” story about an incident or situation that shows rather than tells. It may be as simple as an explanation or example or as involved as the retelling of an incident complete with a sense of plot and dialogue.
NovelA serialized work of fiction that typically tells a story about characters and events in a vivid and engaging way.
JokeVery short text of entertainment that is meant to be humorous and often involves a play on words or a humorous twist.
PoemA type of literature that uses highly structured and stylized language, including meter, rhyme, and other poetic devices.

The category of “news,” for instance, encompasses primarily hard news. However, it was not always immediately apparent whether a given text constituted a factual news item or a fictional narrative. In some instances, purported news stories were entirely fabricated: something that was not evident at the level of the newspaper page and required further contextual analysis to uncover. “When a topic model sorts a corpus into an unexpected arrangement, then, it can defamiliarize our understanding of its organization and surface new connections that are equally as meaningful for that corpus as any of the critical schemas that we may seek to impose on it” (Algee-Hewitt). I found numerous connections not only between news and fiction (see chapter 6), but also between informational texts and promotional literature (see chapter 5). In the case of “lists,” for example, I frequently encountered significant repetition of terms within a single text (as discussed in chapter 3), alongside other texts that provided information on commodity prices such as cattle, vegetables, and fruit, as well as announcements of market arrivals, shipping news, and demographic data such as births, deaths, and marriages. A list did not necessarily have to follow a clearly defined format marked by numbers or hyphens. Instead, texts were classified as lists when they enumerated similar objects ranging from names to items or values in a structured manner. “Listicles,” by contrast, often appeared in the form of how-to guides or instructional descriptions and were more frequently embedded within other genres, particularly advertisements (see chapter 5). I found few train schedules in the dataset, which I assume is due to noisy OCR, i.e. these texts often contain many numbers and follow rigid formatting, making them particularly difficult for OCR systems to detect accurately. Subsequently, they were not necessarily detected through text reuse detection. Similarly, advertisements with distinct layouts have also posed significant challenges for OCR recognition. Such layouts, combined with non-standard fonts and spacing, frequently result in misrecognition or omission (Blome 19).

To create the training set for the classifier, I manually annotated 3.35% of the dataset in the topic model browser, which are 1,502 of 44,750 text reuse clusters. The objective was to use these topic-based hand-coded texts to train the model of the ten categories listed above. “The success of machine learning depends both on gathering data and on condensing it, but the second, subtractive step is the part statisticians call ‘learning’” (Underwood, “Machine Learning”). In supervised learning, data was compiled from these instances, which were labeled into multiple classes, where one class was a text type. The goal of a function called classifier was to predict the “correct” label for a previously unseen instance.[32] As described in chapter 2, the analyses depend on words produced by OCR systems. Therefore, starting with topics has several advantages. First, one uses the properties of the entire corpus for classification, i.e. the topics are used to derive a topic distribution for each text (text reuse cluster). In other words, one calculates how strongly each topic is associated with each reprint cluster. Second, using topics has the benefit of not relying so heavily on these words as the texts are often more readable at the phrase level—rather than word-level. Third, the probability of automatically classifying news as news, or ads as ads, is now higher even though there is an uneven distribution of tagged categories in the training set, as one can see in table 3, simply because, for instance, more ads were found and tagged in comparison to poems. One assumes that the topics are relatively “clean,” which means that there is only one topic per text type such as in the example of topic no. 9. Subsequently, if one expects that five topics about news were found with 10,000 news articles and one poem topic with 50 poems, a poem could be found by only looking at six topics. With a random selection, the chance to find a poem correlates with 1/200 with the above distribution of 50/10,000. This means that one needs to look at 200 texts. However, the topics are not that clean and, therefore, there are only a few topics that correlate simply with one category. For this reason, several texts were annotated per topic.

Considering the morphological structure is also important when working with “messy” OCR-ed data such as digitized newspapers. Even though models based on neural networks have become widely used for text classification (see Conneau et al.), “[m]ost of these techniques represent each word of the vocabulary by a distinct vector, without parameter sharing” (Bojanowski et al. 136). Thereby, they ignore “the internal structure of words, which is an important limitation for morphologically rich languages, such as Turkish or Finnish,” but also when working with OCR noise. Bojanowski et al. “propose to learn representations for character n-grams, and to represent words as the sum of the n-gram vectors” (137). To include morphology, a fastText word-embedding model was trained on the data in the present study for two reasons.[33] First, the embedding model became familiar with the data and its domain. Second, the fastText embedding was useful because it does not only rely on words, but also on “subwords” (character n-grams), which allows word parts to be represented as vectors. This is very advantageous in this context with OCR errors, as even “broken” words (“Franen” instead of “Frauen”) were correctly interpreted. Unlike traditional embeddings or bag-of-words approaches, the classification algorithm was able to work with “Franen” (otherwise, it would simply be an unknown word without meaning).[34]

Subsequently, the machine learning algorithm was fed with this training data that consists of pairs of feature sets (vectors for each text example) and tags (e.g., ads or news) to produce a classification model (see Joulin et al. 427–431). After training on a sufficient number of samples, the model started making predictions: the probability of A, given B, equals the probability of B, given A, multiplied by the probability of A, divided by the probability of B. This implies that any vector representing a text contained data about the likelihood of specific words appearing in texts of a particular category, enabling the algorithm to calculate the probability that a text belonged to that category. If boundaries between textual forms are at root fuzzy, “computational methods can model that uncertainty by assigning probabilities to multiple classified genres” (Fitzgerald and Cordell). For concreteness, a specific example can anchor the discussion of probabilities (degrees of belonging to a particular type). For the learning process, one must understand that a text always consisted of components from different topics, and the classifier learned from what it has seen. A text was labeled as a poem, for instance, if it correlated at least 30% with topic 10, 10% with topic 56, or 5% with topic 70. But there were also topics that were a big indicator for a specific category, in which the learned classification rule was, for instance, joke, if the text consists of at least 40% of topic no. 9. These features based on topics in combination with SVM classification worked best (see Joachims 137–142). After the automatic measurement was established, it was used on the application corpus. In other words, each text was given to the machine one by one, and subsequently it was asked to predict which category the text belonged to, based on its language model that it computed earlier by looking at all the texts.

As Long and So argue, humanistic computational work “will require a method of reading that oscillates or pivots between human and machine interpretation, each providing feedback to the other in the critic’s effort to extract meaning from texts” by bringing “together close reading, cultural history, and machine learning so that they supplement one another” (266–267). Table 3 shows the number of reprint clusters that were annotated and the generated output by the classifier. It is not surprising that most texts were classified as news, ads, or short stories. News had (and still have) a strong transregional character which makes them a useful textual form to be widely shared. Similar to news, but disguised as fiction, short stories must be told in a short and precise manner that appeals to the masses, which made them a useful object to be shared widely. Ads were likely to be printed often in the same newspaper for a specific time frame or if they were nationally branded also in other newspapers (see chapter 5). In addition, many jokes, lists, and notifications were found, but significantly fewer poems.

Table 3: This list shows the number of annotated reprint clusters per category and the prediction made by the classifier for the entire corpus.

Annotated Reuse ClustersPrediction
Total1,502 (3.25%)44,750
News60222,068
Ad4459,142
Short Story1622,818
Factual Text111583
Joke56362
Novel45173
Poem2139
List32125
Fiction1244
Religious Text105
Notification1358

The experiments outlined in this chapter show the overlapping, probabilistic outcomes which underlie classification algorithms. “Here a given text is 83% likely to be poetry and 81% likely to be literary and 52% likely to be an advertisement. Even the most generous humanists might look at such a result and believe the classifier has failed, but we argue this is a deeply limiting response,” as Fitzgerald and Cordell argue. The goal in the present study is not to have a perfectly classified corpus. Since textual forms are social constructs that change over time, it is likely that scholars make different decisions when analyzing and annotating texts. Instead, it is more interesting to find out when texts belong to several categories. These findings emphasize how computational models of text classification can express similar ambiguities and overlaps to those highlighted in common humanistic accounts. As the project progressed, the objective became both to develop an efficient classification method and gain a deeper understanding of how diverse textual forms functioned in nineteenth-century German-language newspapers. As a result of this classification process, the dataset of reused texts was enriched by other metadata, such as ads or short stories. This makes it possible to dig into every category and examine its characteristics to study their similarities and differences at scale and up close.

Computational processes are not more neutral or objective than human decision-making. One cannot guarantee, in advance, that one method is equally suited to all textual forms. If a text and type are particularly hard to pin down to a vocabulary, it might be hard to classify. Poetry, for instance, seems likely to pose special problems since, as shown in the previous chapter, poems were also very often embedded into other textual forms. Given their high adaptability they can be found in various contexts. In addition, typical stylistic features of poems such as repetition of words can also be found in other genres as the following example illustrates:

Complizirt. Warum heirathest Du nicht, Freund? O, Das ist sehr einfach—die, wo ich will, will mich nicht, und die, wo mich will, will ich nicht, und die, wo ich nicht will, will mich, und die, wo mich nicht will, will ich auch nicht, und eine, wo ich will und mich will, gibt’s nicht.[35]

The repetition of “die, wo …” constructs a rhythmic and almost poetic quality, reflecting the cyclical and repetitive nature of the male speaker’s romantic frustrations. Each line builds on the previous one, creating a sense of escalating complexity. The structure mirrors the endless loop the speaker finds themselves in, where every potential romantic scenario ends in a mismatch of desires. In this example, humor stems from the exaggerated and seemingly endless complications of love that are summarized in its title: “complicated.” Yet, beneath the humor lies a poignant commentary on the difficulties of finding mutual affection and the often-paradoxical nature of human relationships. The title almost resembles a relationship status on today’s social media platform. I labeled it as a joke, but according to the classifier it resonates more with a poem being labeled as one with a probability of 0.837019. Looking at different examples of predictions revealed that the majority of poems were poems indeed, but in some cases jokes (like “Complicated”) were also classified as poems. Many of these jokes discuss love and marriage. As this shows, the classifier was not able to distinguish humor from poetry, but instead it brought these texts together based on their shared vocabulary.

This section has presented the preliminary findings of a large-scale classification project involving 44,750 clusters of reused texts from digitized German-American newspapers. While a comprehensive analysis of all text types lies beyond the scope of this study, the focus here is to identify texts that were reprinted, reused, or rewritten (and their corresponding categories) and center women as authors, subjects, or intended audiences. The classification outlined in this chapter should be viewed as a work in progress. A major challenge in developing intertextuality tools is the absence of comprehensive gold standards, also referred to as ground truth, which claim a high degree of completeness (Gruber 274). In the case of intertexts, such standards are inherently incomplete, as every text carries traces of prior texts. For large corpora where intertextual relationships are previously unknown, gold standards remain elusive (275). In the context of text classification in the Digital Humanities, ground truth is not fixed or objective, but interpretive and context-sensitive. Assigning genre labels to historical newspaper texts, as done in this chapter, was shaped by my disciplinary training in History, Literary, and Cultural Studies. There is often no singular correct classification, especially when dealing with ambiguous, layered, or fragmentary sources. Scholars have employed diverse labeling systems to account for the variety of content in historical newspapers reflecting not only individual judgment but also differences in newspaper types, time periods, and historiographical traditions. Many classification systems still in use today derive from twentieth-century understandings of the press.

Ground truth in Digital Humanities is often co-produced through scholarly practice, influenced by research questions, disciplinary norms, and critical source analysis. It may evolve over time and differ across projects. Since this single-author monograph is a result of a dissertation, I did not have the resources to work with a co-tagging team.[36] In other words, the classification results might have differed had multiple people annotated the training dataset or if we had developed alternative annotation guidelines from the outset. However, drawing on my expertise in transnational press history and the feedback I received from (computational) periodical scholars over the years, I am confident that this approach enabled me to engage with the archive in new ways and to surface patterns and features that had previously gone unnoticed. Despite its limitations, the utility of this classification (both the process and its output) will be demonstrated in the final two chapters. They examine how women-centered texts interact with other genres, revealing the German-American newspaper as a dynamic, hybrid medium, where fact and fiction regularly coexisted to inform and entertain. These findings raise important questions: Could nineteenth-century readers clearly distinguish between fabricated and factual content? And might the roots of sensationalism, often attributed to the yellow press at the century’s end, be traced further back, as part of a broader shift shaped by evolving copyright laws and journalistic norms in the second half of the nineteenth century? This period marked not only the emergence of new legal and editorial frameworks, but also the contested formation of “news” before the rise of specialized sections and magazines at the turn of the century.

“A promising feature of supervised models is their transferability: classification is not a one-time job that focuses on a specific corpus and ends with the completion of the task” (Langlais 232). Since these are digital research outcomes, one can share the results of the classifier and the annotation platform with other newspaper projects or scholars in German-American studies to work on texts that cannot be discussed in the present study as the objective is to prioritize texts and types that portrayed discussions about gender. Future research may benefit from refining these patterns, incorporating additional metadata for more precise classification. This approach could also enhance the interpretability of gendered language use across various text types.

4.5. Gender Mining: “Frauen leiden” (“Women Suffer”), “Männer standen” (“Men were Standing”)

The classifier can direct to a particular part of the dataset, but from there one has to explore, continue to read, and consider whether the computational model validates the sense of literary, informational, and commercial forms or whether it challenges it. The aggregate view in figure 15 confirms some existing assumptions such as, for instance, the heterogeneity of content published in the German-American press to inform about the new environment while remembering migrants’ Heimat through news and fiction, the large number of news, ads, and short stories in the dataset, and the difficulty of classifying poems. However, this form of reading at scale does not allow a nuanced analysis of gender (yet). One option for using the topic model browser, without making use of the classifier, to find texts for, by, and about women, is to examine the word view as shown in figure 15 as it allows us to combine topics and keywords. “The word view’s main purpose is to remind us that the topic model tends to divide occurrences of each word among multiple topics,” as Goldstone explains (Goldstone and Underwood, Quiet Transformations). “[S]pace limitations,” he continues, “mean the browser cannot make use of the full information the model provides about which words in each document have been assigned to which topics” (“The word view”). In other words, one can only find terms that are ranked as the top ones even though they would appear in the entire top word list. While the term “Frauen” (women), for instance, can be found in at least two topics, its singular version “Frau” (woman) returns: “There are no topics in which this word is prominent.” In contrast, “Herr” (sir/man) occurs in 30 topics and “Herren” (men) in 22. When zooming into these “Frauen” topics to read the text reuse clusters associated with topic 138, one mainly finds factual texts from the 1890s about how to be thrifty as a family, the true nature of women’s beauty, or advice on wedding planning. While topic 138 predominantly directs to texts about specific women’s topics, topic 215 mainly relates to non-gendered news. “Frauen und Kinder kampieren auf jedem Felde” (“Women and children camp in every field”). This phrase is taken from a text that correlates with topic 215, which in contrast brings to the surface texts from earlier decades, which are news reports about accidents, fires, and explosions. While the texts in topic 138 account for the explosion of women-specific topics at the turn of the century, topic 215 leads to texts that show how women were often used throughout the decades as symbolic narrative elements to trigger shock and empathy in event reporting.

Figure 15: Word view showing the highest-ranking terms for each topic, highlighting the words most strongly associated with individual topics.

Word view of a topic model listing strongly associated words for each topic, showing that topics 138 and 215 prominently feature the term “Frauen.”

“The explicit code system of written language provides a powerful tool for the computational analysis of textual corpora” (Arnold and Tilton i5). Numbers have an impact on what one can find in these corpora. As stated earlier, there are many terms denoting women in the top words list, when using the 30 top words. However, when decreasing the size of these lists, as in the case of the word view here, it seems as if words denoting men are overrepresented in the dataset. These results do not reflect the abundance and complexity of texts for, by, and about women that were encountered in the corpus and later during the annotation process. To systematically find these texts, I explored with frequency analyses to further examine women’s representations in the ten different text types. They were identified by using the classification output of the best candidates, as illustrated in the example of the joke titled “Gute Freundinnen.” By counting the frequency of words like “Frau,” “Weib,” “Dirn” (“woman”), “Fräulein” (“miss”), or “Mädchen” (“girl”) in the corpus, one can identify how frequently terms denoting “women” are mentioned and in what contexts. Additionally, the frequency of these words across different time periods, text categories, or other variables were compared to see if there are any patterns or changes in representation over time. Ads, for instance, account for 13,079,320 of 125,648,463 words in the dataset, which is 10.409%. The term “weiblich” (“female”) returns 3,170 search results compared to “männlich” (“male”) with 392 ones. Findings like these, which result from looking at different words denoting men and women lead me to suggest that information about women was disproportionally featured in ads, jokes, and short stories. In news and lists, no anomalies were found when addressing gender as a collective linguistic endeavor of either women or men. Names of men, nevertheless, were more likely to be represented in all textual forms compared to women’s names. These results do not function as final arguments but instead as interesting directories. To a large degree, they confirm the hypotheses that were made through the annotation process as it seems that the classification helped to examine them at scale and up close.

As the gender mining results outlined above indicate, in the aggregate, gender-specific representations occurred in the German-American press in both fiction and non-fiction, most frequently in jokes, ads, and short stories when focusing on the overrepresentation of women. But “numbers have to be read,” highlights Laura Mandell (5). The computational analysis enables expanding perspectives while bringing things to the surface that are otherwise invisible. “If you want to do research on women, you have to embrace qualitative data. There’s no two ways about it, because the reality of women’s lives is simply not captured in quantitative statistics,” as Valerie Hudson emphasizes (qtd. in D’Ignazio and Klein 170). Frequency analysis alone cannot fully capture the representation of women in a dataset. Absences, as discussed in chapter 3, can also speak volumes about gender. Such analysis should be complemented by close reading and contextual interpretation. Using large-scale classification can be an effective way to identify (un-)usual patterns in textual categories. Once they are identified, one can zoom in and analyze why women are overrepresented in ads, whether more products for or by women were advertised in the newspapers, or whether only products for or by women spread across states (see chapter 5). Additionally, one can investigate why women were highly visible in short stories, what were the topics, themes, and concerns in these stories, and if these short stories were mainly written by women visibly or disguised through pseudonyms by using men’s names (see chapter 6). Gender mining is a computational technique used to reorganize the corpus, enabling the identification of texts that predominantly feature women. This approach is particularly useful for revealing gendered patterns of representation across large datasets, allowing for a more focused analysis of how women are portrayed in different contexts or time periods. By isolating texts based on gendered content, researchers can explore shifts in language, themes, and visibility of women in the corpus. As outlined in the introduction, I understand gender mining not merely as counting women in a dataset, but as an iterative process that integrates data-driven methods with close reading and layout analysis.

Frauen leiden” (women suffer) and “Frauen brauchen” (women need), whereas “Männer standen” (men were standing) and “Männer gingen” (men were walking). These are the most frequent verbs that complement the words “Frauen” and “Männer” using the entire non-classified corpus. “Männer leiden” (men suffer) in contrast is one of the least common combinations. What one can observe is that women were complemented by verbs denoting an emotional or passive state. Men, on the other hand, were accompanied by verbs of action. This insight is significant because while women had a high representation on the pages, it is essential to find out how they were represented in different text types to figure out “what counts as ‘woman’” (Haraway 67). Donna Haraway argues that genre is a way of articulating and managing gendered identities and experiences. According to her, genres work as discursive formations that produce and reinforce gendered norms and expectations. Any gender analysis based on word frequencies paints with a broad brush. However, in combination with close readings they can prompt reconsideration of different visibilities of women and men in different textual forms. These findings formed the basis for the two final chapters. Chapter 5 will demonstrate how advertisements, one of the most frequently reprinted categories in the dataset, drew on techniques from both informational writing and fiction, positioning them as a dynamic, hybrid form in the German-American press that conveyed far more than just objects and prices. At the same time, viewed through the lens of women, these texts represent not only one of the categories with the highest reprint activity (alongside news), but also a major vehicle for spreading misinformation, both about women’s bodies and the advertised products, due to the lack of regulation in ingredient labeling and “the social shame and popular ignorance surrounding sexuality and sexual health” (Robinson 81). Chapter 6 subsequently brings to the fore other reused texts that contributed to transatlantic debates on emancipation. These works highlight the difficulty of distinguishing between news stories and short fiction, and of disentangling fiction from reality—thereby drawing attention to fiction as a valuable historical source for gender history.

Combining genre classification and gender mining has the potential to considerably enhance the findability of these numerous texts neglected by scholars that spread across the states and were read from Scranton to San Antonio. Scalable readings of ads and short stories illustrate that women were already targeted as a readership of the German-language press before the 1890s. Until now, given the abundance of newspapers, such instances have been frequently discovered through sheer luck. Digitized newspapers and advanced search make it possible to examine these observations at scale and up close. In this study, I approach gender as a fluid and socially constructed category that extends beyond a binary system. While the analysis primarily focuses on representations of women, I recognize that many of the texts examined were likely written, selected, or circulated by men, which inevitably shapes the narratives and discourses under scrutiny. This asymmetry is not intended to reinforce a binary framework but instead reflects the historical and archival conditions of the sources. Moreover, the current methodological setup limits the ability to detect more nuanced or non-normative expressions of gender. Future research must therefore work toward developing computational approaches capable of identifying and analyzing queerness and other forms of gender diversity in historical newspaper corpora, moving beyond binary classifications to better reflect the complexity of gendered experience and expression in the past.

Digitized newspapers and advanced search offer new ways to find out more about once popular authors and to examine what masses of women were reading, focusing on non-gendered pages and publications. Since the early 2000s, there has been a widespread “rhetoric of abundance” surrounding digitization, which “has suggested several significant shifts of emphasis in how we think about the creation of collections and of canons,” writes Julia Flanders (5). She further explains that it has become easier to digitize an entire library collection rather than meticulously curating it because the cost of storage has decreased while decision-making remains a time-consuming and expensive process. “The result is,” as she concludes, “that the rare, the lesser-known, the overlooked, the neglected, and the downright excluded are now likely to make their way into digital library collections, even if only by accident” (5). As shown in chapter 2, since digital archives do not provide innovative ways to (re-)search the abundance of newspaper pages, one needs to find novel ways to identify and analyze otherwise unseen patterns of language, form, and narrative. This chapter has shown that automatic text classification in conjunction with statistical analysis are useful to analyze text and gender in the nineteenth-century press. In the present study, unsupervised (topic models, clustering) and supervised (manual annotation) machine learning were combined. The main part of this chapter presented “a modeling strategy grounded in the perspective of cultural history and literary analysis” (Langlais 195), which was then exemplified by the building of a topic-based annotation platform for creating the training corpus for a text classifier. The idea behind such a machine learning model was that without having been explicitly told, it decides based on its training data that a previously unknown text gets classified as belonging to a specific text type with a certain probability. Based on the sample data and feedback, the classifier looked at previously unknown, that is untagged, texts and decided how likely it is that a text belongs to the group of texts in the training dataset that was classified as part of a certain textual category.

This chapter has highlighted the reimagining of classification probabilities as potential tools for analyzing gender, genre, and text reuse (and intertextuality). The second part has shown an exploration of the output data generated by the model through a method of scalable reading. Annotation combined with exploration of quantitative and qualitative data through switching between different datasets proved to be a corrective to initial observations as explained in the previous chapter about the kinds of text types one might expect to find in nineteenth-century German migrant newspapers.

Footnotes
  1. Die Freie Presse für Texas, June 17, 1884, p. 1. The poem literally translates into: “The girl will become a bride / I hit, you hit, he hits / Brides turn into women / We hit, you hit, they hit.”

  2. “The short span of life / I am, you are, he is / In a short time it passes by / We are, you are, they are.”

  3. In the introduction to a collection of essays that discuss celebrity culture in the nineteenth century to provide insights and nuances to the popularity and cultural power of European celebrities in the U.S. and American celebrities in Europe, Finnerty and Rosenquist argue that a transnational literary culture has not been fully examined that predates the twentieth century (2). To fill this gap, the volume provides among other things perspectives on the lives, careers, and activities of various European celebrities who appealed to American audiences as orators, performers, artists, lecturers, and writers. They mainly focus, however, on British and French celebrities.

  4. Horn argues that so far gossip has not been crucially discussed in studies of nineteenth century print culture, “which focus instead on supposedly serious innovations like ‘new journalism,’” as I have shown in the introduction, and popular public figures such as Pulitzer (2). Horn suggests integrating the topic of gossip into the study of periodicals because these sources implement “a form of gossip specifically created for a commercial context, namely a form of gossip that brings personal rapport into public discourse” (3).

  5. Der Nordstern, June 11, 1884, p. 7; Der Westbote, June 5, 1884, p. 3.

  6. Die Gartenlaube (1853–1944), which from 1861 was the first German magazine to reach a circulation of 100,000 copies and grew continuously, finally coming onto the market in 1875 with 382,000 copies (Blaschke, Die Entdeckung 165). For a comprehensive study on Die Gartenlaube and the rise of popular literature in Germany, see Stockinger.

  7. One of the most gratifying achievements of this century of emancipation is undoubtedly the freer attitude of German women to literature.

  8. In 1981, Elke Frederikson argued that scholars in French, English, and American studies started much earlier to outline the role of women writers in the nineteenth century in comparison to German studies (97). She further stated that scholars in German studies have mainly focused on the representation of women, femininity, and emancipation in works by men writers, such as Schiller and Goethe (97). To put women writers into perspective for further research, librarian Elisabeth Friedrichs created a comprehensive study about German women authors of the eighteenth and nineteenth centuries, including information about the libraries and archives that hold these works, bibliographies, and information about the male pseudonyms these women used (vii). To compile this study, she used different sources from letters to magazines.

  9. Additionally, this will make it possible to publish the dataset in a way that scholars can search German immigrant newspapers by using the filtering “genre” as an option, which is not possible when using Chronicling America.

  10. For AI-assisted recognition of handwritten texts, see Transkribus. It is a web-based platform for automated recognition, transcription, and searching of historical documents developed by the READ project (Recognition and Enrichment of Archival Documents). Transkribus uses OCR technology and ML algorithms to transcribe digitized documents, such as manuscripts, letters, and archival materials, into machine-encoded text that can be searched, edited, and analyzed. The platform also includes tools for manual transcription and correction, as well as for training the machine learning models to recognize handwriting and the layout of specific document types.

  11. For years, the Library of Congress has explored the use of different types of machine learning applications for their digital collections. Cordell was asked to write a report for the Library of Congress, which describes the history of the use of these technologies in libraries and explores various approaches. The report highlights the risks and potential advantages of crowdsourcing, improving the searchability of typed and handwritten documents and audiovisual materials, as well as enhancing collection management, conservation, preservation, and creative works. It concludes with suggestions for responsibly integrating machine learning, expanding data access, developing infrastructure, and fostering expertise. See Cordell, “Machine Learning.”

  12. For an overview of the characteristics and functionalities of eleven different digitized newspaper programs as an open access guide to a selection of newspaper databases around the world, see Beals and Bell.

  13. For an introduction to machine learning, see Alpaydin. The textbook covers several topics such as supervised learning, Bayesian decision theory, reinforcement learning, graphical models, and statistical testing.

  14. Bender at al. discuss the consequences of large language models including environmental and financial costs (612–613) and advocate for investing time and costs into the curation and documentation of datasets instead of indiscriminately collecting information from the internet (610). Their work contributes to discussions about the role of large language models that brings together ethical, social, environmental, and economic questions for the use of machine learning and AI in general (see Chun).

  15. Between these poles, there are also semi-supervised approaches to ML which increasingly target human-computer interaction by combining “human expertise with computational serendipity” (Cordell, “Machine Learning” 4).

  16. The article demonstrates how word vector models can be used to track changes in the meaning of concepts over time and across countries by analyzing newspapers (English, Finnish, German, and Swedish newspapers) published between 1840 and 1914. The study defines concepts as key terms or core ideas that have been used historically and, as case studies, examines “nations and national identity” and “illness and health” (3).

  17. While these are examples of textual content classification, there are also projects that work with visual content. Lee et al., for instance, develop a deep learning model to identify visual content such as headlines, photographs, illustrations, maps, comics, editorial cartoons, and advertisements using 16.3 million pages from the Chronicling America corpus. For an interface to search 1.56 million historic newspaper photos (1900–1963), which is called the Newspaper Navigator.

  18. Underwood’s study, which uses digitized books and not newspapers, presents a novel experimental approach for tracing “the codification process of science fiction, detective fiction, and gothic fiction using the probabilities of classification” derived from a vast corpus of numerous novels. It challenges three existing theories of genres: it counterargues that genre is a generational cycle, that genre boundaries “gradually ‘consolidate’ in the early twentieth century,” and that “genre are merely a genealogical thread linking disparate cultural forms” (23).

  19. While Hettinger et al. develop a genre classification model using German novels, Schöch (“Topic Modeling”) experiments with topic modeling on a corpus of French Drama.

  20. In the first step, “preprocessing,” the text data was cleaned and preprocessed to remove noise (stop words, punctuation, and special characters, and to convert the text into a standard format). Removing stopwords (articles, conjunctions, negations, and prepositions) decreases the runtime and the number of iterations needed for the training. Moreover, they are removed to improve the interpretability. Secondly, the preprocessed text was converted into a numerical representation, such as a bag-of-words model, where each document is represented as a vector of word frequencies (vectorization). Subsequently, the algorithm randomly assigned an initial topic distribution for each document (topic initialization). Finally, the algorithm estimated the word distribution of each topic and the topic distribution of each document based on the current assignment of topics (estimation).

  21. The MALLET topic modeling toolkit contains efficient, sampling-based implementations of Latent Dirichlet Allocation, Pachinko Allocation, and Hierarchical LDA. To learn more about this application including slides and tutorials and to access the API, see McCallum.

  22. In an intertopic distance map, each topic is represented by a point, and the distances between the points reflect the degree of similarity or dissimilarity between the topics. Topics that are similar or closely related are located near each other on the map, while topics that are dissimilar or unrelated are located farther apart.

  23. When one clicks on a topic such as 2, a topic list of the top thirty most relevant terms is generated on the right side, detailing their overall frequency (blue) and their estimated frequency in the topic (red).

  24. Support vector machines are a type of supervised machine learning algorithm used for classification and regression analysis. They are based on the concept of finding the hyperplane that best separates the data points into different classes. This hyperplane is chosen in such a way that it maximizes the margin between the classes, that is the distance between the hyperplane and the nearest data points of each class. SVMs are effective in high-dimensional spaces and can handle nonlinear relationships between input features by using kernel functions to transform the data into a higher dimensional space. They are commonly used in text classification, but also in image classification.

  25. See Goldstone and Underwood (Quiet Transformations) for the application and a detailed explanation of the interactive visualization’s different views.

  26. Goldstone developed this browser to accompany an essay that he co-authored with Underwood on a probabilistic topic model of seven literary studies journals to offer an interface for further exploration. They divided 21,367 articles from seven literary studies journals over the 1889–2013 period into 150 topics to study the history of enumeration practices in literary studies. See Goldstone and Underwood, “The Quiet Transformations of Literary Studies”).

  27. Johanna Drucker has written extensively on the subject of visualizations. Her other contributions focus mainly on the theoretical implications of data. She argues that data is not a neutral or objective representation of the world but rather is constructed and interpreted in particular ways by those who collect and analyze it (“Humanities Approaches,” 3). She suggests that data is shaped by cultural, social, and historical contexts, and is always mediated by the tools and methods used to collect and analyze it (7). For a critical reflection on Drucker’s contributions, see Lavin.

  28. For a study on the involvement of German migrants in the Civil War using emigrant letters as source material, see Helbich and Kamphoefner. Their work shows that the letters written by German Americans during the Civil War describe a variety of topics and perspectives including battles, boredom, motives for enlistment and desertion, and their opinions about the war, slavery, and race, thereby challenging the idea that the Civil War erased ethnic divisions and created a melting pot.

  29. To gain information about the U.S., emigrant letters were often read aloud in taverns before open fireplaces or in reading clubs (Hansen 150). According to Jochen Krebber, emigrant letters were considered the most reliable source of information for potential emigrants (21). These writings reported that Americans were in need of strong and eager people to populate and cultivate the land, where there were not only vaster spaces to settle but above all higher wages and personal and economic advancement (Moltmann10). Increasing numbers of German immigrants among other things influenced the rise of German publications in the 1870s (see chapter 5).

  30. In the 1850s and 1870s, two massive immigration waves followed. By the end of the nineteenth century, however, emigrant numbers decreased for several political and economic reasons on both sides of the ocean: to name only a few, the American Civil War (1861–1865), the foundation of the German Reich in 1871, the American Economic Crisis of 1893, and simultaneously the economic recovery in Europe as a result of a successful transition into industrialized nations, which ensured an improvement in employment.

  31. The terms in brackets are the three top words in these topics.

  32. In Natural Language Processing (NLP), machine learning is commonly used to teach the computer to understand natural language. The computer needs to understand the semantics of language. Extracting structure such as syntactic, grammatical, or semantic structure from text helps the computer to understand natural language by learning from the text. The capitalization of a word, for instance, serves as a good feature for predicting nouns to train a classifier to predict POS (Part-of-Speech) tags for the German language.

  33. FastText, which was developed by Facebook AI Research (FAIR), is an open-source, free, lightweight library that performs text classification tasks. It is based on the word embedding technique, which represents words as vectors in a high-dimensional space, allowing the computer to process them as numerical inputs (Bojanowski et al. 135–146).

  34. These are the most similar embeddings to the query “Frauen”: franen 0.660979, männer 0.658383, männern 0.622199, fraueu 0.619326, weiber 0.585595, jungfrauen 0.527914, mütter 0.52627, beifrauen 0.502945, krauen 0.498036, frauenwelche 0.494084.

  35. “Complicated. Why don’t you marry me friend? Oh, that’s very simple—the one that I want doesn’t want me, and the one that wants me, I don’t want her, and the one that I don’t want, wants me, and the one that doesn’t want me, I don’t want her either, and the one that I want and who wants me, does not exist.”

  36. Integrating such annotation work into the classroom could be highly valuable, not only to introduce students to machine learning, but also to train them in digital editorial studies by developing classification categories, labeling data, and exploring the characteristics of historical newspapers through both distant and close reading methods.