Week 6
Computational
Text Analysis

Soci—316

Sakeef M. Karim
Amherst College

SOCIAL RESEARCH

Text as Data (Cont.)—
October 8th

Reminders

Qualtrics Survey Assignment

Qualtrics Survey Assessment Deadline

Your Qualtrics surveys are now due by 8:00 PM on Friday, October 16th.

Reminders

Qualtrics Survey Assignment

Assignment instructions are, of course, online.

Reminders

Your Next Assignment

Annotated Bibliography Assignment

Reminders

Mid-Semester Break

Mid-Semester Break

The mid-semester break is upon us: i.e., we do not have class on Tuesday.

Please, do not show up on Tuesday.

Reminders

Until the Midterm Elections
Election Day — Tuesday, November 3rd

A Class Discussion

The “Woke” Project

On Tuesday, Terrence Chen and Jintae Bae discussed a project that draws on computational text analyses to analyze the relationship between public and personal culture.

How would you summarize the project? Consider the following questions:

Two Orienting Questions
  1. What is the puzzle at the heart of our study?

  2. What is our empirical strategy—and how is it informed by theory?

The “Woke” Project

Data
Methods
Quantities
2025 Survey
Personal CultureN = 2,182
Mainstream Media
Public CultureN = 4,792
Reddit Posts
Public CultureN = 3,042
BERTopic
Inductively identify heterogeneity
Large Language Models
Standardize discourse, find deeper meaning
Clusters of Meaning
What is “woke?”
Semiotic Codes
What does “woke” stand for?

High-Level Overview Positioning Text as Data

A Mainstream Technique

Computational text analysis is in its heyday. The availability of unprecedented volumes of digitized content, the development of sophisticated methods for extracting meaning from such data, and the rapid upgrading of the relevant technical expertise among practitioners—all fueled by the cross-pollination of ideas between computer science, engineering, linguistics, and the social sciences, as well as between academia and industry—have brought this mode of empirical inquiry into the scholarly mainstream.

(Bonikowski and Nelson 2022:1470, EMPHASIS ADDED)

Stylized Example

Topic Models

Choose a passage, then run the model to see its mix of topics.
Topic mixture Share of topic words in the passage

Stylized Example

Word Embeddings

Compare Migration with
0.00
Pick a Word to Compare
Cosine SimilarityA measure of how alike two words are in meaning, with values closer to 1 meaning more alike. Between Word Vectors in HyperspaceThe semantic space (with hundreds, if not thousands, of dimensions) where each word represents a point.

Stylized Example

Using Large Language Models

Prompt Zero-ShotThe LLM sees no labels, only prompts.
Review this snippet and tell me the most likely discursive category it belongs to (IMMIGRATION, ECONOMY, or CLIMATE CHANGE) based on the semantic patterns you detect. {snippet}
TemperatureAt 0, the LLM deterministically produces the most likely class. Above 0, the answer is drawn at random (but weighted), so less likely classes are occasionally chosen.
Predicted Class
No Prediction
–
Probability

Deduction, Induction, and Textual Analysis

Towards an Iterative Model

[R]esearchers often discover new directions, questions, and measures within their quantitative data … If the standard deductive procedure is followed too closely and data is only collected at the very last minute, researchers might miss the opportunity to refine their concepts, develop new theories, and assess new hypotheses. A great deal of learning happens while analyzing the data. Even when a research project starts with a clear question of interest, it frequently ends with a substantially different focus.

(Grimmer, Roberts, and Stewart 2022:14, EMPHASIS ADDED)

Towards an Iterative Model

Regardless of how the research was actually conducted, standard practice of writing articles begins by stating the theory, its observable implications, and the measurement strategy; then the dataset is introduced. This poses a problem for inference if a researcher—even unintentionally—presents a theory as if it is being applied to a fresh dataset when in fact the same data is being used both to develop and to test the theory of interest … Acknowledging that we regularly engage in induction to refine the methods of discovery and then rigorously test these discoveries would greatly improve how we conduct social science.

(Grimmer et al. 2022:14, EMPHASIS ADDED)

Towards an Iterative Model

Deductive and iterative models of research Deductive: theory or model, then hypotheses, data collection and analysis, results. Iterative: theory or model, data collection and data analysis form a repeating loop that then leads to hypotheses, data collection and analysis, and results. Deductive Model Theory / Model Hypotheses Data Collection and Analysis Results Iterative Model Theory / Model Data Analysis Data Collection Hypotheses Data Collection and Analysis Results

Adaptation of Figure 2.1 in Grimmer, Roberts, and Stewart (2022)

Towards an Iterative Model

[T]hinking iteratively is the best approach for analyzing text as data … As in the standard model, we begin with an interesting question, insight, or specific dataset. But rather than suppose that our theories are completely developed before looking at data, we emphasize that iteratively examining data and refining theories help us to clarify our theoretical insights. After this inductive process, the researcher must then obtain new data on which to test the refined theories.

(Grimmer et al. 2022:15, EMPHASIS ADDED)

Towards an Iterative Model

The need for a more inductive model of social science research is not new; nor is it specific to text. However, because of the high informational content and richness of text data, an inductive approach can be helpful at the early stages of a research project—when scholars are formulating their intuitions—as well as at the later stages of the research process … [T]ext as data can contribute to inference at three stages of the research process: discovery, measurement, and inference.

(Grimmer et al. 2022:15, EMPHASIS ADDED)

Group Exercise

Discovery, Measurement, Inference

Grimmer et al. (2022:15–17) argue that treating text as data can facilitate discovery, measurement, and inference throughout the research process.

In small groups, complete the following tasks:

Review and discuss what Grimmer et al. (2022) mean by discovery, measurement, and inference.

Describe how you could use textual data to discover, measure, or draw inferences about your topic of substantive interest.

Six Key Principles

Key Principles

1Social science theories and substantive knowledge are essential for research design.
2Text analysis does not replace humans—it augments them.
3Building, refining, and testing social science theories requires iteration and cumulation.
4Text analysis methods distill generalizations from language.
5The best method depends on the task.
6Validations are essential and depend on the theory and the task.

Adaptation of Table 2.1 in Grimmer et al. (2022)

Key Principles

Three Quick Examples

Radical Politics

Figure 5  from Bonikowski, Luo and Stuhler (2022)

Devaluation of Feminized Work

Figure 3  from Jiang (2025)

Temporalities of Climate Change

Figure 7 from Stuhler, Tavory, and Wagner-Pacifici (2026)

Work Session

Designing a Survey Instrument

For the rest of today’s session, please
work on your Qualtrics surveys.

Enjoy the Break

References

Note: Scroll to access the entire bibliography

Bonikowski, Bart, Yuchen Luo, and Oscar Stuhler. 2022. “Politics as Usual? Measuring Populism, Nationalism, and Authoritarianism in U.S. Presidential Campaigns (1952–2020) with Neural Language Models.” Sociological Methods & Research 51(4):1721–87. doi: 10.1177/00491241221122317.
Bonikowski, Bart, and Laura K. Nelson. 2022. “From Ends to Means: The Promise of Computational Text Analysis for Theoretically Driven Sociological Research.” Sociological Methods & Research 51(4):1469–83. doi: 10.1177/00491241221123088.
Grimmer, Justin, Margaret E. Roberts, and Brandon M. Stewart. 2022. Text as Data: A New Framework for Machine Learning and the Social Sciences. Princeton University Press.
Jiang, Wenhao. 2025. “The Cultural Devaluation of Feminized Work: The Evolution of U.S. Occupational Prestige and Gender Typing in Linguistic Representations, 1900 to 2019.” American Sociological Review 90(5):755–87. doi: 10.1177/00031224251362351.
Stuhler, Oscar, Iddo Tavory, and Robin Wagner-Pacifici. 2026. “Time and Climate Change: U.S. Media Representations of Climate Actions, Horizons, and Events (2000 to 2021).” American Sociological Review 91(1):158–89. doi: 10.1177/00031224251403596.