Skip to content

Commit 4e94ec4

Browse files
acoppockclaude
andcommitted
Back-port 23 of Princeton's copyedits into the source
The printed book was copyedited at Princeton and none of it came back to this repository, so in these places the website has been serving the pre-copyedit manuscript since publication. Checking every candidate against the full text history (350 commits, 879 distinct blobs) and against all four proof stages shows the pattern: the printed wording appears nowhere in the source, and 1P still matches the source in many places while 2P onward matches print. Most of the copyediting therefore happened on the manuscript, before a page was set. These 23 are the subset that can be applied mechanically and safely: the changed span carries no markup, the replacement is sliced from the printed sentence with its punctuation intact, and guards reject anything that would pull a figure caption, table label or heading out of the PDF text stream and into a paragraph. Each was read before committing. Includes two plain typo fixes the website has carried since 2022: "treue" for "true" (ch. 6) and "more likely to response" for "respond" (ch. 8). Roughly 110 further copyedits are known but not mechanical: the changed span touches a citation, cross-reference or math, or the printed sentence is interleaved with figure text. Those need to be done by hand against the print PDF, chapter by chapter. Verified: full book renders 31/31, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VGM3E2QmKNbQpLLByVw2TN
1 parent de627cc commit 4e94ec4

11 files changed

Lines changed: 21 additions & 21 deletions

‎declaration-diagnosis-redesign/crafting-data-strategy.qmd‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -220,7 +220,7 @@ The analogy between sampling and assignment runs so deep because, in a sense, as
220220

221221
@fig-ch8num4 visualizes nine kinds of random assignment, arranged according to whether the assignment procedure is simple, complete, or blocked and according to whether the assignment procedure is carried out at the individual, cluster, or saturation level. In the top left facet, we have simple (or Bernoulli) random assignment, in which all units have a 50% probability of treatment, but the total number of treated units can bounce around from assignment to assignment. In the top center, this problem is fixed: under complete random assignment, exactly $m$ of $N$ units are assigned to treatment and the $N - m$ are assigned to control. While complete random assignment fixes the number of units treated at exactly $m$, the number of units that are treated within any particular group of units (defined by a pre-treatment covariate) could vary. Under block random assignment, we conduct complete random assignment within each block separately, so we directly control the number treated within each block. Moving from simple to complete random assignment tends to decrease sampling variability a bit, by ruling out highly unbalanced allocations. Moving from complete to blocked can help more, so long as the blocking variable is correlated with the outcome. Blocking rules out assignments in which too many or too few units in a particular subgroup are treated.
222222

223-
The second row of @fig-ch8num4 shows clustered designs in which all units within a cluster receive the same treatment assignment. Clustered designs are common for household-level, school-level, or village-level designs, where it would be impractical or unfeasible to conduct individual level assignment. When units within the same cluster are more alike than units in different clusters (as in most cases), clustering increases sampling variability relative to individual level assignment. Just like in individual level designs, moving from simple to complete or from complete to blocked tends to result in lower sampling variability.
223+
The second row of @fig-ch8num4 shows clustered designs in which all units within a cluster receive the same treatment assignment. Clustered designs are common for household-level, school-level, or village-level designs, where it would be impractical or unfeasible to conduct individual level assignment. When units within the same cluster are more alike than units in different clusters (as in most cases), clustering increases sampling variability relative to individual level assignment. Just as in individual level designs, moving from simple to complete or from complete to blocked tends to result in lower sampling variability.
224224

225225
The final row of @fig-ch8num4 shows a series of designs that are analogous to the multi-stage sampling designs shown in @fig-ch8num2 -- but their purpose is subtly different in spirit. Multi-stage sampling designs are employed to reduce costs -- first clusters are sampled but not all units within a cluster are sampled. A saturation randomization design (sometimes called a "partial population design", see @sec-ch18s10) uses a similar procedure to both contain and learn about spillover effects. Some clusters are chosen for treatment, but some units *within* those clusters are not treated. Units that are untreated in treated clusters can be compared with units that are untreated in untreated clusters in order to suss out intra-cluster spillover effects [@betsy2012detecting]. @fig-ch8num4 shows how the saturation design comes in simple, complete, and blocked varieties.
226226

@@ -306,7 +306,7 @@ To make these choices, we depend on methodological research whose main~inquiries
306306

307307
Researchers select several characteristics of a measurement strategy: who collects the measures, the mode of measurement, how often and when measures are taken, how many different observed measures of the latent outcome $Y^*$ are collected, and how they are summarized into a single measure. These design characteristics may affect validity, reliability, cost, or all three.
308308

309-
Data may be collected by researchers themselves, by participants, or by third parties. In some forms of qualitative research such as participant-observation and interview-based research, the researcher may be the primary data collector. In survey research, interviewers are typically a hired agents of the researcher, each of whom may ask questions differently. Participants are sometimes asked to collect data on themselves, either through self-administered surveys, journaling, or taking measurements of themselves using thermometers or scales. A primary concern with self-reports is validity: do respondents report their measurements truthfully? A parallel concern is raised when participants do not collect their own data, but are made aware of the fact that they are being measured by others. Finally, data may be collected by agents of governments or other organizations, yielding so-called "administrative" data.
309+
Data may be collected by researchers themselves, by participants, or by third parties. In some forms of qualitative research such as participant-observation and interview-based research, the researcher may be the primary data collector. In survey research, interviewers are typically hired agents of the researchers, each of whom may ask questions differently. Participants are sometimes asked to collect data on themselves through self-administered surveys, journaling, or taking measurements of themselves using thermometers or scales. A primary concern with self-reports is validity: do respondents report their measurements truthfully? A parallel concern is raised when participants do not collect their own data, but are made aware of the fact that they are being measured by others. Finally, data may be collected by agents of governments or other organizations, yielding so-called "administrative" data.
310310

311311
Most of the variety in measurement strategies is how those data collectors obtain their data. Data collectors can use observation and ask respondents for self-reports. Increasingly, photos, videos, sound recordings, and even water and soil measurements are used for outcome measurement. The translation of raw data, like videos, into coded data, like counts of the number of police stops, that can be used for analysis is part of $Q$ in the measurement strategy.
312312

@@ -359,7 +359,7 @@ Attrition occurs when we do not obtain outcome measures for all sampled units. T
359359

360360
Whether attrition is a problem depends on whether response ($R$) is causally affected by variables other than random sampling. If it is not, we say the missingness is completely at random, just as if we had simply added one more random sampling step to the design. Outside of explicit sampling designs, missingness completely at random is rare, though possible, perhaps due to idiosyncratic administrative procedures or computer error. If attrition is completely at random, precision suffers due to a loss of sample size, but bias is unaffected.
361361

362-
If missingness *is* affected by other variables -- some units are more likely to response because of unobserved background characteristics such as being at home when the survey taker calls -- then inferences may be biased. Attrition is doubly difficult in experiments, because if treatment affects not just how a unit responds, but *whether* it responds, then treatment-control comparisons on the basis of observed data may be biased.
362+
If missingness *is* affected by other variables -- some units are more likely to respond because of unobserved background characteristics such as being at home when the survey taker calls -- then inferences may be biased. Attrition is doubly difficult in experiments, because if treatment affects not just how a unit responds, but *whether* it responds, then treatment-control comparisons on the basis of observed data may be biased.
363363

364364
One approach is to address this weakness in the data strategy via changes to the answer strategy. A bounding approach like the one described in @sec-ch9s2ss4 (interval-estimation) is a design-based answer strategy for drawing inferences despite missingness. Model-based approaches involve reweighting the data by stratum, supposing random missingness within a stratum but not across strata. An example of a data strategy response is to intensively revist a subset of attritors to gather information about them---including information about why they attrited.For more see the disccussion in @sec-ch15s1ss1 as well as the combination of data strategy and answer strategy response in @coppock2017double.
365365

@@ -369,7 +369,7 @@ Excludability means that when we define potential outcomes, we can exclude extra
369369

370370
The figure asserts no effect of sampling $S$ on latent outcome $Y_i^*$. This assumption could be violated if the fact of being *included in the sample* changes your attitudes. For example, if the very act of being asked to be in a focus group causes subjects to reflect on their political beliefs and thereby change them, the sampling excludability assumption would be violated.
371371

372-
Next, we assert no causal effect of assignment $Z$ on outcome $Y^*$ -- except through the treatment $D$. This assumption is constantly under threat! In observational studies "instrumental variables" design, excludability is the assumption of no alternative channels through which the instrument affects outcomes except the treatment variable. In the entertainingly titled "Rain, Rain, Go Away: 176 potential exclusion-restriction violations for studies using weather as an instrumental variable," @mellon_2021 discusses how random variation in rainfall has been misused to study the effects of other treatments.
372+
Next, we assert no causal effect of assignment $Z$ on outcome $Y^*$ -- except through the treatment $D$. This assumption is constantly under threat! In observational studies using the "instrumental variables" design, excludability is the assumption of no alternative channels through which the instrument affects outcomes except the treatment variable. In the entertainingly titled "Rain, Rain, Go Away: 176 potential exclusion-restriction violations for studies using weather as an instrumental variable," @mellon_2021 discusses how random variation in rainfall has been misused to study the effects of other treatments.
373373

374374
We further assume that $Q$ does not affect $Y^*$. Hawthorne effects, in which the fact of being measured changes outcomes, are an example a violation of this kind excludability assumption. If outcomes depend on whether subjects know they are being measured or do not, then we cannot exclude the effect of measurement from our effect estimates.
375375

‎declaration-diagnosis-redesign/declaration-in-code.qmd‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -716,7 +716,7 @@ kable(
716716

717717
### Tidying statistical modelling function output
718718

719-
Let's unpack the `.summary` argument. In this case we sent the `tidy` function from the `broom` package (the default). Understanding what `tidy` does opens a window into the way we match estimates and estimands. The tidy function takes many model fit objects and returns a data frame in which rows represent estimates and columns represent statistics about that estimate. The columns typically include the estimate itself (`estimate`), an estimated standard error (`std.error`), a test statistic of some kind reported by the model function such as a *t*-statistic or *Z*-statistic (`statistic`), a *p*-value based on the test statistic (`p.value`), a confidence interval (`conf.low`, `conf.high`), and the degrees of freedom of the test statistic if relevant (`df`).
719+
Let's unpack the `.summary` argument. In this case we send the `tidy` function from the `broom` package (the default). Understanding what `tidy` does opens a window into the way we match estimates and estimands. The tidy function takes many model fit objects and returns a data frame in which rows represent estimates and columns represent statistics about that estimate. The columns typically include the estimate itself (`estimate`), an estimated standard error (`std.error`), a test statistic of some kind reported by the model function such as a *t*-statistic or *Z*-statistic (`statistic`), a *p*-value based on the test statistic (`p.value`), a confidence interval (`conf.low`, `conf.high`), and the degrees of freedom of the test statistic if relevant (`df`).
720720

721721
A key column in the output of tidy is `term`, which represents which coefficient (term) is being described in that row. We will often need to use the term column in conjunction with the name of the estimator to link estimates to inquiries when there is more than one. If in the regression we pull out two coefficients (e.g., for treatment indicator 1 and for treatment indicator 2), we need to be able to link those to separate inquiries representing the true effect of treatment 1 and the true effect of treatment 2. Term is our tool for doing so. The default is for term to pick the first coefficient that is not the intercept, so for the regression `Y ~ Z` there will be an intercept and then the coefficient on `Z`, which is what will be picked.
722722

‎declaration-diagnosis-redesign/declaring-designs.qmd‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -115,4 +115,4 @@ Example declaration
115115

116116
:::
117117

118-
This design declaration includes a specification of all four design elements: model, inquiry, data strategy, and answer strategy. The next four chapters will describe each of these four design elements in great detail. For now, notice that the declaration does not include a specification of $m^*$ (the true causal model), only $M$, a model we entertain for research planning purposes.
118+
This design declaration includes a specification of all four design elements: model, inquiry, data strategy, and answer strategy. The next four chapters will describe each of these four design elements in great detail. For now, notice that the declaration does not include a specification of $m^*$ (the true causal model), and includes only $M$, a model we entertain for research planning purposes.

‎declaration-diagnosis-redesign/specifying-model.qmd‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -145,7 +145,7 @@ One useful exercise is to return to and question the assumptions you have built
145145

146146
When we are agnostic, we admit we don't know whether the truth is in the set of models we consider reasonable---so we entertain a wider set than we might think plausible. We suggest three guides for choosing these ranges: the logical minimum and maximum bounds of a parameter, a meta-analytic summary of past studies, or best- and worst-case bounds, based on the substantive interpretations of previous work. A design that performs well in terms of power and bias under many such ranges might be labeled "robust to multiple models."
147147

148-
A separate goal is assessing the performance of a research design under different models implied by alternative theories. A good design will provide probative evidence no matter the treue event generating process. A poor design might provide reliable answers only for specific types of event generating processes.
148+
A separate goal is assessing the performance of a research design under different models implied by alternative theories. A good design will provide probative evidence no matter the true event generating process. A poor design might provide reliable answers only for specific types of event generating processes.
149149

150150
An important example is the performance of a research design under a "null model," where the true effect size is zero. A good research design should report with a high probability that there is insufficient evidence to reject a null effect. That same research design, under an alternative model with a large effect size, should with a high probability return evidence rejecting the null hypothesis of zero effect. The example makes clear that in order to understand whether the research design is strong, we need to understand how it performs under different models.
151151

0 commit comments

Comments
 (0)