Methodology rationale: S5 (locked wave)

Why the completed S5 study (N=392, locked) is built the way it is, and how it avoids leading respondents.

Methodology Rationale: S5 (locked wave)

Wave: S5. Locked April 2026, N = 392. Pre-registered on OSF (osf.io/br3u2) before analysis.

Validator-facing companion to the S5 instrument. Written for a survey methodologist auditing the completed S5 design, and for lay validators who want to understand why the survey is built the way it is. The exact wording, item structure, and analysis plan are published with the study materials to download and inspect; this document explains the reasoning behind those choices. S6 differs in several specific ways; those changes are described in the S6 methodology brief and the S5 to S6 changes summary, and are not mixed into this document.

Executive summary

The study asks whether ordinary Americans endorse a set of governance principles and a ballot measure built on them, and it is engineered so that an endorsement, if we find one, is hard to dismiss as an artifact of survey design. Respondents rate the abstract principles first, before they ever see the ballot or any enforcement mechanism, so their principle ratings cannot be contaminated by reactions to the specific policy. Item order within the principles block is randomized per respondent under fixed constraints, which neutralizes primacy and fatigue effects while keeping three diagnostic items in known positions. Those three fixed-position items (a cross-partisan-vocabulary opener, an instructed attention check, and a reverse-coded item) together separate genuine responders from inattentive or reflexively agreeable ones; two of them (B-bal and B-rev) are themselves substantive honesty principles, while one (B-att) is purely an attention check. Comprehension checks confirm that respondents actually understood what they read. Tightly scoped framing guards against demand effects: the study uses a neutral sponsor label and no directive reminders. Crucially, every quality failure is recorded at analysis time rather than used to block a respondent mid-survey, which preserves an honest denominator and prevents the survey from "teaching" people the answers we want. On the locked S5 sample, the headline support number (90.3%, Wilson 95% CI 87.2 to 92.7) and the principle agreement number (99.5%) both survive the obvious methodological objections.

1. Principles-first ordering

Every respondent rates the principles on a dedicated page that comes before the ballot page and before any mechanism or enforcement content.

The reason is contamination control. If a respondent saw the concrete ballot measure first, with its enforcement teeth and its implementation details, their later ratings of the abstract principles would partly reflect how they felt about the specific policy. We want the cleanest possible read on the principles themselves: do people agree that, for example, government officials of both parties should explain major decisions honestly, independent of any particular enforcement scheme? Asking the principles cold, before the policy is introduced, means a high principle-endorsement rate cannot be explained away as "they only agreed because they already liked the ballot measure." The ordering runs from the most abstract (principles) to the most concrete (a specific ballot measure and named mechanisms), so each later, more loaded stimulus cannot reach backward and color the cleaner measure that preceded it.

The principles are not a separate questionnaire placed in front of the ballot. They are the substance the measure enacts. The ballot text makes honest reasoning on the public record the core requirement of the law, and it directs the legislature to set up the system "consistent with the principles of honest governance in this measure." Agreeing or disagreeing with the principles is therefore agreeing or disagreeing with a binding requirement of the measure, not with framing the researchers added on top of it. That is why rating them first is not priming: respondents are evaluating part of the law before they see the rest of it.

2. Stratified randomization of Section B item order

Section B contains 14 items. Thirteen of them are substantive: the 11 governance principles (P1 to P10 plus P6-narrow), the cross-partisan-vocabulary item (B-bal, which is itself a substantive honesty principle), and the reverse-coded honesty item (B-rev, also a substantive principle stated in the negative). The 14th item (B-att) is purely an attention check and the only Section B item that does not measure a view about government. The presentation order is generated separately for each respondent by a fixed, reproducible procedure, so any respondent's ordering can be regenerated and checked after the fact.

The randomization is stratified, not free. Three slots are fixed for every respondent:

  • Position 1: B-bal, the cross-partisan-vocabulary item.
  • Position 7 (roughly 60% through): B-att, the instructed attention check.
  • Position 11: B-rev, the reverse-coded item.

The remaining 11 governance principles are shuffled into the open positions (2 to 6, 8 to 10, 12 to 14), subject to a "no three in a row from the same thematic cluster" rule described below.

Two methodological problems motivate this. First, order and primacy effects: if every respondent saw the principles in the same fixed order, any item-level differences could be confounded with position (early items get fresher attention, later items catch fatigue). Randomizing per respondent averages position out of the per-principle estimates. Second, we still need the quality items to land in comparable places for everyone, because their diagnostic value depends on where they sit. The attention check is most informative mid-section, once a speed-clicker has settled into a rhythm but before end-of-survey fatigue, so it is fixed at position 7. The reverse-coded item is held at position 11, well separated from the attention check, so that failing both is far stronger evidence of inattention or adversarial responding than a single stray click from a momentary lapse.

The no-three-in-a-row cluster rule

Each principle item belongs to one of four thematic clusters:

  • reasoning (P2, P3, P4, P10): items about honest reasoning, evidence on the public record, and citizens being able to challenge stated reasons.
  • private-capture (P1, P5): loyalty to the public over private interests, and conflict-of-interest disclosure.
  • accountability (P6, P6-narrow, P7, P8, P9): consequences for fraud, removal in serious cases, and equal application of ethics rules.
  • neutral (B-att): the attention check, which is not counted toward any thematic streak.

When the order of the 11 principles is drawn for a respondent, any ordering that would place three items from the same cluster back to back is rejected and a fresh ordering is drawn in its place. In practice this rarely requires more than one redraw, and the rule has been checked exhaustively against many thousands of random orderings without a single violation. The rule applies to the shuffled run of principles; the three fixed-position items break any near-streak that could otherwise form around positions 7 and 11.

The reason this matters: a randomization that ignored clusters could still, by chance, hand a respondent three accountability items back to back, which produces a "themed run" that fatigues attention on that theme and invites straight-lining (clicking the same value down the column). Spreading the clusters keeps each principle item perceptually distinct from its neighbors.

3. The role of each fixed-position item

The three fixed-position items are designed to catch different failure modes (and, for two of them, to also stand as substantive principle ratings). They are deliberately complementary rather than redundant.

B-bal (cross-partisan-vocabulary item, position 1). Wording: "Government officials of both parties should explain their decisions honestly to the people they serve." This item sits at position 1 deliberately. It is the very first principle a respondent reads, and its job is to neutralize partisan priming before the rest of Section B begins. By stating, up front, that the survey is about the conduct of officials "of both parties," and by using language that lands naturally for a more conservative or skeptical respondent ("the people they serve"), it signals that this is not a partisan instrument and that the respondent is being asked about a standard that applies symmetrically across the political spectrum. The intent is to suppress the reflex (in either direction) to read what follows as advocacy for or against one party, so that the principle ratings that come after it are answered honestly on the merits.

B-bal is also a substantive honest-government principle in its own right and is counted as one of the 13 substantive items in the agreement index (see section 3a). It is not analyzed as a vocabulary-sensitivity contrast against other items.

B-att (instructed attention check, position 7). Wording: "To confirm you are reading carefully, please select '4 - Neither agree nor disagree' for this item." The correct answer is exactly 4. This is a direct instructed-response check: anyone reading the item will pass, anyone clicking without reading will almost certainly miss it. It catches inattention directly.

B-rev (reverse-coded acquiescence check, position 11). Wording: "Government officials should be allowed to knowingly put false reasons into the official public record of a major government decision." A genuinely engaged respondent of any political stripe should disagree. The item is reverse-coded so that an agreeing answer is a red flag. The word "knowingly" is deliberate: it disambiguates acquiescence from a principled position about free speech or the limits of speech regulation, because agreeing here requires endorsing deliberate fabrication in the official record, not merely defending abstract speech rights.

In S5, B-rev is a hard exclusion: a response is removed from the analysis cohort if B-rev is 3 (slightly disagree) or higher on the 1-to-7 scale. (S6 later relaxed this to a contradiction-based rule; that change belongs to the S6 brief, not here.)

Together these three detect inattentive and acquiescent responders without penalizing genuine ones. An acquiescent "yea-sayer" who agrees with everything will agree with both B-bal and B-rev, which is internally contradictory (you cannot consistently want honest explanations and permit knowing fabrication). A genuine respondent, by contrast, agrees with B-bal, answers 4 on B-att, and disagrees with B-rev, a pattern that costs them nothing. Because the items are spread across the section and target distinct behaviors (vocabulary sensitivity, instructed attention, directional consistency), passing them is hard to fake and easy for an honest reader to satisfy.

3a. What counts as a principle in the agreement index

The presentation structure above (11 shuffled principle items plus three fixed-position items) describes how Section B is shown. The agreement index, the number behind the headline "principle agreement" figure, counts all 13 substantive items: the 11 governance principles (P1 to P10 plus P6-narrow), the cross-partisan-vocabulary principle (B-bal), and the reverse-coded honesty principle (B-rev). Only the instructed attention check (B-att) is left out, because it measures attention rather than a view about government.

B-rev is reverse-scored before it enters the mean (the index uses 8 minus the raw response). Agreeing with B-rev means endorsing officials knowingly filing false reasons into the public record, so on the principle scale a low raw answer is high principle support. Reverse-scoring puts it on the same direction as the other twelve.

This is a deliberate, disclosed deviation from the preregistration, which scored the index over the 11 governance items only. B-bal and B-rev are genuine honest-government principles, not mere filler, so they belong in the count. The deviation does not move the result: principle agreement on the S5 analysis cohort is 99.5% under both the 11-item and 13-item definitions. The reverse-scoring of B-rev and the 13-item index are fixed in the published analysis plan and applied uniformly, so the reported number cannot drift from the documented rule.

4. Comprehension checks, not only instructed checks

Instructed checks like B-att and the ballot method-check on the S5 ballot pages confirm that a respondent is paying attention. They do not confirm that the respondent understood the material. A person can dutifully select the instructed box and still have skimmed past what the ballot measure actually does.

For that we use comprehension checks tied to the content. The S5 ballot includes a fact-recall question about the implementation window (the "Two years" item). The number appears once in the ballot text, so answering correctly requires having read it. A final attention check at the demographics stage asks what the survey was about, with the correct answer being "Principles for how government should explain major decisions." The distractor options are chosen to be plausible and frame-neutral, so the check does not prime a positive or negative reaction to the measure itself.

The reason we use both kinds is that they fail independently. Instructed checks catch the click-without-reading respondent; comprehension checks catch the read-without-understanding respondent and confirm that an endorsement of the ballot measure is an informed one. A support number from respondents who both attended and understood is far more defensible than one resting on attention alone.

5. Guarding against demand characteristics

The study guards against demand characteristics (the tendency of respondents to give the answer they think the researcher wants) through structural choices rather than researcher framing. The S5 instrument that was fielded ran with no pre-framing reminder: respondents rate the principles cold and first, proceed to the ballot, and then see the mechanism options, with nothing steering them toward support. Nothing in the flow tells the respondent which answer the researchers prefer. The neutral study sponsor label ("American Institutions Study 2026") avoids signaling a partisan or advocacy origin.

6. Why exclusions are applied at analysis time

Every quality failure in S5 is recorded, not blocking. The survey does not stop a respondent who fails the attention check, the reverse-coded item, the ballot method-check, the comprehension check, or who speeds through a section. They complete the survey normally, and the failures are computed afterward from the stored responses.

This is deliberate, and it matters for validity in three ways. First, an honest denominator: if we silently bounced people who failed a check, we could never report the true failure rate or audit whether our exclusion rules were too aggressive. Recording everything lets us state, for example, the attention-check and reverse-coded pass rates as reported quantities. Second, no teaching effect: a survey that blocks you and makes you retry until you "get it right" effectively coaches respondents toward a desired answer, which would manufacture the very consensus we are trying to measure. Non-blocking recording avoids that entirely. Third, reproducible, pre-committed exclusions: because the rules are fixed before data collection and applied uniformly afterward, anyone can re-run them on the raw data and reproduce our analyzed sample. There is no analyst discretion at the moment of exclusion.

S5 reports a single analysis sample. A response is removed if it fails any S5 hard check: the instructed attention check (B-att), the reverse-coded acquiescence item (B-rev at 3 or higher), a ballot-page method check, the implementation-window comprehension check, the final demographics attention check, or an implausibly fast completion (the principles or mechanism page under 45 seconds, or the arm page under 30 seconds). For the headline U.S.-adult cohort, S5 also excluded respondents who indicated they were not eligible to vote in 2024, used as a citizenship proxy. (S6 replaces both the single-sample design and the proxy; see the S6 brief.)

7. How this avoids biasing the result

Pulling the pieces together against the standard threats to survey validity:

  • Acquiescence (yea-saying). B-rev is reverse-coded, so a reflexive agreer reveals themselves by endorsing knowing fabrication. In S5 the item is a hard exclusion at 3 or higher, so indiscriminate agreement is detected and removed rather than absorbed into "support."
  • Order and primacy effects. Per-respondent stratified randomization of Section B averages item position out of the principle estimates, while the no-three-in-a-row cluster rule blocks straight-lining runs. The three quality items stay in fixed, comparable positions.
  • Demand characteristics and leading. Principles are rated cold and first; the fielded S5 instrument used no pre-framing reminder. The neutral sponsor label avoids signaling a partisan or advocacy origin. Nothing in the flow tells the respondent which answer the researchers prefer.
  • Comprehension masquerading as agreement. Content comprehension checks (the implementation-window item and the final "what was this about" check) confirm that ballot support comes from people who understood what they read, not just people who clicked attentively.
  • Partisan priming and vocabulary bias. B-bal opens Section B with a principle stated in cross-partisan vocabulary ("officials of both parties," "the people they serve"), so the respondent's first cue about the instrument is that it is not a partisan exercise.
  • Selection and denominator integrity. Non-blocking, analysis-time exclusions keep the true completion and failure rates visible and prevent the survey from coaching respondents into the desired answer.

8. What a skeptic might still object to, and our answer

"You primed people with the principles, so of course they liked the ballot." The principles block is informational, not persuasive: it asks respondents what they believe, not what they should believe. Respondents who disagree simply disagree. The principles are phrased in neutral, cross-partisan language (tested by the B-bal vocabulary-balance item), and the ballot measure is presented afterward without any instruction to be consistent with prior answers. Moreover, each principle item stands alone with a "no opinion" escape, so respondents are never forced into a pattern. The principles are also not external to the measure: the ballot makes honest reasoning the binding requirement and directs implementation consistent with these principles, so rating them is rating part of the law, not a separate persuasion step. Any analyst can compare principle-endorsement patterns with ballot support to check whether the ordering produced a mechanical consistency effect.

"Attention checks just measure compliance, not real opinion." Agreed, which is why we pair instructed checks with content comprehension checks and a reverse-coded consistency item. Passing all of them requires reading, understanding, and answering consistently, not merely following an instruction.

"You threw out the respondents who disagreed." No. Exclusions target inattention, acquiescence, speeding, and failed comprehension, defined before data collection and applied uniformly. Genuine disagreement is not a failure condition. Because failures are recorded rather than blocked, the screening is fully auditable and reproducible on the raw data.

"One reverse item is thin evidence of acquiescence." The S5 results change inconsequentially little when the exclusion rules are varied (within 0.2 percentage points on the S5 dataset), which is the kind of robustness check a reviewer can run directly on the frozen data. S6 went further and reports a two-tier sample so reviewers can see the result under both lenient and strict screening; that design is described in the S6 brief.