Methodology Rationale: S6 (current wave, fielding)
Wave: S6 (with the s6-short variant). This is the current wave, now fielding with Verasight as a nationally representative replication. This brief describes the design; there are no published S6 results yet. S6 result numbers will be published only after the wave closes and a frozen snapshot exists. Until then, the locked S5 snapshot remains the confirmatory archive.
Validator-facing companion to the S6 instrument. Written for a survey methodologist auditing the current design, and for lay validators who want to understand why the survey is built the way it is. The exact wording, item structure, and analysis plan are published with the study materials to download and inspect; this document explains the reasoning behind those choices. The completed S5 wave differs in several specific ways; those are described in the S5 methodology brief and the S5 to S6 changes summary, and are not mixed into this document. Where this brief refers to evidence from the locked S5 dataset, it says so explicitly.
Executive summary
The study asks whether ordinary Americans endorse a set of governance principles and a ballot measure built on them, and it is engineered so that an endorsement, if we find one, is hard to dismiss as an artifact of survey design. Respondents rate the abstract principles first, before they ever see the ballot or any enforcement mechanism, so their principle ratings cannot be contaminated by reactions to the specific policy. Item order within the principles block is randomized per respondent under fixed constraints, which neutralizes primacy and fatigue effects while keeping three diagnostic items in known positions. Those three fixed-position items (a cross-partisan-vocabulary opener, an instructed attention check, and a reverse-coded item) together separate genuine responders from inattentive or reflexively agreeable ones; two of them (B-bal and B-rev) are themselves substantive honesty principles, while one (B-att) is purely an attention check. Comprehension checks confirm that respondents actually understood what they read. Tightly scoped framing guards against demand effects: the study uses a neutral sponsor label and no directive reminders. Every quality failure is recorded at analysis time rather than used to block a respondent mid-survey, which preserves an honest denominator and prevents the survey from "teaching" people the answers we want. S6 reports a two-tier sample (a headline ballot sample and a stricter confirmatory clean sample) so a reviewer can see how the result behaves under lenient and strict screening.
1. Principles-first ordering
Every respondent rates the principles on a dedicated page that comes before the ballot page and before any mechanism or enforcement content.
The reason is contamination control. If a respondent saw the concrete ballot measure first, with its enforcement teeth and its implementation details, their later ratings of the abstract principles would partly reflect how they felt about the specific policy. We want the cleanest possible read on the principles themselves: do people agree that, for example, government officials of both parties should explain major decisions honestly, independent of any particular enforcement scheme? Asking the principles cold, before the policy is introduced, means a high principle-endorsement rate cannot be explained away as "they only agreed because they already liked the ballot measure." The ordering runs from the most abstract (principles) to the most concrete (a specific ballot measure and named mechanisms), so each later, more loaded stimulus cannot reach backward and color the cleaner measure that preceded it.
The principles are not a separate questionnaire placed in front of the ballot. They are the substance the measure enacts. The ballot text makes honest reasoning on the public record the core requirement of the law, and it directs the legislature to set up the system "consistent with the principles of honest governance in this measure." Agreeing or disagreeing with the principles is therefore agreeing or disagreeing with a binding requirement of the measure, not with framing the researchers added on top of it. That is why rating them first is not priming: respondents are evaluating part of the law before they see the rest of it.
2. Stratified randomization of Section B item order
Section B contains 14 items. Thirteen of them are substantive: the 11 governance principles (P1 to P10 plus P6-narrow), the cross-partisan-vocabulary item (B-bal, which is itself a substantive honesty principle), and the reverse-coded honesty item (B-rev, also a substantive principle stated in the negative). The 14th item (B-att) is purely an attention check and the only Section B item that does not measure a view about government. The presentation order is generated separately for each respondent by a fixed, reproducible procedure, so any respondent's ordering can be regenerated and checked after the fact.
The randomization is stratified, not free. Three slots are fixed for every respondent:
- Position 1:
B-bal, the cross-partisan-vocabulary item. - Position 7 (roughly 60% through):
B-att, the instructed attention check. - Position 11:
B-rev, the reverse-coded item.
The remaining 11 governance principles are shuffled into the open positions (2 to 6, 8 to 10, 12 to 14) by a plain per-respondent shuffle.
Two methodological problems motivate this. First, order and primacy effects: if every respondent saw the principles in the same fixed order, any item-level differences could be confounded with position (early items get fresher attention, later items catch fatigue). Randomizing per respondent averages position out of the per-principle estimates. Second, we still need the quality items to land in comparable places for everyone, because their diagnostic value depends on where they sit. The attention check is most informative mid-section, once a speed-clicker has settled into a rhythm but before end-of-survey fatigue, so it is fixed at position 7. The reverse-coded item is held at position 11, well separated from the attention check, so that failing both is far stronger evidence of inattention or adversarial responding than a single stray click from a momentary lapse.
Randomization is a plain shuffle in S6 (changed from S5)
S5 layered an extra constraint on top of the shuffle: it rejected any ordering that placed three principle items from the same thematic cluster back to back, redrawing until the order satisfied that rule. S6 drops this rule. The order is now a plain per-respondent shuffle of the 11 principles into the open positions, with only the three fixed diagnostic slots preserved.
The reason for the change is cross-instrument fidelity. S6 is fielded on Qualtrics (including through Verasight), and Qualtrics' question randomization cannot reproduce a "no three from the same cluster in a row" constraint. Rather than run two subtly different randomization schemes on the live site and the Qualtrics version, S6 uses the plain shuffle in both, so the two instruments are identical on this point. The cost is small: with three fixed slots already breaking up the principle run, and per-respondent randomization averaging position out of the estimates, an occasional same-theme adjacency for an individual respondent does not bias the aggregate principle estimates. (The locked S5 wave keeps its cluster rule; only S6 changes.)
3. The role of each fixed-position item
The three fixed-position items are designed to catch different failure modes (and, for two of them, to also stand as substantive principle ratings). They are deliberately complementary rather than redundant.
B-bal (cross-partisan-vocabulary item, position 1). Wording: "Government officials of both parties should explain their decisions honestly to the people they serve." This item sits at position 1 deliberately. It is the very first principle a respondent reads, and its job is to neutralize partisan priming before the rest of Section B begins. By stating, up front, that the survey is about the conduct of officials "of both parties," and by using language that lands naturally for a more conservative or skeptical respondent ("the people they serve"), it signals that this is not a partisan instrument and that the respondent is being asked about a standard that applies symmetrically across the political spectrum. The intent is to suppress the reflex (in either direction) to read what follows as advocacy for or against one party, so that the principle ratings that come after it are answered honestly on the merits.
B-bal is also a substantive honest-government principle in its own right and is counted as one of the 13 substantive items in the agreement index (see section 3a). It also plays a second role in the S6 acquiescence rule (see the B-rev item below).
B-att (instructed attention check, position 7). Wording: "To confirm you are reading carefully, please select '4 - Neither agree nor disagree' for this item." The correct answer is exactly 4. This is a direct instructed-response check: anyone reading the item will pass, anyone clicking without reading will almost certainly miss it. It catches inattention directly. In S6 it is also an in-survey gate (one same-page retry with an inline notice, then a screen-out terminate); see section 7 for why a content-neutral instructed item can be gated without coaching the substantive result.
B-rev (reverse-coded acquiescence check, position 11). Wording: "Government officials should be allowed to knowingly put false reasons into the official public record of a major government decision." A genuinely engaged respondent of any political stripe should disagree. The item is reverse-coded so that an agreeing answer is a red flag. The word "knowingly" is deliberate: it disambiguates acquiescence from a principled position about free speech or the limits of speech regulation, because agreeing here requires endorsing deliberate fabrication in the official record, not merely defending abstract speech rights.
S6 uses a contradiction-based rule for B-rev. A response is flagged only when the respondent agrees with B-rev (5 or higher on the 1-to-7 scale) and also agrees with B-bal (5 or higher). Agreeing that officials should be allowed to knowingly fabricate the record while also agreeing that officials should explain honestly is internally contradictory, which is the signature of inattentive yea-saying. This is deliberately narrower than the S5 rule (which excluded at B-rev of 3 or higher): it protects a respondent who genuinely disagrees with both propositions, or who holds a coherent libertarian-adjacent position, and fires only on the incoherent pattern.
Together these three detect inattentive and acquiescent responders without penalizing genuine ones. A genuine respondent agrees with B-bal, answers 4 on B-att, and disagrees with B-rev, a pattern that costs them nothing. Because the items are spread across the section and target distinct behaviors (vocabulary sensitivity, instructed attention, directional consistency), passing them is hard to fake and easy for an honest reader to satisfy.
3a. What counts as a principle in the agreement index
The presentation structure above (11 shuffled principle items plus three fixed-position items) describes how Section B is shown. The agreement index, the number behind the "principle agreement" figure, counts all 13 substantive items: the 11 governance principles (P1 to P10 plus P6-narrow), the cross-partisan-vocabulary principle (B-bal), and the reverse-coded honesty principle (B-rev). Only the instructed attention check (B-att) is left out, because it measures attention rather than a view about government.
B-rev is reverse-scored before it enters the mean (the index uses 8 minus the raw response). Agreeing with B-rev means endorsing officials knowingly filing false reasons into the public record, so on the principle scale a low raw answer is high principle support. Reverse-scoring puts it on the same direction as the other twelve.
This is a deliberate, disclosed deviation from the preregistration, which scored the index over the 11 governance items only. B-bal and B-rev are genuine honest-government principles, not mere filler, so they belong in the count. The reverse-scoring of B-rev and the 13-item index are fixed in the published analysis plan and applied uniformly, so the reported number cannot drift from the documented rule. (On the locked S5 dataset, the 11-item and 13-item definitions agreed to the first decimal place; the S6 figure will be reported once the S6 snapshot is frozen.)
4. Comprehension checks, not only instructed checks
Instructed checks like B-att confirm that a respondent is paying attention. They do not confirm that the respondent understood the material. A person can dutifully select the instructed box and still have skimmed past what the ballot measure actually does.
For that we use comprehension checks tied to the content. The S6 ballot includes a fact-recall question about the implementation window (the "Two years" item), treated as a hard exclusion because the number appears in the ballot text and the question is trivially easy for a reader. A final attention check at the demographics stage asks what the survey was about, with the correct answer being "Principles for how government should explain major decisions." The distractor options are chosen to be plausible and frame-neutral, so the check does not prime a positive or negative reaction to the measure itself.
The reason we use both kinds is that they fail independently. Instructed checks catch the click-without-reading respondent; comprehension checks catch the read-without-understanding respondent and confirm that an endorsement of the ballot measure is an informed one. A support number from respondents who both attended and understood is far more defensible than one resting on attention alone.
5. Guarding against demand characteristics
The study guards against demand characteristics (the tendency of respondents to give the answer they think the researcher wants) through structural choices rather than researcher framing. No pre-framing or directive reminders appear in the S6 or s6-short instrument; respondents rate the principles cold and first, proceed to the ballot, and then see the mechanism options. Nothing in the flow tells the respondent which answer the researchers prefer. The neutral study sponsor label ("American Institutions Study 2026") avoids signaling a partisan or advocacy origin.
6. Citizenship handling
S5 relied on a vote-2024 "not eligible" answer as a proxy for non-citizenship. S6 removes that proxy and replaces it with an explicit screener. The welcome page asks "Are you a U.S. citizen?"; a "No" hard-blocks participation, and Prolific's recruitment pre-screener also filters at the source. As a backstop, any response that indicates the respondent is not a U.S. citizen is excluded from the official U.S.-adult counts in both reported samples, so a non-citizen never enters the headline. This is both more direct than the old proxy (eligibility to vote and citizenship are not the same thing) and more transparent, because the gating signal is an explicit answer rather than an inference.
7. Why exclusions are applied at analysis time, and the two-tier sample
With one exception, every quality failure in S6 is recorded, not blocking. The reverse-coded acquiescence rule, the comprehension checks, the speed floors, and the final attention check are all computed afterward from the stored responses, and a respondent who fails them completes the survey normally. The exception is the instructed attention check B-att: S6 enforces it as an in-survey gate. A respondent who does not select the instructed answer stays on the same page with their answers intact and sees an inline notice that one item on the page requires a specific response and was not answered as instructed, with one remaining attempt; a second miss ends the survey as a screen-out. The screened-out row is stored with a dedicated screened-out status, so it never counts as a completion and the failure stays in the record. Both fielding paths enforce this identically: the live instrument in SectionBPrinciples.tsx and the Qualtrics/Verasight survey via a branch-and-terminate flow in the survey flow. The post-hoc B-att exclusion in computeResponseFlags() is retained as a backstop for any row that reaches analysis without having passed the gate (for example a campaign-path response, or an S5 row scored under the S6 rules).
This is deliberate, and it matters for validity in three ways. First, an honest denominator: if we silently bounced people who failed a check, we could never report the true failure rate or audit whether our exclusion rules were too aggressive. Recording every outcome, including the B-att screen-outs, keeps the B-att and B-rev rates reportable. Second, no teaching effect on the substantive result: a survey that blocks you and makes you retry a substantive item until you "get it right" would coach respondents toward a desired answer and manufacture the very consensus we are trying to measure, so the substantive checks (B-rev, the comprehension checks, the speed floors) are never blocking. The B-att gate is exempt from this concern because its instructed answer is content-neutral, "4 - Neither agree nor disagree." Gating it removes respondents who are not reading the items at all, but it cannot nudge anyone toward agreeing or disagreeing with the ballot or any principle, so it cannot inflate measured support. Third, reproducible, pre-committed exclusions: because the rules are fixed before data collection and applied uniformly afterward, anyone can re-run them on the raw data and reproduce our analyzed sample. There is no analyst discretion at the moment of exclusion.
The S6 exclusion architecture is two-tier:
- The headline ballot sample excludes on every hard failure except the
B-revcontradiction flag, so the primary support number does not hinge on one acquiescence item. - The confirmatory clean sample excludes on all hard failures, including the
B-revcontradiction flag, for the tightest analyses where acquiescence is the most serious threat.
Reporting both lets a reviewer see whether the result is robust to how strictly we screen. (On the locked S5 data, applying the S6 rules recovered respondents at essentially the same support level, within a few tenths of a percentage point; the S6 figures will be reported once the S6 snapshot is frozen.)
8. How this avoids biasing the result
Pulling the pieces together against the standard threats to survey validity:
- Acquiescence (yea-saying).
B-revis reverse-coded, so a reflexive agreer reveals themselves by endorsing knowing fabrication. The confirmatory sample screens on theB-revandB-balcontradiction; the headline sample is reported without it so the main number does not depend on a single item. Either way, indiscriminate agreement is detected rather than absorbed into "support." - Order and primacy effects. Per-respondent stratified randomization of Section B averages item position out of the principle estimates. The three quality items stay in fixed, comparable positions (1, 7, 11). S6 uses a plain shuffle for the remaining principles, dropping the S5 no-three-in-a-row cluster rule so the live and Qualtrics instruments randomize identically.
- Demand characteristics and leading. Principles are rated cold and first; no pre-framing or reminders appear in the S6 or s6-short instrument. The neutral sponsor label avoids signaling a partisan or advocacy origin.
- Comprehension masquerading as agreement. Content comprehension checks (the implementation-window item and the final "what was this about" check) confirm that ballot support comes from people who understood what they read, not just people who clicked attentively.
- Partisan priming and vocabulary bias.
B-balopens Section B with a principle stated in cross-partisan vocabulary ("officials of both parties," "the people they serve"), so the respondent's first cue about the instrument is that it is not a partisan exercise. - Selection and denominator integrity. Exclusions on the substantive checks are non-blocking and applied at analysis time, which keeps the true completion and failure rates visible and prevents the survey from coaching respondents into the desired answer. The one in-survey gate,
B-att, screens on a content-neutral instructed answer, so it filters non-readers without touching the substantive result, and its screen-outs are recorded rather than hidden. The explicit citizenship screener keeps the U.S.-adult denominator clean.
9. What a skeptic might still object to, and our answer
"You primed people with the principles, so of course they liked the ballot." The principles block is informational, not persuasive: it asks respondents what they believe, not what they should believe. Respondents who disagree simply disagree. The principles are phrased in neutral, cross-partisan language (tested by the B-bal vocabulary-balance item), and the ballot measure is presented afterward without any instruction to be consistent with prior answers. Moreover, each principle item stands alone with a "no opinion" escape, so respondents are never forced into a pattern. The principles are also not external to the measure: the ballot makes honest reasoning the binding requirement and directs implementation consistent with these principles, so rating them is rating part of the law, not a separate persuasion step.
"Attention checks just measure compliance, not real opinion." Agreed, which is why we pair instructed checks with content comprehension checks and a reverse-coded consistency item. Passing all of them requires reading, understanding, and answering consistently, not merely following an instruction.
"You threw out the respondents who disagreed." No. Exclusions target inattention, acquiescence, speeding, and failed comprehension, defined before data collection and applied uniformly. Genuine disagreement is not a failure condition. The substantive failures are recorded rather than blocked, so that screening is fully auditable and reproducible on the raw data; the one in-survey gate, B-att, screens only on a content-neutral instructed answer and its screen-outs are recorded too, so no substantive position is ever filtered at the door.
"B-rev could exclude principled libertarians who distrust speech regulation." The S6 rule was tightened precisely for this. It treats B-rev as a failure only in combination with contradictory agreement on B-bal (agreeing both that officials must explain honestly and that they may knowingly fabricate the record), which is incoherence, not a coherent libertarian stance. And the headline sample does not screen on B-rev at all.
10. The s6-short variant
s6-short is not a separate study. It is a trimmed, low-cost arm of S6 designed for large-N fielding at a much lower per-complete cost than full S6. It reuses the S6 ballot face, mechanism page, and panel questions, and its responses are kept in a separate partition so they never mix with full-S6 responses.
What s6-short keeps and what it drops
s6-short keeps the full Section B principles battery and its quality checks: the instructed attention check (B-att) and the reverse-coded B-rev both remain, and the final demographics attention check is retained. So the principle-agreement measure and the core attention and acquiescence screens are identical to full S6.
It trims content that the headline estimates do not need:
- mechanism rating becomes opt-in (rate at least one) rather than a forced rate-all-four; the forced choice is kept.
- the panel-composition page is optional or collapsed (skip or expand), and the rate at which people opt in is itself recorded.
- the re-vote drops the self-reported "did anything change" yes/no item, but keeps the one-line re-vote reason (the merged "main reason, and what changed if anything" open text), which is required on both arms. Net movement is still computable from the vote versus the re-vote.
- the post-mechanism reflection page (recall, pull-pause, attention open text) is dropped.
- per-mechanism open text and mitigation feedback are dropped.
- the region field is dropped, and is reconstructable from the state field.
For data quality, s6-short additionally relies on a completion-time floor (its threshold is preregistered after the pilot median is known), plus Prolific's approval and attention infrastructure and standard straight-lining and duplicate checks.
How s6-short and full S6 combine (pooling rule)
- Pre-divergence measures are poolable. Principle agreement and initial ballot support are measured under identical context in both arms (everything before the mechanism page is the same), so S6 and s6-short responses may be pooled for these headlines, with the arm reported as a subgroup and a preregistered arm-effect check run first.
- Post-divergence measures are pool-with-check. The forced-choice mechanism preference and especially the post-mechanism re-vote are preceded by different mechanism and panel engagement across the two arms. These are reported primary-from-S6, with s6-short shown alongside as large-N corroboration, and pooled only if the preregistered arm-effect check shows no material difference.
- S6-only measures (per-mechanism rate-all-four means, full-sample panel composition, the self-reported "did anything change" yes/no, mitigation feedback, reflection and recall) are reported from full S6 only. The one-line re-vote reason is collected on both arms and reported for each.
Full detail is in the S6 pre-registration, where the s6-short arm is consolidated in Appendix A.
11. The S6-Verasight wave (panel edition)
The s6-verasight arm is the S6 instrument prepared for fielding on a representative or quota-balanced online panel (planned: Verasight). It is registered in Appendix B of the S6 pre-registration package; this section explains the reasoning behind its four deltas from the base S6 design.
Two pre-exposure items, asked first and on their own pages
Before any study content, the panel wave asks two questions: agreement with the statement "The American people have a right to honest government service." and trust in government. Both are worded and scaled identically to the project's earlier public-website surveys, so the panel result joins to those runs on the exact same question rather than through a fuzzy match across differently worded instruments.
They are asked first because they are pre-exposure measures: the value of the agreement item is that it is answered before the respondent has seen any principle, ballot text, or mechanism content that could color it. They are asked on separate pages, agreement first, so the statement is rated with nothing else on screen and cannot be primed by the trust question. The agreement item offers no "No opinion" option because the website-survey item it bridges to offers none; adding one would change the response distribution and break the comparison. Both items lock read-only if the respondent navigates back, the same protection the rest of the instrument uses against answer revision after later content is seen.
The trust item wording
The S5/S6 welcome page asks about trust in "the federal government." The panel wave asks the question with "government" instead, matching the earlier website surveys exactly, and asks it once, early, on its own page. The welcome-page version is not asked in addition, because two near-identical trust questions minutes apart would prime each other and make both readings worse. The consequence is disclosed in the registration: any trust comparison between this wave and S5/S6 crosses that wording change and must say so.
Demographics rely on the panel's profile variables
The panel provider samples to representativeness on census-style variables (age, gender, state, region, race and ethnicity, education) and delivers those variables with the data export. Asking them again inside the survey would add length without adding information, so the panel arm's demographics page keeps only the political items (party affiliation, the independent-lean follow-up, and 2024 presidential vote), worded identically to S5/S6 so party subgroup comparisons stay direct. The end-of-survey recall check is also not fielded on this arm: the provider runs its own respondent quality control, and the instrument already carries the instructed attention check, the reverse-worded pair, and the ballot comprehension check. Citizenship is still screened in-survey on the welcome page, because the provider does not collect it. The debrief's optional paid-interview opt-in is likewise not offered: it is a recruitment feature of the project's Prolific fielding and collects no study measure.
Why a panel wave at all
The S6 package names the nonprobability recruitment frame as the design's main generalization caveat. Fielding the same locked instrument on a representative panel, with the provider's sampling targets and weighting published alongside, is the direct answer to that caveat, and the two pre-exposure items let the panel wave also anchor the headline agreement number from the earlier website surveys on a representative sample.