Pew Lets AI Write Its Code, Not Its Conclusions
Pew Research Center published its AI rules: no synthetic respondents, no AI-generated photos, no AI-written reports. The disclosure threshold is its own call.
The line Pew drew
On September 28, 2026, Pew Research Center republished its statement on how it does and does not use artificial intelligence, an update of a post that first appeared on Aug. 27, 2024. The most consequential sentence is a refusal: the Center says it does not use AI to create or model synthetic public opinion. Its U.S. survey results, it says, come from a multimode, probability-based panel of roughly 10,000 adults selected at random from across the country. Real people, answering.
That is a specific commitment at a moment when the alternative is cheap and getting cheaper. The rest of the post, written by Claudia Deane, describes where AI does run inside the organization: code, data preparation, text classification, copy editing, and social posts.
Why "synthetic respondents" is the fight worth having
A probability-based panel means every adult in the target population had a known, non-zero chance of being selected. That property is what licenses the leap from 10,000 people to 260-odd million — and what makes a margin of error mean anything. It is also slow and expensive: recruitment, incentives, multiple contact modes for people who don't answer online.
Large language models offer a tempting shortcut. Prompt a model to answer as a 58-year-old rural independent, repeat across a demographic grid, and you get a dataset that looks like a poll in minutes. The output is not opinion. It is a reconstruction of what text patterns in the training corpus suggest such a person might say — which means it reproduces the past, smooths over minorities of view, and cannot detect that anything has changed. Pew says it isn't doing this. Worth noting that the claim is verifiable in principle: synthetic data leaves fingerprints in the variance structure, and Pew publishes its methodology.
The uses that actually touch the numbers
Two of the disclosed uses are low-drama. Engineers use AI code assistants to build pewresearch.org, where bugs are visible and fixable. Copy editing for grammar and punctuation is described as experimental, with humans approving final copy.
The interesting one is textual data. Pew says AI assists with coding open-ended survey responses into categories and scraping websites for data. Open-ended coding is the part of survey research where a human reads thousands of free-text answers — "what worries you most about the economy?" — and sorts each into a scheme. It is laborious, and historically it is quality-controlled by having multiple coders work the same responses and measuring how often they agree.
Hand that to a classifier and the labour cost collapses. So does the natural error check. A model that systematically drops ambiguous answers into the wrong bucket doesn't produce an obvious bug; it produces a percentage that is slightly wrong and looks completely normal. Pew's stated safeguard is that researchers design the analysis plans and interpret the results, and that the Center will describe meaningful AI use in a report's methodology section.
What is shipped, and what is a policy
Be clear about what this document is. It is a statement of intent, not an audit. No tools or models are named. No error rates are published for AI-assisted coding against a human baseline. No count is given of reports that have already used these methods. Two of the four disclosed uses — copy editing and derivative products like social posts — are described as experiments, which is honest but means the settled practice is narrower than the list suggests.
The disclosure trigger is also self-defined. "Meaningful use" is Pew's judgment call, and the post commits to revisiting disclosure if AI use expands in ways that "materially affect" external products — again, Pew's assessment. Set against most research shops, which have published nothing, this is above the line. Set against an evidence standard, it is a promise.
Questions You Should Be Asking
- When a model codes open-ended responses, what share is double-checked by a human, and how often do they disagree?
- Who decides that a given use is "meaningful" enough to disclose, and is that decision recorded anywhere a reader can see?
- If a classifier's behaviour shifts after a model update, what in the pipeline would catch a two-point change in a reported figure?
- Which published reports have already used AI-assisted coding, and does their methodology section say so?
- For your own organisation: could you write this list of refusals today, and would anyone outside be able to check it?
What To Watch Next
Watch the methodology sections. Pew has committed to describing meaningful AI use there, so the signal is a Pew report whose methods appendix names a model, a task, and — ideally — an agreement rate against human coders. If that paragraph appears with numbers in it, the policy has teeth. If AI-assisted coding expands and the appendices stay silent, the threshold was doing the work all along.
- 1Publish an explicit AI-use statement that names where AI runs (code, cleaning, classification, copy) and where it never does, like generating survey responses.
- 2Never substitute LLM-generated 'synthetic respondents' for probability-sampled humans when your claim is about what the public actually thinks.
- 3When AI classifies text or writes analysis code, validate against a human-coded sample and report the agreement rate alongside your findings.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
