The Inquirer's AI news scout works because an editor graded it for months
Scrape cut the noise from hyperlocal news by training on one editor's daily feedback. The prompts, not the model, did the work.
One editor at The Philadelphia Inquirer was spending about 12 hours a week just finding items for local newsletters, according to a case study published September 30, 2026 by the Lenfest Institute. The fix was an AI tool called Scrape, which the institute says has since become "load-bearing and critical infrastructure" for six newsletters, with two more planned and nearly 20 journalists following its output.
What Scrape actually does
Scrape runs daily over a curated list of sources: municipal meetings, school calendars, Facebook pages and other small, scattered outlets. It produces a tip sheet of potential newsletter items. The Inquirer's hyperlocal newsletters each cover a specific community, and the news for those communities comes from, in the words of Lenfest AI Fellow Kevin Hoffman, "a lot of really small sources, and so it's a lot to sift through."
The source describes no novel model or architecture. What it describes is a way of building around a model.
The method: the editor's own work as the answer key
The lead editor had already compiled the source list and was producing daily tip sheets by hand. Hoffman treated those handmade sheets as the benchmark. Each day the AI produced its own digest, and the team compared it with the editor's: did it extract the same insights, find things the editor missed, or miss things it should have caught?
Over three or four months, the editor annotated the digests every other day, marking what was useful and what was noise. Hoffman describes the effort as "pretty intense." Bullet points were added and removed from the prompt, and the output shifted away from rescheduled meetings and routine road closures toward community news that mattered.
Note what "training" means here. The source describes iterative prompt refinement, which is rewriting the instructions given to the model, not retraining the model itself.
Why encoding judgment was the hard part
"Newsworthiness" is not something a model knows in the abstract. The first design was region-based: scrape everything about a place like Lower Merion, then filter with one generic newsworthiness prompt. That worked "manageably" for one editor and a few newsletters but did not scale, according to Hoffman.
The replacement is a brief per newsletter, written by that newsletter's editor. For the Chester County newsletter, the brief covers what readers want (municipal and local government, localized state news, things affecting daily life such as restaurants and retail), what to exclude (straight business news like stock exchange coverage or personnel changes), who the readers are, which school districts count, and which neighboring communities do not belong. Scrape now takes all of that as input.
This is the transferable idea. The model is the same; the editorial standard is written down, owned by editors, and fed in explicitly.
Cost, and what the source leaves out
Searching many small sources is expensive. The team found that prompting can implicitly steer how deeply and how widely the system searches, but only through trial and error. Hoffman's hindsight advice is to set acceptable cost-quality trade-offs on day one and design prompts and architecture around them.
The case study is published by the Lenfest Institute, whose AI program is supported by OpenAI and Microsoft, and its results are reported by the people who ran the project. It gives no dollar costs, no measurement of how many hours the editor now saves, and no error rate.
Questions You Should Be Asking
- What did the editor's 12 weekly hours drop to, and who measured it? The case study states the starting figure but no result.
- What does Scrape miss? The comparison against the editor's tip sheets checked for gaps, but the source reports no miss rate, and a story that never reaches a tip sheet is invisible.
- What does it cost to run daily across all those sources, and who decided that was acceptable? The source admits cost was a lesson learned, not a solved problem.
- Who maintains each newsletter's brief when the editor leaves? If judgment lives in a prompt one person wrote, what happens to quality when that person moves on?
- How does a brief that excludes certain communities get audited? Explicit exclusions are a choice about whose news counts.
What To Watch Next
The signal is the expansion from six newsletters to eight and beyond. The region-based design failed to generalize from one editor to many, and the per-newsletter briefs are the fix. If new editors can write effective briefs without months of daily annotation, the method transfers. If each launch needs its own intense tuning period, Scrape is a craft project, not a repeatable tool.
- 1Train AI tools with human editorial review for weeks to ensure quality before full deployment to newsrooms.
- 2Use AI to monitor fragmented local sources like municipal meetings and social media to save journalists' research time.
- 3Start AI adoption with one specific workflow pain point before scaling to multiple teams and publications.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
