Case Study
How What Works for Health Built Evi, a Tool Grounded Only in the Evidence
September 2026
The Challenge
The University of Wisconsin Population Health Institute has spent 15 years researching community health and equity strategies through its County Health Rankings & Roadmaps program, including a treasure trove of information in the What Works for Health database. There, community leaders can access hundreds of evidence-informed policies, programs, and systems changes, each with an evidence rating and a disparity rating.
But the people who need information about the policies and programs rarely arrive with a clear search term. Often, they arrive with a problem. A county health officer knows youth mental health is getting worse and doesn't know where to begin. A coalition wants to improve food access but can't name the intervention they're looking for. In short, database users don't necessarily know the research vocabulary. While keyword searches describing the problem help, they return a list of results that requires the user to stitch together the evidence on their own.
This left an opportunity for What Works for Health to increase accessibility and usability by moving beyond what traditional search could provide. AI-powered chat and search was a strong candidate, but What Works for Health's credibility is deeply rooted in the quality of their own highly researched work. An AI assistant that improvised or quietly reached out to the open web would undermine the exact thing that makes the database worth using: that every claim traces back to evidence the team has evaluated. They needed a solution with built-in guardrails and controls to meet their strict requirements for data accuracy.
The standards went beyond sourcing. The What Works for Health team has spent years refining how findings are framed. For example, they deliberately moved away from using rankings in their language, since they know each community's conditions are different and interact in their own ways: there is no single "best" strategy. Similarly, they thoughtfully describe what a strategy has the potential to do rather than what it will do and always use person-first construction. Any tool speaking on their behalf would have to carry all of that, not just the underlying facts.
The Solution
What Works for Health partnered with the digital agency Forum One and Dewey to build Evi — short for evidence — as a natural language guide to their database that responds using only their own body of work.
Forum One brought the project to Dewey and handled the code work to launch the experience on the County Health Rankings & Roadmaps site. Dewey designed the embed, and built and owns Evi's answer behavior: what it says, how it says it, and where it stops. The What Works for Health team owned corpus selection, with close coaching from Dewey on what belongs in an answer engine's knowledge base and what doesn't.
Curate the corpus, not the archive
The build began with a deliberately narrow library. Rather than pointing the system at everything the organization had ever published, Dewey worked with the content team to assemble only what Evi should draw on: the What Works for Health database itself, plus internal resources explaining the reasoning behind how the team evaluates evidence. Today that library holds roughly 400 curated documents, and what was left out was decided as carefully as what went in.
Webinars are the clearest example. Evi can distinguish a guest presenter's words from the organization's own position, but even correctly attributed, a guest's framing carries organizational weight once Evi surfaces it. So individual webinar recordings stay out of the knowledge base, and Evi points people to the webinar library instead.
Teach it their judgment, not just their content
Curation determines what Evi knows. It doesn't determine how Evi talks. Dewey interviewed the What Works for Health team extensively to surface the editorial standards that experienced staff apply without thinking, then encoded them as explicit rules:
- No ranking language. Asked which county is healthiest, Evi doesn't answer the question as posed. Instead, it explains that it can help find health strategies, then routes to the Health Data portal and three other places on the site built for exactly that.
- Person-first construction. For example, "individuals with low incomes," never "low-income individuals."
- Nuance over guarantees. Evi describes what a strategy has the potential to do and what the evidence suggests. It does not recommend, promise, or guarantee outcomes.
- Ratings travel with the claim. Every answer cites the evidence rating and disparity rating where available, and always links back to the underlying strategy page.
That last rule is what makes Evi checkable. A user never has to take its word for anything; the source is one click away, with its rating attached.
Design the edges
Much of the work in a trustworthy answer engine is in what it declines to do. Evi's guardrails are custom-built for this partnership, not generic safety filters. When a question falls outside the reviewed evidence, Evi acknowledges what the person is looking for, says plainly that it can't answer, and directs them to the right tool or page.
The experience is designed so people rarely hit a dead end in either direction. After each answer, Evi suggests follow-up questions that pull users deeper into the evidence rather than leaving them with a single result.
Test it like credibility depends on it
Before launch, Evi went through hundreds of test queries across three tracks: on-topic accuracy, adversarial red-team prompts, and dedicated hallucination testing. The What Works for Health leadership team reviewed answer samples and gave feedback across two full cycles before going live.
Testing didn't stop at launch. Dewey owns ongoing testing as a standing responsibility, watching for drift as the corpus and the underlying models change.
“Dewey took the time to work with us until we felt comfortable with how Evi represented our work. Initially, I was skeptical about an AI tool standing up to the rigor, nuance, and level of detail that the What Works for Health database is known for. The Dewey team always took our feedback seriously! Ultimately, the testing, training, and guardrails alongside our ability to monitor and continuously adapt convinced us.”
The Impact
Today, Evi is live on the County Health Rankings & Roadmaps site, giving communities a way to talk about real problems in their own words and get solutions drawn entirely from reviewed evidence. Since launch, communities have asked Evi more than 2,800 questions across nearly 700 conversations.
The questions show who is reaching for the evidence. People arrive mid-problem, in plain language, often with the constraints already attached. This can sound like "we need very low budget strategies for engaging with youth, specifically middle school age girls" or "there's a lot of obesity in my county. What programs teach people to eat healthier?" These are practitioners in the middle of the work, and the range is wide: youth and schools, air and water quality, equity and disparities, mental health, food access, and rural communities.
Users also don't stop at one question. In fact, conversations run more than four questions deep on average, which is exactly what the follow-up design was built to do.
Not every question Evi receives can be answered by What Works for Health's research, which means the 89% answer rate is right on target. That's because it was built to answer only what What Works for Health's own reviewed evidence can support, so when a question falls outside that body of work, the custom guardrails are designed to politely decline and point the person toward where they can look instead. A tool like this reaching 100% would be the warning sign, because it would mean it had started stretching beyond the evidence to fill gaps, which is precisely the failure the team designed Evi to avoid. In a field where a wrong answer can send a community down the wrong path, knowing when to stop is worth as much as knowing the answer.
The usage data is only half the story. Evi is also a research instrument pointed at the organization's own audience, giving them a continuous record of what practitioners need, in their own words, at a granularity no survey or search log delivers.
Mental health is behaving like a top-level concern, not a subcategory. The site currently files it under Clinical Care, but the questions coming into Evi suggest it's one of the largest areas of interest in its own right. The volume is itself a signal, since people reach for a chat interface most often when the navigation didn't get them there. Dewey surfaced the pattern to the team as a candidate for a future structural change to the site's information architecture.
Equity framing is core to how users think. It shows up in the phrasing of questions and in how topics cluster together, not as a separate category people ask about. That's independent evidence for a direction What Works for Health is already committed to deepening.
The questions Evi can't answer carry their own signal. While some fall outside the database by design, others point to real demand that What Works for Health hasn't covered yet. Every unanswered question is captured and reviewed, becoming a running list of candidates for new strategies to research and add. The team receives the data in weekly reports and meets quarterly with Dewey to turn it into decisions.
“We've spent years continuously refining how people search and filter What Works for Health strategies to find what they need or explore the breadth of offerings, so when Dewey came into the picture, our job was making sure Evi fit into that experience seamlessly. We focused on where in the user journey audiences would find the most value, how Evi surfaced relevant questions, and whether the language matched what this audience expects. The Dewey team took our audience and UX knowledge seriously at every step, and that productive collaboration helped sharpen the final product.”
What's Next
The near-term work is widening Evi's reach in two directions.
The first is coverage: continuing to expand the library with new research that meets Evi's evidence requirements, so more questions land on grounded answers and fewer need a redirect.
The second is discoverability, since many people find County Health Rankings & Roadmaps by landing on a specific strategy page from a general internet search, instead of their homepage that features Evi. The teams are exploring letting Evi live on those more granular pages, meeting people at the moment a question actually occurs to them.
Underneath both is the same principle that shaped the project from the start. Evidence only helps communities if they can find it, understand it, and trust it. Evi is built to make the first two effortless without ever putting the third at risk.
Want to build your own?
Connect with us to see how Dewey helps organizations turn trusted expertise into everyday tools.
Book a Consult