What two pilots at RHR can tell you about running AI-interview pilots with Savo
Executive summary
RHR International's leadership assessment is built on interviews. Interviews provide rich insights, but they do not scale, so RHR ran two scalable, science-backed, AI-moderated 360 pilots on the Savo™ platform, weeks apart, and measured what worked and what did not. Raters talked candidly and at length. Most preferred it to a survey. Three problems showed up in the first pilot. Each got a fix, and RHR measured whether the fix worked instead of assuming it did. The next round will test the output against human scorers. If your listening or assessment work depends on interviews you have not been able to scale, RHR's pilot with Savo is a model you can copy.
1. Where RHR started
RHR International's leadership assessment is built on interviews. The output is rich. The problem is that it only scales one way, by adding consultant hours.
Every assessment firm knows the two constraints, and so does every people-analytics team. Interviews do not scale. The scaled alternative, a survey, sacrifices depth. Scale, rigor, depth. Historically, you pick two.
"Interviews give the richest signal on people, and they've always been the hardest thing to scale. The Savo platform gives us an opportunity to scale that insight in a way that wasn't possible before. The team has been great to work with, and we're looking forward to learning more."
Matt Betts, PhD, Head of Assessment and Product Innovation, RHR International
2. Three tests before trusting it
An AI-moderated interview promises scale, rigor, and depth. To a team of industrial-organizational (I-O) psychologists, that sounded too good to be true. So RHR set three tests before trusting it with client work.
-
User experience. Do raters stay engaged and candid? Is it good enough to do it again?
-
Flexibility. Can it be shaped to RHR's own leadership model?
-
Measurement precision. Is the output accurate and traceable, not just plausible?
3. The pilot
RHR ran two internal 360s on the Savo platform, weeks apart. About 29 Signal Event™ interview sessions in total, with 23 raters surveyed afterward on how it felt. Internal means RHR's own people, not client work.
One caveat to keep in mind. The platform was changing while the pilot ran. New reporting views, a discovery event type, and voice-pipeline updates all landed mid-pilot. The results below reflect a product that has improved a great deal since, as a result of the design partnership.
4. What worked
-
People will talk to it, at length and candidly. Ease of use was the top-rated item in both pilots. 71% of raters said they were comfortable giving unfiltered feedback. Several said candor was easier than with a person.
-
A majority preferred it to a traditional survey, both times. The open-ended record came back more specific than a typical 360 comment field.
-
Configuration changes shipped between and during the pilots. Sessions got shorter as a result, and fewer people said the session was too long.
-
RHR's own leadership model fit the platform. The team brought its existing model into Savo Event Studio™ as a Signal Event interview, and Savo and RHR worked out the configuration together. Nothing had to be rebuilt to fit the tool.
|
Measure |
Pilot 1 |
Pilot 2 |
|---|---|---|
|
Ease of use (out of 5) |
4.4 |
4.7 |
|
"Follow-ups helped me expand my thinking" (out of 5) |
3.7 |
4.2 |
|
Session length (minutes) |
~31 |
~27 |
Source: RHR internal pilot data, two 360 rounds.
5. What we changed, and how it landed
Three problems showed up in pilot 1. Each got a fix. RHR measured whether the fix worked instead of assuming it did.
-
The interviewer kept pressing after "I don't know." The median rater got four more questions after the first "I don't know." For the second pilot, instructions were expanded to account for different roles and to refine definitions. The share of topics clearing the bar rose from 45% to 57%, but the system still did not take no for an answer. The real fix is structural: a confirmed "not observed" closes the topic. That is in spec for pilot 3.
-
The interviewer interrupted and left little time to think. Turn detection was rebuilt mid-pilot. It did not land, and complaints held steady. End-of-turn patience shipped after pilot 2. It is untested until the next run.
-
The interviewer asked the wrong questions for the rater, with a cold open. Peers were asked about behaviors only a manager or direct report can see. Raters wanted some rapport before being probed. The proposed fix for pilot 3: guides tiered by vantage point, a softer strengths-first open, and a pre-brief that "I haven't seen that" is a legitimate answer.
6. What we learned
Comfort is a user experience question. Trust is a measurement question. The two pilots taught us something about each.
-
Where control lives may not be where you think. Instructions shape what the interviewer says. A second system behind the scenes, one the participant never sees, decides what the interviewer does. For complex behaviors, the only way to find the real controls is to run the interview and measure what happens.
-
Raters can only report what they can see. More than a third of the rater-by-construct cells came back unscored. That was not reluctance. Some raters could not see the behavior from where they sat. Others saw it and described it, but never gave a specific incident, and the model was scoring incidents. This is the oldest problem in multi-source feedback. It does not go away because the interviewer is software.
-
Ease of use and trust are different things. Raters gave the experience its highest ratings for ease of use. Two asked to read their summaries to check the AI got it right. Both matter, and they are answered differently. We were able to provide both.
-
A design partnership runs both ways. Savo had to be honest about boundaries. Pause-and-resume was declined because it would turn a conversation back into a survey. RHR had to be honest about pain points, or the probes could not be tuned. In a space this new, no one has all the answers yet, including the people building the tools. That is why design partners matter so much.
7. What this means for your organization
You do not have to run a leadership assessment for this to apply. The same three things carry over to any engagement, listening, or onboarding program an HR or people-analytics team runs.
Interview depth at survey scale. Asynchronous, no scheduling, and every participant gets a follow-up question. It closes the gap between the comment box and the focus group.
Evidence you can trace. Every claim links to a transcript turn. Corroboration is reported as N of M, with the denominator. Disconfirming evidence is shown. Neither surveys nor interviewer notes can give you that audit trail.
8. Where to start
If your assessment or listening work depends on interviews you have not been able to scale, you can run a pilot the way RHR did. Pick one instrument you already trust. Set your own three tests before you start. Measure whether the changes land. Expand from there.
Scale, rigor, depth. You no longer have to pick two.
Talk to us about scoping your pilot.
Savo™, Savo Signal Science™, Signal Event™, Savo Event Studio™, Savo Session™, and Savo Insights™ are trademarks of Savo, Inc. The Savo logo is a trademark of Savo, Inc. All other trademarks and registered trademarks are the property of their respective owners. Savo is Patent-Pending