The Savo Two-Out-of-Three Problem
Why Every Way We Measure What People Think Makes You Give Something Up
Answering the hardest questions in business, policy, and research asks for the three things at once: depth, scale, and rigor. For a century we have had to choose which two out of the three we could live with, because having all three was impossible. Narrative Intelligence changes that dynamic, read below to learn why.
By Dr. Anthony S. Boyce, Chief Science & Product Officer, Savo · 4-minute read
Executive Summary
We inherited two great measurement instruments, the interview and the survey, and a gap between them. The hardest questions in business, policy, and research keep falling into that gap: engagement, brand, customer intent, transformation readiness, research signal, risk.
Each demands three things at once: depth (what each person really thinks and why), scale (enough respondents to generalize), and rigor (results you can trust and compare across people, units, and time). No available instrument has ever delivered all three. Interviews gave depth and rigor but not scale. Surveys gave scale and rigor but not depth. Every hybrid, from focus groups to open-text fields, traded one demand for fragments of another.
Savo's Narrative Intelligence methodology is a patent-pending end-to-end system for designing, conducting, and analyzing AI-facilitated interviews. It delivers depth, scale, and rigor together for the first time. Each interview is engineered to close a specific intelligence gap, runs as an adaptive conversation at survey scale, and produces results with the rigor of a scientific instrument. Every result is traceable to the respondent's own words.
The payoff is decision-grade evidence: measurement you can defend to a board, a regulator, or a peer reviewer. Until now, no available method could deliver it on the questions that demand depth, scale, and rigor at once. This paper explains what Narrative Intelligence is, why it matters, and what makes it possible.
The trade-off we have been forced to make
The questions that most need answering in business, policy, and research hinge on what is going on, why, and what comes next. What a market believes about a brand. What a customer base is going to do. How engagement is moving across a workforce. Whether a transformation is taking root. What a research population is really telling us. Where risk is quietly accumulating.
The answer that can effectively inform a decision has to do three things at once: it must have the depth to capture what each person really thinks, the scale to generalize across enough of them, and the rigor to compare across people and time. Every instrument we’ve been able to bring to these questions has delivered two of the three, but not all of them.
Two instruments in particular have carried the weight of these questions for a century: the interview and the survey, with focus groups and open-text fields as partial compromises between them. Each is scientifically defensible for what it was designed to do, and each is still indispensable for the narrower problems it was built to solve. Neither was built to do all three at once.
We have tacitly accepted the trade-off as the unavoidable cost of measuring what people think. Every applied research team, every product group, every brand and pricing study, every transformation review has worked around the same constraint, picking which of the three demands to under rotate on this time.
Three demands in tension
Measurement problems that inform a significant decision asks for the same three things at once. Historically, the instruments available to us have delivered two.
Depth. Rich, complete information about what each respondent thinks and why.
Scale. A representative sample large enough to generalize with confidence to the population you are trying to understand.
Rigor. Results you can trust and compare across people and across time, produced by a method you can defend.
Structured interviews deliver depth and rigor, not scale. Surveys deliver scale and rigor, not depth. Focus groups and open-text fields give you pieces of each, never all three. Narrative Intelligence delivers all three at once.
The interview
When well-designed, the interview delivers two of the three demands. It gives us depth, the rich, behavioral specificity that comes from adaptive probing and the chance to follow a thread. When the protocol is properly structured, it also gives us rigor: scores anchored to operationalized constructs, comparable across respondents. What the interview cannot give us, regardless of how well it is designed, is the third demand: scale.
The scale limit cascades into two practical problems. The first is dependence on the interviewer: the quality of the data is bounded by the quality of preparation, the discipline of the protocol, and the skill of the person executing it. The second is miscalibration across interviewers: each brings their own bias, their own interpretation, their own style. Across ten interviewers, you get ten calibrations, and those differences accumulate into interviewer drift. Every interviewer added in an attempt to scale is a new source of variance.
One common workaround has been the focus group: buy a little scale by putting more respondents in the room at once. A trade that favors scale at the expense of hearing what's true.. You gain marginally more people per hour of interviewer time, but the depth you bought with the interview format is now compromised by social-psychological distortions that don’t show up in the transcript. Conformity pressure reshapes what gets said (Asch, 1951; Cialdini & Goldstein, 2004). Status and dominance dynamics suppress minority voices. Groupthink amplifies what is already shared and quiets what is not (Janis, 1972; Stasser & Titus, 1985). Respondents say different things in groups than they do alone. The focus group does not resolve the three-demands problem. It trades interview-grade depth for slightly more respondents and group-level contamination.
Structured interviews deliver depth and rigor for the few. Focus groups trade part of that depth for slightly broader reach. Neither reaches the scale the question requires for adequate.
The survey
Surveys invert the trade. Where the interview delivers depth and rigor for the few, the survey delivers scale and rigor across the many. Tens of thousands of respondents at a cost per response that makes large sample and population-level measurement possible. Standardized items with validated psychometric properties, letting us say with confidence that engagement rose four points, or that a division's inclusion score diverged from baseline. What they cannot deliver is the demand the interview owned: depth.
The depth gap cascades into two practical problems. The first is that a score tells you what but not why. Two respondents who mark the same point on the Likert scale may mean radically different things, and the survey has no way to ask. The second is that the survey can only find what the designer thought to ask for. If a question is not on the instrument, the signal is not in the data. If a question is poorly worded, the answer is anchored to the flaw, and the survey format gives no way to detect or repair it (Schwarz, 1999; Tourangeau, Rips, & Rasinski, 2000).
The field's workaround is to staple open-text boxes onto the survey. Open text asks respondents to do the interviewer's job: synthesize, prioritize, and articulate, unprompted. That work is hard, and engagement with open-ended items is self-selected, skewed toward the motivated (Holland & Christian, 2009) and the aggrieved (Poncheri et al., 2008).
Surveys deliver scale and rigor across the many. Open text recovers only a sliver of depth, and only from those willing to write. Neither delivers depth at scale.
Narrative Intelligence
What has been impossible until now is measurement with the depth of an interview, the reach of a survey, and the rigor of a psychometric instrument at once. Savo's Narrative Intelligence platform makes it possible: an end-to-end system for designing, conducting, and analyzing AI-facilitated interviews. Each interview is engineered to close a specific intelligence gap, runs as an adaptive conversation at survey-grade scale, and produces decision-grade evidence: results scored against pre-defined constructs, with every score traceable to the respondent's own words.
The three demands, resolved
Depth. Adaptive, AI-facilitated probing applies the interview's defining strength, follow-the-thread elicitation by a skilled interviewer, to every respondent in the sample. AI removes the limit that once made this kind of elicitation a human-only craft.
Scale. Interview research has historically been capped at small samples by a single constraint: one human interviewer per respondent. An AI interviewer is not bound by that constraint. Each respondent receives the one-to-one attention of a skilled interviewer, in parallel with thousands of others, without the group-level contamination of focus groups or the closed-ended survey's inability to probe beyond its own question list.
Modern AI now makes depth and scale buildable together. A large language model fluent enough to sustain a conversation, paired with basic prompt scaffolding for context and guardrails, can probe adaptively and run in parallel across thousands of respondents. Any AI tool with a competent build can claim these properties. What good engineering does not deliver is rigor. Conversational AI without measurement science produces outputs that are fluent, confident, and plausible, but that lack the properties required for measurement: accuracy, comparability across respondents and over time, and defensibility as the basis for a decision.
Rigor. For measurement, rigor means valid quantification: results that are accurate, comparable across respondents and over time, and defensible under real scrutiny. That takes techniques honed over a hundred years of measurement science, which most AI conversation tools skip. Narrative Intelligence incorporates these techniques in four areas.
Defining the construct. Construct operationalization specifies, in advance, what is being measured. Behavioral anchors specify what evidence in the respondent's language will count as signal. Together they are the precondition of construct validity (Cronbach & Meehl, 1955; Messick, 1995) and of reliability, the property that the same construct yields the same score across raters and across occasions (AERA, APA, & NCME, 2014). What "competitive pressure" means in one respondent's conversation has to mean the same thing in the next, or the scores are not comparable.
Structuring the elicitation. What's defined in advance must be captured through a structured approach: unstructured conversation produces data that doesn't generalize. Interview Modes specify the strategy appropriate to the kind of information being captured: memory, perception, self-knowledge, and classification each have their own (e.g., Fisher & Geiselman, 1992; Geiselman, Fisher, MacKinnon, & Holland, 1985). Facilitation guides specify how the AI probes turn by turn, adapting to what the respondent just said rather than following a fixed script, with a separate signal-quality monitor running in parallel to flag any gaps in evidence as the conversation unfolds. Meta-analytic evidence shows structured interviews produce roughly twice the accuracy of unstructured ones (Sackett, Zhang, Berry, & Lievens, 2022).
Producing the insight. Qualitative narrative becomes quantitative measurement through two disciplined steps. Calibrated scoring turns captured language into construct scores via a pre-specified rubric. The scoring methodology has been calibrated against human-scored reference sets and shown to produce reliable and valid scores. The Standards for Educational and Psychological Testing set the expectations for instruments that reports quantitative scores from qualitative material: documented procedures, rater calibration, and inter-rater reliability (AERA, APA, & NCME, 2014). The methodology satisfies these requirements. Evidence Gating keeps the system from confabulating when evidence is thin. The scoring system abstains, returning "insufficient evidence" as its own outcome, distinct from a low score (which would imply the construct is absent) or a neutral score (which would imply ambivalence).
Defending the result. Traceability points every insight and construct score back to the specific respondent language that produced it. The evidence chain is auditable at the sentence level. Separation of roles keeps the agent that conducted the conversation from being the agent that scores it. Each role is accountable for mitigating a different source of error.
Treated as a measurement instrument, each AI-facilitated interview is tailored, before any participant joins, to close a specific intelligence gap: the constructs to be measured are specified in advance with behavioral anchors that define what evidence counts as signal, facilitation guides structure the elicitation, and a priori hypotheses are registered for the claims the interview is meant to test. Once deployed, the same instrument runs across any number of participants with the protocol held invariant, produces insights and construct scores traceable to the language that produced them, and reports emergent patterns through the same structured procedure as construct scoring.
None of this was possible even two years ago. The foundation is new: language models fluent enough to hold an adaptive conversation and to read open-ended language at scale. On its own, that foundation produces fluent, plausible output. Turning it into evidence you can stake a decision on takes a disciplined layer on top, one that requires measurement expertise to build. The models are now common. The discipline and human expertise to turn their output into evidence are not.
What changes
Resolving the depth/scale/rigor trade-off changes what kinds of measurement work can be done at decision-grade.
Causal inquiry, the work of asking why something is happening, has historically come with a choice: a quantitative answer at scale that did not contain the why, or an explanatory answer from the few who could be interviewed at depth. With the trade resolved, the choice does not hold. When a key cohort score moves twelve points, the why now lives in the same evidence chain as the score, traceable to specific respondent language at the sentence level, across whatever sample the question requires.
Longitudinal tracking has had a quiet problem: when a number moves, the respondent's internal yardstick may have moved with it. Response shift, the literature's term for this drift, comes in three forms (recalibration, reprioritization, reconceptualization), which standard survey measures cannot readily separate from real change (Sprangers & Schwartz, 1999). With every score traceable to the respondent's own language, the drift becomes visible. When three quarters of trend data on engagement, transformation readiness, or brand position show a shift, you can see which themes intensified, which faded, and which were scored differently this time.
Defensible inference has historically required a choice: a survey's psychometric properties or an interview's depth of evidence, never both with the same instrument. With both now in one instrument, findings can stand up to skeptical scrutiny from leadership, a regulator, or the workforce it describes. Every score has documented behavioral anchors and calibrated scoring. Every claim points to the respondent language that produced it. When evidence is insufficient, the instrument abstains rather than inventing a plausible answer it cannot support.
Open exploration has historically required choosing depth with ten participants or scale across ten thousand. Adaptive probing and structured scoring now run together at the same sample size. Whether the research is a jobs-to-be-done study, a concept test that evolves as participants react, or a first look at a new market segment, respondents answer in their own words, the instrument follows the threads each participant opens, and whatever emerges is held to the same evidence standard as any score.
Decision-grade evidence used to require choosing which of depth, scale, or rigor to underspend on. That choice is no longer required. Measurement that meets the rigor a serious decision deserves can now be done at the speed and scale of conversation. Plausible is no longer the bar. Valid is.
FAQ
A capable AI can hold a fluent conversation and probe at scale. What general-purpose tools lack is the discipline around it: a method fixed before anyone joins, every finding tied to what a respondent actually said, abstention when the evidence is thin, and results that hold up when someone checks them. That discipline, not the fluency, is what makes a result you can act on.
Because the measurement discipline is built in at every step. Constructs are defined in advance with behavioral anchors, scoring is calibrated against human-scored reference sets, the system abstains when evidence is thin, and every score traces to the respondent's own words. Together these meet the expectations the Standards for Educational and Psychological Testing set for defensible measurement (AERA, APA, & NCME, 2014).
Two things work in its favor. Every respondent meets the same instrument, so the interviewer variance that accumulates across a human team disappears. And a non-judgmental interviewer often draws out candor that people withhold face to face. The instrument also flags thin or evasive evidence rather than scoring it, so low-signal conversations do not pass themselves off as findings.
As many as the question requires. Because one AI interviewer speaks with every respondent in parallel, the same instrument runs from a handful of participants to tens of thousands, with the protocol held invariant so the scores stay comparable. Sample size becomes a decision about the evidence you need, not a limit the method imposes.
When the qualities that matter can be named in advance, the instrument measures them and reports scores. When the goal is to map a situation before anyone can name those measures, the same instrument gathers and organizes what participants report in their own words. The method changes; the standard of proof does not. Scored or not, every finding must trace to specific things real people said, unsupported claims are held back, and the procedure is documented and repeatable. What makes a finding trustworthy is the method behind it, whether or not it ends in a number.
Whenever a question needs only two of the three demands. A short pulse check with very well-defined rating scales and questions across a large population is well served by a survey. A deep conversation with a handful of experts is well served by an interview. This approach is built for the harder case, when a question needs depth, scale, and rigor at the same time.
References
Psychometric foundation
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957
Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. https://doi.org/10.1037/0003-066X.50.9.741
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.aera.net/publications/books/standards-for-educational-psychological-testing-2014-edition
Interview validity
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
Group dynamics
Asch, S. E. (1951). Effects of group pressure upon the modification and distortion of judgments. In H. Guetzkow (Ed.), Groups, leadership and men: Research in human relations (pp. 177–190). Carnegie Press.
Cialdini, R. B., & Goldstein, N. J. (2004). Social influence: Compliance and conformity. Annual Review of Psychology, 55(1), 591–621. https://doi.org/10.1146/annurev.psych.55.090902.142015
Janis, I. L. (1972). Victims of groupthink: A psychological study of foreign-policy decisions and fiascoes. Houghton Mifflin.
Stasser, G., & Titus, W. (1985). Pooling of unshared information in group decision making: Biased information sampling during discussion. Journal of Personality and Social Psychology, 48(6), 1467–1478. https://doi.org/10.1037/0022-3514.48.6.1467
Survey limitations
Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54(2), 93–105. https://doi.org/10.1037/0003-066X.54.2.93
Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The psychology of survey response. Cambridge University Press. https://doi.org/10.1017/CBO9780511819322
Open-ended response bias
Holland, J. L., & Christian, L. M. (2009). The influence of topic interest and interactive probing on responses to open-ended questions in web surveys. Social Science Computer Review, 27(2), 196–212. https://doi.org/10.1177/0894439308327481
Poncheri, R. M., Lindberg, J. T., Thompson, L. F., & Surface, E. A. (2008). A comment on employee surveys: Negativity bias in open-ended responses. Organizational Research Methods, 11(3), 614–630. https://doi.org/10.1177/1094428106295504
Cognitive interview
Fisher, R. P., & Geiselman, R. E. (1992). Memory-enhancing techniques for investigative interviewing: The cognitive interview. Charles C. Thomas.
Geiselman, R. E., Fisher, R. P., MacKinnon, D. P., & Holland, H. L. (1985). Eyewitness memory enhancement in the police interview: Cognitive retrieval mnemonics versus hypnosis. Journal of Applied Psychology, 70(2), 401–412. https://doi.org/10.1037/0021-9010.70.2.401
Response shift
Sprangers, M. A. G., & Schwartz, C. E. (1999). Integrating response shift into health-related quality of life research: A theoretical model. Social Science & Medicine, 48(11), 1507–1515. https://doi.org/10.1016/S0277-9536(99)00045-3