Blog

Narrative Intelligence: The Measurement Layer That Was Missing

Written by Dr. Anthony S. Boyce | 08/20/2026

The Savo Two-Out-of-Three Problem

Why Every Way We Measure What People Think Makes You Give Something Up

Answering the hardest questions in business, policy, and research asks for the three things at once: depth, scale, and rigor. For a century we have had to choose which two out of the three we could live with, because having all three was impossible. Narrative Intelligence changes that dynamic, read below to learn why.

By Dr. Anthony S. Boyce, Chief Science & Product Officer, Savo · 4-minute read

Executive Summary

We inherited two great measurement instruments, the interview and the survey, and a gap between them. The hardest questions in business, policy, and research keep falling into that gap: engagement, brand, customer intent, transformation readiness, research signal, risk.

Each demands three things at once: depth (what each person really thinks and why), scale (enough respondents to generalize), and rigor (results you can trust and compare across people, units, and time). No available instrument has ever delivered all three. Interviews gave depth and rigor but not scale. Surveys gave scale and rigor but not depth. Every hybrid, from focus groups to open-text fields, traded one demand for fragments of another.

Savo's Narrative Intelligence methodology is a patent-pending end-to-end system for designing, conducting, and analyzing AI-facilitated interviews. It delivers depth, scale, and rigor together for the first time. Each interview is engineered to close a specific intelligence gap, runs as an adaptive conversation at survey scale, and produces results with the rigor of a scientific instrument. Every result is traceable to the respondent's own words.

The payoff is decision-grade evidence: measurement you can defend to a board, a regulator, or a peer reviewer. Until now, no available method could deliver it on the questions that demand depth, scale, and rigor at once. This paper explains what Narrative Intelligence is, why it matters, and what makes it possible.


The trade-off we have been forced to make

The questions that most need answering in business, policy, and research hinge on what is going on, why, and what comes next. What a market believes about a brand. What a customer base is going to do. How engagement is moving across a workforce. Whether a transformation is taking root. What a research population is really telling us. Where risk is quietly accumulating.

The answer that can effectively inform a decision has to do three things at once: it must have the depth to capture what each person really thinks, the scale to generalize across enough of them, and the rigor to compare across people and time. Every instrument we’ve been able to bring to these questions has delivered two of the three, but not all of them.

Two instruments in particular have carried the weight of these questions for a century: the interview and the survey, with focus groups and open-text fields as partial compromises between them. Each is scientifically defensible for what it was designed to do, and each is still indispensable for the narrower problems it was built to solve. Neither was built to do all three at once.

We have tacitly accepted the trade-off as the unavoidable cost of measuring what people think. Every applied research team, every product group, every brand and pricing study, every transformation review has worked around the same constraint, picking which of the three demands to under rotate on this time.

Three demands in tension

Measurement problems that inform a significant decision asks for the same three things at once. Historically, the instruments available to us have delivered two.

Depth. Rich, complete information about what each respondent thinks and why.

Scale. A representative sample large enough to generalize with confidence to the population you are trying to understand.

Rigor. Results you can trust and compare across people and across time, produced by a method you can defend.

Structured interviews deliver depth and rigor, not scale. Surveys deliver scale and rigor, not depth. Focus groups and open-text fields give you pieces of each, never all three. Narrative Intelligence delivers all three at once.


The interview

When well-designed, the interview delivers two of the three demands. It gives us depth, the rich, behavioral specificity that comes from adaptive probing and the chance to follow a thread. When the protocol is properly structured, it also gives us rigor: scores anchored to operationalized constructs, comparable across respondents. What the interview cannot give us, regardless of how well it is designed, is the third demand: scale.

The scale limit cascades into two practical problems. The first is dependence on the interviewer: the quality of the data is bounded by the quality of preparation, the discipline of the protocol, and the skill of the person executing it. The second is miscalibration across interviewers: each brings their own bias, their own interpretation, their own style. Across ten interviewers, you get ten calibrations, and those differences accumulate into interviewer drift. Every interviewer added in an attempt to scale is a new source of variance.

One common workaround has been the focus group: buy a little scale by putting more respondents in the room at once. A trade that favors scale at the expense of hearing what's true.. You gain marginally more people per hour of interviewer time, but the depth you bought with the interview format is now compromised by social-psychological distortions that don’t show up in the transcript. Conformity pressure reshapes what gets said (Asch, 1951; Cialdini & Goldstein, 2004). Status and dominance dynamics suppress minority voices. Groupthink amplifies what is already shared and quiets what is not (Janis, 1972; Stasser & Titus, 1985). Respondents say different things in groups than they do alone. The focus group does not resolve the three-demands problem. It trades interview-grade depth for slightly more respondents and group-level contamination.

Structured interviews deliver depth and rigor for the few. Focus groups trade part of that depth for slightly broader reach. Neither reaches the scale the question requires for adequate.

The survey

Surveys invert the trade. Where the interview delivers depth and rigor for the few, the survey delivers scale and rigor across the many. Tens of thousands of respondents at a cost per response that makes large sample and population-level measurement possible. Standardized items with validated psychometric properties, letting us say with confidence that engagement rose four points, or that a division's inclusion score diverged from baseline. What they cannot deliver is the demand the interview owned: depth.

The depth gap cascades into two practical problems. The first is that a score tells you what but not why. Two respondents who mark the same point on the Likert scale may mean radically different things, and the survey has no way to ask. The second is that the survey can only find what the designer thought to ask for. If a question is not on the instrument, the signal is not in the data. If a question is poorly worded, the answer is anchored to the flaw, and the survey format gives no way to detect or repair it (Schwarz, 1999; Tourangeau, Rips, & Rasinski, 2000).

The field's workaround is to staple open-text boxes onto the survey. Open text asks respondents to do the interviewer's job: synthesize, prioritize, and articulate, unprompted. That work is hard, and engagement with open-ended items is self-selected, skewed toward the motivated (Holland & Christian, 2009) and the aggrieved (Poncheri et al., 2008).

Surveys deliver scale and rigor across the many. Open text recovers only a sliver of depth, and only from those willing to write. Neither delivers depth at scale.

Narrative Intelligence

What has been impossible until now is measurement with the depth of an interview, the reach of a survey, and the rigor of a psychometric instrument at once. Savo's Narrative Intelligence platform makes it possible: an end-to-end system for designing, conducting, and analyzing AI-facilitated interviews. Each interview is engineered to close a specific intelligence gap, runs as an adaptive conversation at survey-grade scale, and produces decision-grade evidence: results scored against pre-defined constructs, with every score traceable to the respondent's own words.

The three demands, resolved

Depth. Adaptive, AI-facilitated probing applies the interview's defining strength, follow-the-thread elicitation by a skilled interviewer, to every respondent in the sample. AI removes the limit that once made this kind of elicitation a human-only craft.

Scale. Interview research has historically been capped at small samples by a single constraint: one human interviewer per respondent. An AI interviewer is not bound by that constraint. Each respondent receives the one-to-one attention of a skilled interviewer, in parallel with thousands of others, without the group-level contamination of focus groups or the closed-ended survey's inability to probe beyond its own question list.

Modern AI now makes depth and scale buildable together. A large language model fluent enough to sustain a conversation, paired with basic prompt scaffolding for context and guardrails, can probe adaptively and run in parallel across thousands of respondents. Any AI tool with a competent build can claim these properties. What good engineering does not deliver is rigor. Conversational AI without measurement science produces outputs that are fluent, confident, and plausible, but that lack the properties required for measurement: accuracy, comparability across respondents and over time, and defensibility as the basis for a decision.

Rigor. For measurement, rigor means valid quantification: results that are accurate, comparable across respondents and over time, and defensible under real scrutiny. That takes techniques honed over a hundred years of measurement science, which most AI conversation tools skip. Narrative Intelligence incorporates these techniques in four areas.

Defining the construct. Construct operationalization specifies, in advance, what is being measured. Behavioral anchors specify what evidence in the respondent's language will count as signal. Together they are the precondition of construct validity (Cronbach & Meehl, 1955; Messick, 1995) and of reliability, the property that the same construct yields the same score across raters and across occasions (AERA, APA, & NCME, 2014). What "competitive pressure" means in one respondent's conversation has to mean the same thing in the next, or the scores are not comparable.

Structuring the elicitation. What's defined in advance must be captured through a structured approach: unstructured conversation produces data that doesn't generalize. Interview Modes specify the strategy appropriate to the kind of information being captured: memory, perception, self-knowledge, and classification each have their own (e.g., Fisher & Geiselman, 1992; Geiselman, Fisher, MacKinnon, & Holland, 1985). Facilitation guides specify how the AI probes turn by turn, adapting to what the respondent just said rather than following a fixed script, with a separate signal-quality monitor running in parallel to flag any gaps in evidence as the conversation unfolds. Meta-analytic evidence shows structured interviews produce roughly twice the accuracy of unstructured ones (Sackett, Zhang, Berry, & Lievens, 2022).

Producing the insight. Qualitative narrative becomes quantitative measurement through two disciplined steps. Calibrated scoring turns captured language into construct scores via a pre-specified rubric. The scoring methodology has been calibrated against human-scored reference sets and shown to produce reliable and valid scores. The Standards for Educational and Psychological Testing set the expectations for instruments that reports quantitative scores from qualitative material: documented procedures, rater calibration, and inter-rater reliability (AERA, APA, & NCME, 2014). The methodology satisfies these requirements. Evidence Gating keeps the system from confabulating when evidence is thin. The scoring system abstains, returning "insufficient evidence" as its own outcome, distinct from a low score (which would imply the construct is absent) or a neutral score (which would imply ambivalence).

Defending the result. Traceability points every insight and construct score back to the specific respondent language that produced it. The evidence chain is auditable at the sentence level. Separation of roles keeps the agent that conducted the conversation from being the agent that scores it. Each role is accountable for mitigating a different source of error.

Treated as a measurement instrument, each AI-facilitated interview is tailored, before any participant joins, to close a specific intelligence gap: the constructs to be measured are specified in advance with behavioral anchors that define what evidence counts as signal, facilitation guides structure the elicitation, and a priori hypotheses are registered for the claims the interview is meant to test. Once deployed, the same instrument runs across any number of participants with the protocol held invariant, produces insights and construct scores traceable to the language that produced them, and reports emergent patterns through the same structured procedure as construct scoring.

None of this was possible even two years ago. The foundation is new: language models fluent enough to hold an adaptive conversation and to read open-ended language at scale. On its own, that foundation produces fluent, plausible output. Turning it into evidence you can stake a decision on takes a disciplined layer on top, one that requires measurement expertise to build. The models are now common. The discipline and human expertise to turn their output into evidence are not.

What changes

Resolving the depth/scale/rigor trade-off changes what kinds of measurement work can be done at decision-grade.

Causal inquiry, the work of asking why something is happening, has historically come with a choice: a quantitative answer at scale that did not contain the why, or an explanatory answer from the few who could be interviewed at depth. With the trade resolved, the choice does not hold. When a key cohort score moves twelve points, the why now lives in the same evidence chain as the score, traceable to specific respondent language at the sentence level, across whatever sample the question requires.

Longitudinal tracking has had a quiet problem: when a number moves, the respondent's internal yardstick may have moved with it. Response shift, the literature's term for this drift, comes in three forms (recalibration, reprioritization, reconceptualization), which standard survey measures cannot readily separate from real change (Sprangers & Schwartz, 1999). With every score traceable to the respondent's own language, the drift becomes visible. When three quarters of trend data on engagement, transformation readiness, or brand position show a shift, you can see which themes intensified, which faded, and which were scored differently this time.

Defensible inference has historically required a choice: a survey's psychometric properties or an interview's depth of evidence, never both with the same instrument. With both now in one instrument, findings can stand up to skeptical scrutiny from leadership, a regulator, or the workforce it describes. Every score has documented behavioral anchors and calibrated scoring. Every claim points to the respondent language that produced it. When evidence is insufficient, the instrument abstains rather than inventing a plausible answer it cannot support.

Open exploration has historically required choosing depth with ten participants or scale across ten thousand. Adaptive probing and structured scoring now run together at the same sample size. Whether the research is a jobs-to-be-done study, a concept test that evolves as participants react, or a first look at a new market segment, respondents answer in their own words, the instrument follows the threads each participant opens, and whatever emerges is held to the same evidence standard as any score.

Decision-grade evidence used to require choosing which of depth, scale, or rigor to underspend on. That choice is no longer required. Measurement that meets the rigor a serious decision deserves can now be done at the speed and scale of conversation. Plausible is no longer the bar. Valid is.

References

Psychometric foundation

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. https://doi.org/10.1037/0003-066X.50.9.741

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. https://www.aera.net/publications/books/standards-for-educational-psychological-testing-2014-edition

Interview validity

Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994

Group dynamics

Asch, S. E. (1951). Effects of group pressure upon the modification and distortion of judgments. In H. Guetzkow (Ed.), Groups, leadership and men: Research in human relations (pp. 177–190). Carnegie Press.

Cialdini, R. B., & Goldstein, N. J. (2004). Social influence: Compliance and conformity. Annual Review of Psychology, 55(1), 591–621. https://doi.org/10.1146/annurev.psych.55.090902.142015

Janis, I. L. (1972). Victims of groupthink: A psychological study of foreign-policy decisions and fiascoes. Houghton Mifflin.

Stasser, G., & Titus, W. (1985). Pooling of unshared information in group decision making: Biased information sampling during discussion. Journal of Personality and Social Psychology, 48(6), 1467–1478. https://doi.org/10.1037/0022-3514.48.6.1467

Survey limitations

Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54(2), 93–105. https://doi.org/10.1037/0003-066X.54.2.93

Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The psychology of survey response. Cambridge University Press. https://doi.org/10.1017/CBO9780511819322

Open-ended response bias

Holland, J. L., & Christian, L. M. (2009). The influence of topic interest and interactive probing on responses to open-ended questions in web surveys. Social Science Computer Review, 27(2), 196–212. https://doi.org/10.1177/0894439308327481

Poncheri, R. M., Lindberg, J. T., Thompson, L. F., & Surface, E. A. (2008). A comment on employee surveys: Negativity bias in open-ended responses. Organizational Research Methods, 11(3), 614–630. https://doi.org/10.1177/1094428106295504

Cognitive interview

Fisher, R. P., & Geiselman, R. E. (1992). Memory-enhancing techniques for investigative interviewing: The cognitive interview. Charles C. Thomas.

Geiselman, R. E., Fisher, R. P., MacKinnon, D. P., & Holland, H. L. (1985). Eyewitness memory enhancement in the police interview: Cognitive retrieval mnemonics versus hypnosis. Journal of Applied Psychology, 70(2), 401–412. https://doi.org/10.1037/0021-9010.70.2.401

Response shift

Sprangers, M. A. G., & Schwartz, C. E. (1999). Integrating response shift into health-related quality of life research: A theoretical model. Social Science & Medicine, 48(11), 1507–1515. https://doi.org/10.1016/S0277-9536(99)00045-3