Skip to main content
Knowledge Base
concepts7 min read

How Narrative Identity Is Measured

Identity assessment through life stories: open-ended prompts, answers kept word for word, a coding manual, two independent raters.

#identity assessment#narrative coding#life story interview#coherence#research methods

Someone tells you about the worst year of their life. It takes twenty minutes. There is a job that ended badly, a parent who got sick, a drive home in the rain where something finally gave way. You listen. By the end you understand something about this person that no questionnaire would have handed you.

Now put a number on it.

That is the problem narrative psychology had to solve before identity assessment could call itself empirical, and the answer is unglamorous. Narrative identity is measured by asking open-ended questions about specific scenes from a life, capturing the answers word for word — transcribed from an interview, or typed by you — and having trained raters score that text against a written manual that defines every feature in advance (Adler et al., 2017). A second rater scores the same material independently, so the agreement between them can be checked and reported.

No scales. No self-report. Text, fixed definitions, and two people trained until they stop disagreeing.

The questions come first, and they are deliberately narrow

The best-known protocol opens by asking you to imagine your life as a book and divide that book into chapters. Then it asks for specific scenes — a high point, a low point, a turning point, an earliest memory, scenes from childhood, adolescence and adult life (McAdams et al., 2001). The full interview runs one to two hours and produces an enormous volume of text (Adler et al., 2017).

Three of those prompts have received extensive research attention: high point, low point, turning point. A study that wants narrative data without aiming at the whole life story is advised to collect those three, as long as they fit the question being asked (Adler et al., 2017).

The prompt is not a conversation starter. It is a stimulus, and it has to arrive in the same form for every participant. What a well-made one contains is its own subject.

That constraint matters more than it sounds. The measurement has not started yet, and the design has already decided what will be measurable.

Why nobody simply asks you to rate your own story

There is an obvious shortcut. You told the story, so why not score it yourself? Rate your turning point for how much you grew and how much it changed you.

Researchers tried. Panattoni and McLean report that when you rate your own account, your rating may not line up closely with the score a trained coder gives the same text. Two readings of that mismatch are available: either one of the raters is unreliable, or the two methods are reaching different things.

Panattoni and McLean take the second. McAdams says he agrees with them — while adding that it is too soon to know how far that reading holds (McAdams, 2018).

McAdams draws the distinction precisely (McAdams, 2018). When you rate your own account, you are rating the memory — everything you recall about that night, including what you never said out loud. The coder cannot reach any of it. The coder has only the text. That looks like a handicap, and in one sense it is. But narrative identity is not the memory; it is the story told about the memory, and the coder is the one looking straight at the story.

The coder's disadvantage is the entire point. Restricted to what you actually said, a coder measures your story rather than your private impression of it.

What the coder is actually marking

Open a coding manual and the surprise is how mechanical it reads.

Redemption. The code marks an explicit move in the account from a decidedly negative emotional state to a decidedly positive one, or to a positive outcome (McAdams et al., 2001). The negative state has to be clear and stated — suffering, pain, fear, grief. The scene either takes the point or it does not.

Contamination. The same rule reversed. A clearly positive scene must be followed by a clearly negative outcome, stated explicitly, or it scores zero.

Coherence. In Baerger and McAdams' system it splits into four separate indices, each rated on a seven-point scale (Baerger & McAdams, 1999). Orientation asks whether the account introduces its characters and places itself in time. Structure asks whether the scene contains an initiating event, an internal response, an attempt and a consequence, ordered so that each plausibly leads to the next. Affect asks whether the narrator makes an evaluative point rather than reciting facts. Integration asks whether contradictions get reconciled and connected to a larger life theme.

Habermas and Bluck mapped the same territory differently, naming four kinds of global coherence a life story can carry: temporal, causal, thematic, and one they call the cultural concept of biography — the sequence of events a given culture treats as a normal life (Habermas & Bluck, 2000).

Notice what none of these ask. Not whether the year was genuinely terrible. Not whether the narrator handled it well. Only whether the text says the state was negative, whether one event is linked to the next by cause, whether the account reconciles what it contradicts. The coder is reading structure, not judging a life.

Two coders, and the number between them

Narrative coding cannot be done by one person. Adler and colleagues are blunt about why: a single coder makes it impossible to determine how much interrater agreement exists, and they call the training phase crucial to the scientific soundness of the coding (Adler et al., 2017).

So two raters train together on a portion of the data — often ten to twenty-five percent (Adler et al., 2017) — then score a fresh subset independently and compare. They rarely agree well enough on the first attempt. They discuss the disagreements, sharpen the manual, and go again. Once agreement holds, they divide the remaining material, with periodic re-checks for what the field calls drift: the slow private redefinition of a code inside one rater's head.

In the redemption study, both coders were blind to identifying information about participants, and on the contamination code a third trained coder — also blind — settled the disagreements (McAdams et al., 2001).

Interrater agreement demonstrates that two trained readers applied the same definition to the same text. It says nothing about whether that definition was worth defining.

The reason this stayed inside universities

Add it up. One to two hours of interview per person (Adler et al., 2017), and two to three in the study that built the coherence indices (Baerger & McAdams, 1999). Verbatim transcription when the answers are spoken, done professionally or double-checked by hand, because one dropped "not" reverses a sentence. De-identification before anyone reads (Adler et al., 2017). Then two trained raters working through the entire dataset one construct at a time.

Adler and his coauthors name it the standing limitation of the whole approach: collecting and coding narratives is time and labor intensive (Adler et al., 2017).

The barrier was never conceptual. It was arithmetic. A method costing hours of expert attention per person can produce excellent findings across a few hundred people and never reach anyone else.

What this does not establish

Several things, and they should be said plainly.

These coding systems were built and validated on modest, specific samples. Baerger and McAdams developed the coherence indices with fifty adults in a single American city (Baerger & McAdams, 1999). The redemption work draws on American midlife adults and American undergraduates (McAdams et al., 2001). A redemption code marks a direction — bad to good — and is not a grade for a life or a marker of health. Habermas and Bluck name one of their four coherence types after cultural expectation itself (Habermas & Bluck, 2000), which means part of the yardstick is a local norm about how a life is supposed to go.

Whether automated coding reproduces what trained human coders produce on these constructs is an open question. This project has not answered it and has not been independently validated as an instrument. Establishing that would take a study of exactly that shape — machine scores against trained human scores, on the same transcripts, at scale.

Nothing in a coded transcript describes a condition, and none of this is diagnostic.

What an identity assessment can actually tell you

The machinery isn't mysterious, and that's the useful part. An identity assessment built on this method can only report what is present in the words you gave it. A code for causal linking means causal language appeared — nothing more. A low coherence score on one scene means that scene, told on that day, was missing the pieces the manual asks about. It is a mirror held up to a transcript, not a verdict on a life.

So tell the hardest scene the way it actually went, including the parts that still do not resolve. The method is built to read what is there — and the scenes you smooth over are the ones it will have nothing to say about.

Answer in your own words

A structured interview asks about specific episodes in your life, and returns a written report on how each account is built.

Last updated: 2026-08-16