Battle Staff Ride participants visit the Florence American Cemetery and Memorial in Florence, Italy, on Sept. 24, 2020. The three-day battle staff ride consisted of short historica. How to Analyze Open-Ended Survey Responses Without Losing Your Mind
Photo by U.S. Army photo by Spc. Meleesa E Gutierrez on Wikimedia Commons, Public domain

Operations

How to Analyze Open-Ended Survey Responses Without Losing Your Mind

To analyze open-ended survey responses, watch the mistakes that hide: an early codebook, an unchecked residual bucket, and counts read as customer shares.

What to take away

  • The costliest coding failures start with a codebook that closes too early and never gets reopened.
  • A residual bucket that swallows rare, severe complaints hides longest, often for two or three quarters.
  • Two coders agreeing on easy sentences is not reliability; without a disagreement check, theme shares measure the coder.
  • Reports should carry the codebook version, the coder count and the size of the residual bucket.
  • Prevention is cheap: read the whole residual each wave, log decisions, and store verbatims apart from identifiers.

Open-ended questions collect the part of a survey that nobody pre-coded, and the overview of open-ended questions explains why they resist tidy counting. A US consumer survey can return a few thousand comments in plain English, with typos and three complaints inside one box. The mistakes below survive review because each one looks like good practice at the time.

The costly one

One analyst builds the codebook from the first 50 responses, declares it finished, and codes the rest with a residual bucket labeled Other.

The bucket holds 5 to 8 percent of responses and looks trivial, so nobody opens it. For two or three quarters the theme shares move slightly and the report reads as stable. Then a pattern of billing complaints turns out to have sat in Other since the first wave. This mistake stays invisible for months, and the earlier reports cannot be corrected after the fact.

Prevention: read every response in the residual bucket before publishing, and print its size in the report. If Other runs above roughly one in ten responses, the codebook is not finished. The skill underneath is ordinary, and qualitative research covers the version that holds up under deadline.

The ones that look fine at first

Two coders, no overlap sample. Two people split 2,000 responses and a spot check shows agreement, because they agree on easy first clauses and split on hedged ones. Theme shares then move with staffing. Prevention: a fixed overlap sample of at least 10 percent, with disagreements counted.

One response, one code. A comment naming price and wait time gets a single tag, so counts understate both issues and the smaller one disappears. Prevention: allow up to three codes per response and mark one as primary.

Themes read as customer shares. People with a complaint type more often than contented ones, so 12 percent of comments gets quoted as 12 percent of customers. Self-selection starts with wording, and the notes on market research surveys cover how phrasing decides who answers. Prevention: write share of comments, and keep the base in the same sentence.

Signals worth checking each wave:

  • Every response carries exactly one code.
  • Theme shares are written as percentages of customers.
  • Nobody can name the codebook version behind last quarter's numbers.

The ones that only show up later

Codebook drift across waves. Nobody versions the codebook, so a label like fees narrows in March and widens in June. The failure stays invisible for months, until a chart of Q1 against Q3 shows movement that is really coder turnover. Primary market research treats that paper trail as part of the method. Prevention: version the codebook and stamp every wave file.

Idiom and translation. Labels written in standard English get applied to responses from Spanish-dominant respondents, and whole categories land in the wrong bucket. This surfaces only when someone rereads raw rows. Prevention: code in the language of the response, then translate the labels.

Verbatims kept beside identifiers. Free text often contains names and account numbers. The mistake stays quiet until a deletion request or partner review arrives months later. Prevention: separate files and a retention schedule. Federally funded human subjects work follows 45 CFR 46 on consent and review, and the same wording usually governs what a commercial panel may quote.

Mistake Surfaces after First symptom
Unversioned codebook Two or three quarters A trend that tracks coder turnover
Ignored residual bucket Months Complaints sitting in Other
One code per response Same day, quietly Understated counts
Verbatims beside identifiers Months, at audit Rework and re-quoting

What they have in common

Each failure above takes the raw text out of view too early and replaces it with numbers nobody can audit. A decision log fixes most of it: one line per judgment call, with a date and a name. None of this needs dedicated software, and market research analysis makes the same point about credibility sitting in the record rather than the tool. Skipping that record is what turns a working study into a dispute six months later.

Common questions

How many responses before coding beats reading? Below about 30 per wave, read everything and quote directly. Past 50, a draft codebook saves time, as long as the residual bucket gets read.

Should two people code every response? No. Code once, then have a second reader code a fixed overlap sample of about 10 percent and count the disagreements.

Can software tag themes for me? It can group similar text. It does not know your categories, so a human still reads the residual each wave.

What if the residual bucket keeps growing? The categories came from an early read. Split the bucket, name the patterns inside it, and version the codebook.

More in Operations

Operations

What Is Conjoint Analysis? A Plain-English Guide for Consumer Researchers

Conjoint analysis in consumer research turns choices between bundles into decision weights, then shows you when those weights are worth acting on.

Latest from Method Desk

Strategy

Focus Groups vs. In-Depth Interviews: Which Fits Your Consumer Research?

Focus groups versus in-depth interviews: how US consumer and B2B researchers pick a qualitative method, with a criteria table and where each one wins.