Designing a Trademark Survey That Survives: A Methodology Checklist
By Casey Scott McKay ·
A trademark survey is the rare witness that can speak for thousands of consumers at once, and the rare witness whose entire testimony can be excluded before trial because of a choice made months earlier. This methodology checklist walks counsel and survey experts through building a confusion survey that survives cross-examination and a Rule 702 or Daubert challenge: gatekeeping the decision to survey at all, defining the load-bearing universe for the actual theory of confusion (forward, reverse, sponsorship, or dilution), selecting the Eveready or Squirt format, engineering a control cell so you report net rather than raw confusion, drafting neutral stimuli with marketplace realism, administering double-blind, sizing the sample, and coding verbatims objectively. It maps the sister formats for secondary meaning, fame, and genericness (Teflon and Thermos), and shows how each design choice doubles as a future cross-examination question. WHY notes, trap warnings, and worked examples accompany every phase, and a closing interpreter's guide explains what a net number actually means among the Polaroid and Sleekcraft factors. Survey design is fact-specific; engage a qualified survey expert and trademark counsel, such as Rightsy's virtual trademark attorneys, before you build.
Intellectual Property → Trademark Litigation | Published 28 June 2026 | rightsy.io
A consumer survey is the strangest witness in a trademark trial. Every other witness speaks for one person: what she saw on the shelf, what he meant to buy, what they remember from the ad. A survey claims to speak for thousands of people at once, compresses their collective state of mind into a single percentage, and delivers it with the quiet authority of arithmetic. Done right, it is the closest thing trademark law has to a window into the consumer's head, the one piece of evidence that can answer the decisive question in an infringement case directly rather than by inference: are people actually confused?
But the survey is also the only witness whose entire testimony can be thrown out before the jury hears a syllable, and thrown out not because the witness lied, but because of a choice somebody made in a conference room six months earlier. A leading question drafted on a Tuesday, a universe defined too broadly in a kickoff call, a missing control cell nobody budgeted for: any one of them can convert a six-figure scientific instrument into an exhibit the court strikes on a paper motion.
That paradox is what this checklist is built around. The phrase in the title is deliberate. We are not designing a survey that merely exists, or one that produces a headline number a press release can use. We are designing a survey that survives: survives the motion to exclude under Federal Rule of Evidence 702 and Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993); survives a cross-examiner who has read every word of the questionnaire backward; and survives a judge who has seen a hundred surveys and trusts none of them on faith.
Here is the organizing secret of survey design, and the reason the phases below run in the order they do: the sequence in which you make design decisions is the sequence in which the other side will attack them. A cross-examiner does not start with your coding protocol. She starts with whom you interviewed, then why you showed them what you showed them, then what you compared it against, then how you asked. Every checkbox in this document is therefore also a future deposition question. Tick it now, in the calm of the design phase, or answer it later under oath, in the worst possible light.
Two framing reminders before we build. First, a survey is one factor, not the whole case. Confusion is proven through the full multifactor inquiry, the Polaroid set in the Second Circuit and the Sleekcraft set in the Ninth, and a strong survey is powerful evidence of the "actual confusion" factor and a useful proxy for the "strength of mark" and "consumer sophistication" factors, but it is never the entire ballgame. For the larger picture, see Running the Likelihood-of-Confusion Analysis and Likelihood of Confusion: A Brand Owner's Field Map. Second, a survey is never mandatory. Plenty of cases are won with anecdotal actual confusion, the parties' own documents, and the other factors. The decision to commission one is itself a strategic choice, and the first phase of this checklist is about making it honestly.
A practical note on groundwork. Before you can define a universe or build a realistic stimulus, you have to know what the contested marks actually look like at the point of sale, who is really selling under them, and who owns the rights on each side. A clearance-grade trademark and logo search on Rightsy, paired with a chain-of-title check in Rightsy's assignment records, grounds the whole exercise in marketplace reality rather than in the lawyers' assumptions about it. Surveys fail far more often from a bad picture of the market than from bad statistics.
Phase 0 — Decide whether to commission a survey at all
This is the phase everyone wants to skip. Skip it at your peril.
- [ ] Confirm the disputed issue genuinely turns on consumer perception rather than on a pure question of law or undisputed fact.
- [ ] Confirm the budget and the calendar can absorb a rigorous design. A defensible confusion survey is usually a multi-month, five- or six-figure undertaking, not a weekend SurveyMonkey poll.
- [ ] Confirm that a competent, independent expert actually believes a defensible survey can be built for these marks in this market, and that it stands a real chance of finding what your theory predicts.
- [ ] Retain the survey expert early, before the pleadings harden and before anyone has promised a number to the client.
WHY. Surveys are persuasive precisely because they are expensive and rigorous; that same rigor is what makes a half-built one dangerous. Courts and juries give weight to surveys because they look scientific, which means a sloppy survey borrows credibility it has not earned, and the borrowed credibility cuts against you the instant the cross-examiner exposes the flaw.
TRAP — the survey that finds nothing. A mediocre survey is frequently worse than no survey. Run a weak design, get a low confusion number, and you have manufactured your opponent's best exhibit at your own expense. The defense will wave your own 4% net confusion figure at the jury for the rest of the trial. If a credible expert tells you, in good conscience, that she cannot build a survey likely to find meaningful confusion on these facts, that is not a failure. That is data about your case. Listen to it.
TRAP — the late expert. Retaining the expert after you have filed a complaint that pleads a specific theory of confusion, or after you have taken positions in discovery that define the market a certain way, can box the survey into a universe that no longer fits. The expert should help shape the theory, not inherit a frozen one.
FIELD NOTE. Whether a survey is worth it often depends on the forum. A bench trial before a trademark-fluent judge, a jury trial, a preliminary-injunction sprint, and a Trademark Trial and Appeal Board proceeding all reward different things. Surveys are admissible and used at the TTAB but are comparatively rare there given cost and the Board's expertise; if your fight is an opposition or cancellation, weigh that before spending. See The TTAB Practice Toolkit and, on the forum's effect on every evidentiary choice, Judge or Jury: Choosing Your Factfinder.
Phase 1 — Define the universe (the load-bearing wall)
If there is one decision that holds up the entire structure, it is this one. Define the wrong population to interview and nothing built on top of it can be salvaged. You did not measure confusion badly; you measured the wrong people's confusion, and no clever cross-examination repair, no eloquent closing argument, no re-weighting can fix a foundation poured in the wrong place.
- [ ] Pin down the theory of confusion in writing before anything else: forward confusion, reverse confusion, confusion as to sponsorship or affiliation, post-sale confusion, or initial-interest confusion.
- [ ] For forward confusion, set the universe to the prospective purchasers of the junior user's (typically the defendant's) goods or services. The question is whether the newcomer's customers are confused into thinking they are getting, or getting something approved by, the senior brand.
- [ ] For reverse confusion, flip it: the universe is the prospective purchasers of the senior (plaintiff's) goods or services. The harm is that a big junior user has so saturated the market that the senior user's own customers now think the small original is the copycat.
- [ ] For sponsorship or affiliation theories, define the universe by the goods on which the confusion would occur, and frame the question around approval, connection, or permission rather than identical source.
- [ ] For dilution (fame), sample the general consuming public of the United States (more on this below).
- [ ] For secondary meaning, sample the relevant purchasers of the goods at issue.
- [ ] Draft screening (filter) questions that actually capture that population: correct product category, realistic purchase timeframe (past and/or prospective), and the right price band and channel.
WHY. The universe is the population your sample is supposed to represent. Get it right and a clean sample of a few hundred people can speak for millions. Get it wrong and the most elegant questionnaire in the world is describing strangers. Courts treat a misdefined universe as a foundational, often fatal, defect rather than a quibble that goes to weight; the canonical statement is Bristol-Myers Squibb Co. v. McNeil-P.P.C., Inc., 973 F.2d 1033 (2d Cir. 1992), where the survey of the wrong population could not support the inferences drawn from it.
TRAP — over-inclusive and under-inclusive universes. Two named errors, both lethal. An over-inclusive universe sweeps in people who are not real prospective purchasers (interviewing the general public about a $40,000 industrial 3D printer), which dilutes and distorts the result. An under-inclusive universe excludes relevant buyers, so the result cannot be generalized to the market that matters. Either way, a sophisticated opponent will probe your screeners line by line in deposition, because loose or tight screeners are the back door through which a universe problem walks right back in after you thought you had closed the front.
TRAP — letting the format pick the universe. The format you choose (next phase) does not change the universe. A Squirt survey run on the wrong population is exactly as broken as an Eveready survey run on the wrong population. Universe first, always; format second.
Worked example — forward confusion (FATHOM). Suppose FATHOM is a well-known maker of premium dive watches, an arbitrary, commercially strong mark. A newcomer launches FATHOMLESS smartwatch bands and accessories. Under a forward-confusion theory, the universe is prospective purchasers of FATHOMLESS bands, not FATHOM watch buyers, because the alleged harm is that the newcomer's customers think the senior watchmaker stands behind the bands. Screeners should capture people who have bought or intend to buy smartwatch accessories in a realistic window and price range.
Worked example — reverse confusion (VAULT). Now suppose VAULT is a tiny regional scheduling app sold to a handful of local gyms. A national fitness conglomerate rolls out a VAULT workout platform with a Super Bowl ad and blanket sponsorships. The senior user's grievance is reverse confusion: its own prospective customers now assume little VAULT is an offshoot of, or rip-off of, the famous newcomer. The universe therefore flips to prospective purchasers of the senior (small) VAULT's services. Reverse-confusion theory traces to Big O Tires, Inc. v. Goodyear Tire & Rubber Co., 561 F.2d 1365 (10th Cir. 1977); the universe flip is the single most-missed move in this corner of survey design. If priority and ownership are contested, confirm who actually holds the senior rights with a chain-of-title check in Rightsy's assignment records before you build a universe around "the senior user."
Phase 2 — Choose the format that fits the marks
Two workhorse formats dominate confusion surveys, and choosing between them is choosing how closely your laboratory will mimic the way real consumers encounter the marks. The choice is not a matter of taste or habit. It is dictated by the strength of the senior mark and by whether shoppers genuinely meet the two marks together in the wild.
- [ ] Choose the Eveready format when the senior mark is strong or famous. Show respondents only the junior mark or product, on its own, and ask open-ended questions from memory: Who do you think puts out this product? What makes you say that? What other products do you think this company makes? Do you think they needed permission or approval from anyone to put this out?
- [ ] Choose the Squirt format when the senior mark is weaker, or when consumers genuinely encounter both marks together in the real marketplace. Show both marks and ask a neutral same-source or affiliation question.
- [ ] For any Squirt design, build a lineup or array that salts the two contested marks among several unrelated marks, so respondents are not staring at a naked two-way matchup that begs them to find a connection.
- [ ] Match the format to the facts on the ground, and write down why you chose it.
WHY — Eveready. The Eveready format, from Union Carbide Corp. v. Ever-Ready, Inc., 531 F.2d 366 (7th Cir. 1976), is widely regarded as the gold standard because it mirrors how confusion actually happens for a strong brand: a shopper sees the junior product alone, and the senior mark is so embedded in memory that it surfaces unprompted. Because it relies on the respondent's own recall and shows nothing suggestive, it minimizes the risk that the survey created the very association it claims to measure.
TRAP — Eveready on an obscure mark. Eveready has a built-in floor: it only counts confusion among people who already carry the senior mark in their heads. Run it on a senior mark few consumers have ever heard of and it will under-count confusion for the wrong reason, because respondents cannot name a brand they do not know, even if the marks are genuinely confusable side by side. That is a mismatch between tool and fact, not proof of no confusion.
WHY — Squirt. The Squirt format, from SquirtCo v. Seven-Up Co., 628 F.2d 1086 (8th Cir. 1980), shows both marks and is appropriate where consumers really do see them in proximity, on the same shelf, in the same search results, in the same category. It rescues cases where Eveready would unfairly under-count a weaker senior mark.
TRAP — Squirt's suggestion problem. Showing two marks side by side can manufacture a relationship no real shopper would ever perceive, simply because the human mind, handed two things at once, goes looking for how they connect. This is the demand effect at the level of the whole design. A bare two-way comparison is one of the most exploitable artifacts in trademark litigation, which is why the salted array exists and why the control cell (next phase) has to be especially muscular for Squirt. If a plainly famous senior mark's expert quietly chose Squirt over Eveready, the most common backstage reason is that Eveready did not show enough confusion. Expect the cross-examiner to ask, and expect the judge to wonder.
Worked example — Squirt done honestly (EMBERHOUSE vs. EMBER & ASH). Two regional hot-sauce makers, EMBERHOUSE (senior) and EMBER & ASH (junior), sit in the same refrigerated specialty-foods cooler in the same stores. Consumers really do encounter both at once, so Squirt fits the facts. The honest version shows both bottles inside an array of four or five other hot sauces of similar style, photographed as they appear on the shelf, and asks a neutral affiliation question. The dishonest version puts the two bottles alone on a white background and asks, in effect, "don't these go together?" Same format name, opposite credibility.
Phase 3 — Build the control and plan the net
Here is a number that means nothing by itself: "34% of respondents were confused." Here is a number that means something: "the test cell showed 34% confusion, the control cell showed 9%, for net confusion of 25%." The difference between those two sentences is the difference between a survey that survives and one that gets shredded, and it all lives in the control.
- [ ] Build a control cell identical to the test cell in every respect except the one allegedly infringing feature. Same universe, same screeners, same format, same questions, same administration. The only thing that changes is the thing in dispute.
- [ ] Choose the control stimulus thoughtfully: a plainly non-infringing alternative mark, or the test stimulus with the contested feature altered or removed, so the control isolates the effect of that feature and nothing else.
- [ ] Commit, in writing and in advance, to reporting net confusion = test cell minus control cell, not the raw test figure.
- [ ] Make the control especially robust for Squirt designs, where juxtaposition inflates raw confusion and the control is doing the heavy lifting of subtracting that inflation back out.
WHY. Every survey generates background noise: people guess, people yea-say to be agreeable, people carry pre-existing beliefs into the room, and the instrument itself nudges. The control measures that noise floor by running the same gauntlet without the infringing element. Subtract it and you isolate the confusion actually attributable to the defendant's mark, the way a clinical trial subtracts the placebo response to isolate the drug. The first question a sophisticated reader, or a sharp judge, asks about any confusion survey is not "what was the rate?" It is "what was the control, and what is the net?" If your answer is "there was no control," you have already lost the witness.
TRAP — the controlless Squirt. A Squirt survey with no control is close to uninterpretable, because the format's own suggestiveness pumps up the raw number and there is nothing to subtract it back down. Judges and good cross-examiners know this cold. The thorough opinion in Simon Property Group, L.P. v. mySimon, Inc., 104 F. Supp. 2d 1033 (S.D. Ind. 2000), is a master class in how a court walks through control design and net confusion, and it rewards a careful read for anyone building one of these.
TRAP — the strawman control. A control so obviously different from the test stimulus that it would never confuse anyone produces an artificially low control number and an artificially high net. The cross-examiner will argue you rigged the baseline. The control should be a fair, realistic non-infringing comparator, not a punching bag.
Worked example. Back to FATHOMLESS bands against famous FATHOM watches. The test cell shows the actual FATHOMLESS product. The control cell shows the same kind of product under a coined, plainly unrelated name (say, "OAKHELM") with the same look, price, and packaging style. If the test cell yields 31% "made or approved by the watch company" and the control yields 8% (the people who would attribute any upscale band to a watchmaker), your net is 23%, and you can defend every point of it.
Phase 4 — Draft the questions, build the stimuli, run it blind
This is the craftsmanship phase, where good intentions die in the wording. The governing principle is restraint: the questionnaire must extract the respondent's genuine, unguided belief without ever hinting at the answer you are hoping for.
- [ ] Use open-ended, non-leading questions. "Who do you think makes this?" not "Do you think Brand X makes this?"
- [ ] Avoid demand effects: any cue in wording, order, or emphasis that telegraphs the "right" answer.
- [ ] Offer a genuine "don't know / no opinion" option, and make it as easy and respectable to choose as any other answer, so that uncertainty is recorded as uncertainty rather than miscoded as confusion.
- [ ] Probe spontaneous answers neutrally: "What makes you say that?" with no follow-on nudging.
- [ ] Rotate the order of stimuli and of answer choices across respondents to cancel out order and primacy effects.
- [ ] Replicate marketplace realism in the stimulus: full trade dress, packaging, color, logo treatment, and point-of-sale context as the consumer truly encounters it, not a bare word mark in black Helvetica on a white screen.
- [ ] Administer double-blind: neither the respondent nor the interviewer (live or programmed) knows who sponsored the survey or which outcome the sponsor wants.
WHY — neutral wording. A leading question generates confusion that exists nowhere but inside the instrument. Ask "Do you think the same company makes both of these?" and you have planted the idea of one company; some fraction of agreeable respondents will salute it regardless of the marks. The open-ended form lets confusion surface on its own or not at all, which is the only kind a court cares about.
WHY — the "don't know" valve. Without an easy escape hatch, respondents who genuinely have no idea will guess, and guesses contaminated by the test stimulus often land on the plaintiff. A real "don't know" option keeps non-opinions out of your numerator. (It also models honesty for the court: a survey unafraid of "don't know" looks like one built to find truth, not a number.)
WHY — marketplace realism. The legal question is whether confusion is likely in the marketplace, so the further your stimulus drifts from the real shopping experience, the less your result measures the thing in dispute. An artificial side-by-side that no shopper would ever see measures something other than real-world confusion. This has special bite in trade dress cases, where the look and feel is the mark, and stripping it down to a logo on a card guts the very thing you are testing. Surveys have been excluded for stimuli unmoored from how consumers actually meet the products; THOIP v. Walt Disney Co., 690 F. Supp. 2d 218 (S.D.N.Y. 2010), is a useful cautionary tale.
WHY — double-blind. Single-blind is not enough. An interviewer who knows the desired answer can signal it without meaning to, through tone, pacing, an encouraging "mm-hm" at the right moment. Double-blind closes that channel, and just as important, it lets you tell the court, truthfully, that no human in the chain could have steered the result. The Reference Manual on Scientific Evidence (Federal Judicial Center), and within it Shari Seidman Diamond's Reference Guide on Survey Research, is the standard articulation of these controls and the document your expert should be able to cite from memory.
TRAP — the tidy white background. The single most common realism failure is presenting word marks alone, decontextualized, because it is cheap and easy to program. It also strips away exactly the marketplace cues that drive or dispel confusion. If the defendant's house brand, dress, or disclaimer would be visible at the point of sale, hiding it in your stimulus is not neutrality; it is a thumb on the scale your opponent will expose.
TRAP — the leading "filter that isn't." Sometimes the leading happens up in the screeners, where a poorly worded qualifier primes the category or the brand before the test question is ever asked. Audit the whole instrument for priming, not just the headline question.
Phase 5 — Sample and quantify
Now the statistics. A beautiful questionnaire administered to too few of the wrong people, with no honest accounting of error, is still a broken survey. This phase is where the result earns the right to be called a measurement.
- [ ] Choose and justify the sampling method. Probability sampling is the theoretical ideal; in practice, well-run non-probability approaches (validated online panels, properly executed mall intercepts) are widely accepted if the expert documents and defends the methodology and its limits.
- [ ] Size each cell, test and control, independently for a usable margin of error. Two hundred completes per cell is a common working target; the right number flows from the precision you need and the effect size you expect, not from the budget alone.
- [ ] Report the cell sizes, the margin of error, and the confidence level (commonly 95%). Do this proactively, in the report, before anyone asks.
- [ ] Validate completed interviews (re-contacts, attention checks, speeder and straight-liner removal) and document the disposition of every respondent who entered the funnel.
WHY. Sample size and margin of error are not decoration; they are the difference between a number and a guess. A net confusion figure of "20%" computed from 40 people per cell carries a margin of error so wide it could honestly be anything from trivial to overwhelming, and the cross-examiner will read your own confidence interval back to you. Reporting the error candidly is not a weakness to hide; it is the signature of a real measurement, and it inoculates you against the accusation of overstatement.
TRAP — the unjustified panel. Online panels are accepted, but "we used a panel" is not a methodology; how you sourced, screened, and validated respondents is. An undocumented panel of professional survey-takers who blow through the questionnaire for points is a gift to the defense. Build and keep the validation paper trail from the first respondent.
TRAP — the disappearing denominator. Be able to account for everyone: how many were invited, how many qualified, how many completed, how many were tossed and why. A survey that cannot reconstruct its own funnel invites the inference that inconvenient respondents quietly vanished.
Phase 6 — Code, document, and report
The last mile, and the place where mischief most easily hides, because coding open-ended answers is a human act of judgment performed on thousands of ambiguous sentences. "I think it's the watch people, but honestly I'm just guessing" is confused, uncertain, or both, depending entirely on who holds the pen and what rules they follow.
- [ ] Code open-ended responses objectively: a written coding protocol fixed in advance, blind coders who do not know which cell a response came from, and a measured inter-coder reliability check.
- [ ] Preserve and produce the verbatims. Every raw open-ended answer, kept and disclosed, so the other side and the court can audit your coding.
- [ ] Document the design rationale in writing as you go: why this universe, this format, this control, these questions. Contemporaneous reasoning is worth a great deal more than a litigation-eve reconstruction.
- [ ] Build the report to be the expert report. Assume it will be attached to a Rule 702 motion and pulled apart sentence by sentence, and write it accordingly.
WHY. Blind, protocol-driven coding with a reliability statistic is how you convert subjective interpretation into a reproducible measurement. Without it, the coding step becomes an unfalsifiable judgment call, and an unfalsifiable judgment call is not science; it is advocacy in a lab coat. Producing the verbatims is the ultimate proof of good faith: it lets the court check your work. A report that resists producing its verbatims is signaling, loudly, that the verbatims would not survive the daylight.
TRAP — coding toward the conclusion. When coders know which cell they are scoring, ambiguous answers drift, ever so slightly, toward the hoped-for result. Blind coding removes both the temptation and the appearance of it. If your inter-coder reliability is poor, do not paper over it; fix the protocol and recode.
TRAP — the orphaned methodology. A survey whose design choices were never written down until the expert report was drafted looks reverse-engineered, because it often was. Memorialize the why contemporaneously, in working memos and the expert's notes, so the rationale predates the result.
Phase 7 — Build it to survive Rule 702 and Daubert from day one
Admissibility is not a hurdle you clear at the end; it is a design constraint you honor from the first decision. The discipline is simple to state and hard to live: make every choice as though the judge will read the questionnaire word for word, because in a serious case, she will.
- [ ] Treat Federal Rule of Evidence 702 (as amended December 1, 2023) as the spec sheet. The proponent must show, by a preponderance of the evidence, that the opinion rests on sufficient facts, is the product of reliable principles and methods, and reflects a reliable application of those methods to the facts.
- [ ] Map your design to the recognized reliability factors for surveys (proper universe, appropriate format, neutral questions, adequate control, double-blind administration, sound sampling, objective coding, candid error reporting) and keep evidence for each.
- [ ] Anticipate the gatekeeping trilogy: Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993); General Electric Co. v. Joiner, 522 U.S. 136 (1997); and Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999), which extends gatekeeping to all expert testimony, surveys included.
- [ ] Mind the disclosure rules. A survey expert's full report and underlying data must be disclosed under Federal Rule of Civil Procedure 26(a)(2)(B), and a failure can trigger automatic exclusion under Rule 37(c)(1). Late or incomplete disclosure can lose you a perfectly good survey on a technicality.
- [ ] Keep Rule 403 in view. Even an admissible survey can be limited if its prejudicial or misleading potential substantially outweighs its probative value, a particular risk when a raw (uncontrolled) number might mislead a jury.
WHY. The 2023 amendment to Rule 702 was a course correction. For years, courts waved methodological problems through as going to "weight, not admissibility," letting the jury sort it out. The amended rule makes explicit what was always implicit: the proponent carries the burden, by a preponderance, and the court must find the method was reliably applied, not merely that the expert is credentialed and the method exists in the abstract. The practical upshot for surveys is that fundamental defects increasingly draw exclusion rather than a footnote. A wrong universe, an absent control, leading questions, an unrealistic stimulus: these used to be cross-examination fodder; now they are exclusion arguments. Designing for survival is no longer optional polish. It is the price of admission.
WHY — hearsay is handled, if you do it right. A confusion survey is technically a mountain of out-of-court statements, yet surveys come in routinely. Courts treat properly conducted surveys as admissible either as non-hearsay evidence of the respondents' present states of mind under Federal Rule of Evidence 803(3), or as the reliable basis for expert opinion under Rule 703. The throughline is trustworthiness: the more rigorously you followed the methodology in this checklist, the more comfortably the hearsay objection dissolves. Sloppiness, by contrast, reopens the door.
TRAP — the Daubert mirror. Everything that makes your survey survive a challenge is also a blueprint for attacking the other side's. The defensive checklist and the offensive checklist are the same document read in two directions. For the attack version, see Keeping the Survey Out: Daubert Challenges to Trademark Survey Experts. For the affirmative build, our companion piece Building a Bulletproof Consumer Survey in Trademark Cases goes deeper on the construction side.
The cross-examiner's checklist (read your survey through enemy eyes)
Before you call a survey finished, run it through the exact sequence a hostile expert and a hostile lawyer will use. If you cannot give a confident, documented answer to each, you have found your homework.
- [ ] "Whom did you interview, and why those people?" (universe) — Can you defend the population and the screeners that built it?
- [ ] "Why that format?" (Eveready vs. Squirt) — Does the format fit the mark's strength and the real marketplace, or did you pick the one that happened to score higher?
- [ ] "What did you compare it against?" (control) — Is there a fair control, and is the headline number the net?
- [ ] "How did you ask?" (leading questions, demand effects, the missing "don't know") — Walk through every question word by word.
- [ ] "What did they actually see?" (marketplace realism) — Does the stimulus look like the shelf, or like a lab?
- [ ] "Who knew the answer?" (blinding) — Single-blind, double-blind, or unblinded?
- [ ] "How many, and how sure?" (sample size, margin of error) — Are the cells large enough to mean anything?
- [ ] "Who coded the words, and by what rules?" (objective coding, verbatims) — Blind coders, written protocol, reliability check, produced verbatims?
- [ ] "Where did you write the rationale down?" (contemporaneous documentation) — Does the why predate the result?
Notice the order. It is Phases 1 through 7, run in sequence. The cross-examination is the design checklist. That is not a coincidence; it is the whole point.
The sister surveys: secondary meaning, fame, and genericness
Confusion is the headline use, but the same methodological discipline powers surveys aimed at other questions. The instrument is one discipline; only the target moves. Each of the following must still satisfy the universal commandments, right universe, neutral questions, blind administration, objective coding, before its specialized design matters at all.
Secondary meaning (acquired distinctiveness)
When a descriptive term or a product look can only be protected on a showing of acquired distinctiveness under Lanham Act § 2(f), 15 U.S.C. § 1052(f), a survey can test whether the relevant public has come to treat the term as identifying a single source. (For product design trade dress, secondary meaning is always required; Wal-Mart Stores, Inc. v. Samara Bros., Inc., 529 U.S. 205 (2000).)
- [ ] Universe: the relevant purchasers of the goods at issue.
- [ ] Test whether the term signifies one source ("one company" vs. "many different companies"), not merely whether respondents recognize the words.
- [ ] Control for the respondents who would attribute any term, descriptive or not, to a single source, so you net out reflexive single-source answers.
- [ ] Keep the wording neutral; do not feed respondents the brand name you hope they will produce.
Worked example (FRESHPRESS). A juice company claims FRESHPRESS for cold-pressed beverages, an obviously descriptive term, and must prove the market sees it as a brand rather than a description. A clean survey asks whether FRESHPRESS juice comes from one company or several, controls with a comparably descriptive non-claimed term, and reports the gap. For the doctrine this evidence serves, see From Descriptive to Distinctive: How Marks Acquire Secondary Meaning and The Abercrombie Spectrum.
Fame (for dilution)
Dilution under 15 U.S.C. § 1125(c) is reserved for marks that are famous, meaning "widely recognized by the general consuming public of the United States," a deliberately demanding standard the Trademark Dilution Revision Act set in 2006.
- [ ] Universe: the general consuming public, not a niche buyer pool. This is the rare survey that should not be narrowed to category purchasers, because the legal test is nationwide household-name recognition.
- [ ] Measure unaided and aided recognition with neutral prompts; document the sampling frame's national reach.
TRAP. Sampling category buyers for fame is a classic mismatch: you may prove a mark is well known to enthusiasts, which is precisely not the legal question. Niche fame is not fame for dilution. For where dilution sits among brand rights, see Trademark Overview: Infringement, Dilution, and Related Rights.
Genericness — the Teflon format
When the fight is whether a term has become the common name for the product itself (a generic), the primary-significance test asks what the term means to the relevant public. Congress codified that focus in the 1984 Trademark Clarification Act, directing courts to the primary significance of the mark to the relevant public and rejecting any "purchaser motivation" detour. The Teflon survey, named for E.I. DuPont de Nemours & Co. v. Yoshida International, Inc., 393 F. Supp. 502 (E.D.N.Y. 1975), is the preferred tool.
- [ ] Tutor respondents on the difference between a brand name (one company) and a common name (a type of product), using clear, neutral examples.
- [ ] Present a salted list mixing known brand names, known common names, and the disputed term, in rotated order, and ask respondents to classify each.
- [ ] Keep the tutorial scrupulously even-handed; a slanted lesson contaminates every later answer.
WHY courts prefer Teflon. The mini-tutorial plus forced classification produces cleaner, more interpretable data than open-ended approaches, which is why it has become the dominant genericness instrument. Its modern star turn came in USPTO v. Booking.com B.V., 591 U.S. 549 (2020), where survey evidence of consumer perception helped establish that even a "generic.com" term could function as a brand if the public so perceived it. The lesson cuts both ways: a well-built Teflon survey can save a mark from the genericide graveyard or push it in.
Worked example (GLIDEBOARD). A maker of self-balancing scooters wants to keep GLIDEBOARD from sliding into generic use. A Teflon survey tutors respondents, then asks them to sort a rotated list, KLEENEX-style known brands, plainly generic terms like "stapler," and the disputed GLIDEBOARD, into "brand" or "common name." The brand percentage, properly controlled, is the headline. On how a once-strong mark slides toward the public domain, see Use It or Lose It: How Trademarks Are Abandoned.
Genericness — the Thermos format
The Thermos format, from King-Seeley Thermos Co. v. Aladdin Industries, Inc., 321 F.2d 577 (2d Cir. 1963), comes at the same question from the other side: it asks what respondents would spontaneously call the product if they were shopping for one, capturing generic usage in the wild.
- [ ] Ask an open, unaided question: "If you were going to buy one of these, what would you ask the store for?"
- [ ] Code the spontaneous answers objectively, exactly as you would confusion verbatims.
- [ ] Consider running Teflon and Thermos together; convergent results from two different instruments are far harder to dismiss than either alone.
WHY. Thermos catches the consumer who reaches for the disputed word as the ordinary name of the thing, with no brand prompting at all. It is the linguistic snapshot of a mark in the act of going generic. Paired with Teflon's structured classification, it gives a court two independent angles on primary significance.
The interpreter's guide: what the number actually means
Suppose you did everything right and you have a clean net confusion figure. What does it prove? Less than the press release wants, and more than the defense will admit.
- [ ] Read the net number, never the raw one. (If anyone quotes you a raw figure, ask for the control.)
- [ ] Use the conventional, soft benchmarks as a frame, not a verdict: net confusion around 15% and up tends to support a likelihood of confusion; below roughly 10% tends to cut against it; the 10–15% band is a genuine gray zone where the other factors do the deciding.
- [ ] Treat the result as one factor within the full Polaroid (Second Circuit) or Sleekcraft (Ninth Circuit) analysis, weighed alongside mark strength, proximity of goods, intent, actual-confusion anecdotes, channels, and sophistication.
- [ ] Calibrate to the forum and the standard: what persuades at a preliminary-injunction hearing differs from what carries a jury or convinces a bench.
WHY the thresholds are soft. Those percentages are rules of thumb distilled from decades of cases (collected at length in McCarthy on Trademarks §§ 32:158 et seq.), not statutory lines. Courts have found likely confusion below 15% and rejected it above, depending on the marks, the goods, and the quality of the survey. A pristine 12% from a flawless design can outweigh a shaky 25% from a contested one. Methodology is destiny: a number is only worth the design that produced it. On how these factors play out at the dispositive-motion stage, see The Polaroid Factors at Summary Judgment in the Second Circuit. For where the survey sits in the broader campaign, from watching to verdict to appeal, see The Trademark Enforcement Toolkit.
FIELD NOTE — the modern survey moment. Surveys are not a relic. A confusion survey was squarely in the record in Jack Daniel's Properties, Inc. v. VIP Products LLC, 599 U.S. 140 (2023), the "Bad Spaniels" dog-toy dispute, and Booking.com turned in significant part on consumer-perception survey evidence. The instrument is alive, contested, and frequently decisive, which is exactly why building it to survive matters now more than ever.
Common fatal flaws, in one place
A field guide to the wreckage, in roughly the order an opponent will find it:
- The wrong universe. The single most common fatal flaw, and the least curable. You measured the wrong people.
- No control cell. The raw figure overstates confusion and cannot be netted; the result is uninterpretable.
- Format mismatch. Squirt on a famous mark to inflate the number, or Eveready on an unknown mark that guarantees a floor.
- Leading questions and demand effects. Confusion that exists only inside the instrument.
- No real "don't know" option. Non-opinions miscoded into the numerator.
- Artificial stimuli. Bare word marks on white, side-by-sides no shopper would ever see, dress and disclaimers stripped away.
- Single-blind or unblinded administration. A channel for unconscious steering, and an easy story for the cross.
- Tiny samples, unreported error. A "number" with a margin of error wide enough to mean anything.
- Subjective or non-blind coding; withheld verbatims. The place where bias hides and audits are refused.
- Reverse-confusion universe not flipped. Measuring the defendant's customers when the theory is about the plaintiff's.
- Late or incomplete Rule 26 disclosure. Losing a good survey to Rule 37(c)(1) on a calendar mistake.
A modular toolkit: reusable components you can lift into any design
Think of these as the standard parts you bolt onto every survey, regardless of the legal question. Keep templates of each so you are not reinventing them under deadline.
- [ ] Universe-definition memo. A one-page statement of the theory of confusion, the resulting population, and the screener logic that captures it, signed off by counsel and the expert before fieldwork. This memo is your first and best answer to the first and hardest cross-examination question.
- [ ] Screener battery. A reusable block of qualifying questions, category, timeframe, price band, channel, decision-maker status, with priming audited out.
- [ ] Control-design worksheet. A short rationale for the chosen control stimulus and an explicit statement that the reported metric is net confusion.
- [ ] Neutral question stem library. Pre-vetted open-ended source, affiliation, and permission questions, plus the standard "don't know" valve, so wording is consistent and defensible across matters.
- [ ] Stimulus realism spec. A checklist for the stimulus: full dress, color, packaging, point-of-sale context, rotation, and a note on anything deliberately included or excluded and why.
- [ ] Blinding protocol. A written description of how respondent and administrator are kept blind to sponsor and purpose.
- [ ] Sampling and power note. Target cell sizes, expected effect, margin of error, confidence level, and panel-validation steps.
- [ ] Coding protocol and reliability plan. The written code frame, blind-coder instructions, and the inter-coder reliability method, fixed before any answer is read.
- [ ] Documentation discipline. A running design log capturing each major decision and its rationale, contemporaneously, so the why always predates the result.
Reusing vetted components does more than save time. It builds a consistent, defensible house style across your matters, so that when an opponent's expert tries to paint a particular choice as ad hoc or outcome-driven, you can show it is your standard practice, applied the same way every time. Consistency is credibility.
A short worked case study, end to end
To see the phases lock together, follow one hypothetical from intake to interpretation.
The dispute. EMBERHOUSE, a regional craft hot-sauce maker, sues EMBER & ASH, a newcomer whose bottles sit in the same specialty coolers. EMBERHOUSE has moderate but not household-name recognition. The theory is straightforward forward confusion: are EMBER & ASH's prospective buyers misled into thinking the senior brand makes or backs it?
Phase 0. Counsel and an independent expert agree the issue truly turns on perception, the budget supports roughly 200 completes per cell, and the expert believes a fair design can find real confusion if it exists. They retain her before the complaint locks in a theory.
Phase 1 (universe). Forward confusion sets the universe to prospective purchasers of EMBER & ASH (the junior) hot sauce: adults who have bought or intend to buy craft hot sauce in the next three months in the relevant price band, captured by audited screeners.
Phase 2 (format). Because the senior mark is only moderately strong and shoppers genuinely meet both bottles in the same cooler, the expert chooses Squirt, and salts the two contested bottles among four other craft sauces, photographed as they appear on the shelf, to avoid a suggestive two-way matchup.
Phase 3 (control). The test cell shows the real EMBER & ASH bottle in the array. The control cell is identical except that EMBER & ASH is replaced with a coined, plainly unrelated name (say "CINDERWICK") in the same style and price. The team commits in advance to reporting net confusion.
Phase 4 (questions and administration). Open-ended source and affiliation questions, a real "don't know" option, neutral probes, rotated order, full-dress stimuli, double-blind online panel.
Phase 5 (sampling). About 200 completes per cell, 95% confidence, margin of error reported, panelists validated, the funnel fully documented.
Phase 6 (coding and report). Blind coders work a written protocol, inter-coder reliability is measured, every verbatim is preserved, and the report explains each design choice with its contemporaneous rationale.
The result. Test cell: 28% attribute EMBER & ASH to, or as approved by, the EMBERHOUSE company. Control cell: 7%. Net confusion: 21%. Above the soft 15% line, generated by a design that answers every cross-examination question in advance, and offered as one strong factor within the full Polaroid analysis rather than as the whole case.
The afterlife. Because the survey was built to survive, the Rule 702 motion fails, the verbatims withstand audit, and the cross-examination becomes a tour of choices the expert can defend point by point. That is what "a survey that survives" means in practice: not a louder number, but a number nobody can take away from you.
When to bring in counsel and an expert
Survey design sits at the seam of three disciplines, trademark law, survey methodology, and statistics, and the costly mistakes usually happen where the seams meet: a lawyer who picks a format for litigation advantage without regard to fit, or a methodologist who designs an elegant instrument around the wrong legal theory. The fix is to pair an experienced, independent survey expert with trademark litigation counsel early, before the theory hardens and before anyone promises the client a number.
Rightsy's virtual trademark attorneys can help you make the threshold call (is a survey worth it here, and what is it likely to show?), align the universe to the right theory of confusion, and connect the survey strategy to the rest of the dispute, from clearance and TTAB proceedings through trial and the remedies that follow a win. Grounding the work in Rightsy's trademark and logo search, brand-watch, and assignment records keeps the design tethered to the real marketplace, which is where confusion is, in the end, the only place it legally counts.
Primary authority
- Reliability and admissibility: Fed. R. Evid. 702 (as amended Dec. 1, 2023), 703, 803(3), and 403; Daubert v. Merrell Dow Pharms., Inc., 509 U.S. 579 (1993); Gen. Elec. Co. v. Joiner, 522 U.S. 136 (1997); Kumho Tire Co. v. Carmichael, 526 U.S. 137 (1999).
- Expert disclosure: Fed. R. Civ. P. 26(a)(2)(B); Fed. R. Civ. P. 37(c)(1).
- Confusion-survey formats: Union Carbide Corp. v. Ever-Ready, Inc., 531 F.2d 366 (7th Cir. 1976) (Eveready); SquirtCo v. Seven-Up Co., 628 F.2d 1086 (8th Cir. 1980) (Squirt).
- Genericness formats: E.I. DuPont de Nemours & Co. v. Yoshida Int'l, Inc., 393 F. Supp. 502 (E.D.N.Y. 1975) (Teflon); King-Seeley Thermos Co. v. Aladdin Indus., 321 F.2d 577 (2d Cir. 1963) (Thermos); USPTO v. Booking.com B.V., 591 U.S. 549 (2020) (consumer-perception survey evidence).
- Universe and methodology: Bristol-Myers Squibb Co. v. McNeil-P.P.C., Inc., 973 F.2d 1033 (2d Cir. 1992); Simon Prop. Grp., L.P. v. mySimon, Inc., 104 F. Supp. 2d 1033 (S.D. Ind. 2000); THOIP v. Walt Disney Co., 690 F. Supp. 2d 218 (S.D.N.Y. 2010).
- Doctrinal anchors: Lanham Act §§ 2(f), 14(3), 43(a), 43(c), 15 U.S.C. §§ 1052(f), 1064(3), 1125(a), 1125(c); Wal-Mart Stores, Inc. v. Samara Bros., Inc., 529 U.S. 205 (2000); Big O Tires, Inc. v. Goodyear Tire & Rubber Co., 561 F.2d 1365 (10th Cir. 1977) (reverse confusion); Jack Daniel's Props., Inc. v. VIP Prods. LLC, 599 U.S. 140 (2023).
- Reference works: Federal Judicial Center, Reference Manual on Scientific Evidence (Shari Seidman Diamond, Reference Guide on Survey Research); McCarthy on Trademarks and Unfair Competition §§ 32:158 et seq.; Restatement (Third) of Unfair Competition §§ 20–23; International Trademark Association (INTA) survey guidelines.
Survey design and admissibility are intensely fact-specific. The benchmarks, formats, and thresholds above are starting points, not rules; consult a qualified survey expert and trademark counsel about any particular matter.
Related Resources
- Building a Bulletproof Consumer Survey in Trademark Cases
- Keeping the Survey Out: Daubert Challenges to Trademark Survey Experts
- Running the Likelihood-of-Confusion Analysis: A Factor-by-Factor Checklist
- Likelihood of Confusion: A Brand Owner's Field Map
- The Polaroid Factors at Summary Judgment in the Second Circuit
- From Descriptive to Distinctive: How Marks Acquire Secondary Meaning
- The Abercrombie Spectrum: From Generic to Fanciful
- Trade Dress: Protecting Brand Identity Without Tripping Over Functionality
- Trademark Overview: Infringement, Dilution, and Related Rights
- The Trademark Enforcement Toolkit: From Watching to Verdict and Appeal
- The TTAB Practice Toolkit: Oppositions, Cancellations, and Appeals from Pleading to Decision
- Judge or Jury: Choosing Your Factfinder in Trademark Litigation
This checklist is general information, not legal advice, and does not create an attorney-client relationship. Consult qualified trademark litigation counsel and a qualified survey expert about any particular matter.