Designing a Trademark Survey That Survives: A Methodology Checklist

By ·

A trademark survey is the rare witness that can speak for thousands of consumers at once, and the rare witness whose entire testimony can be excluded before trial because of a choice made months earlier. This methodology checklist walks counsel and survey experts through building a confusion survey that survives cross-examination and a Rule 702 or Daubert challenge: gatekeeping the decision to survey at all, defining the load-bearing universe for the actual theory of confusion (forward, reverse, sponsorship, or dilution), selecting the Eveready or Squirt format, engineering a control cell so you report net rather than raw confusion, drafting neutral stimuli with marketplace realism, administering double-blind, sizing the sample, and coding verbatims objectively. It maps the sister formats for secondary meaning, fame, and genericness (Teflon and Thermos), and shows how each design choice doubles as a future cross-examination question. WHY notes, trap warnings, and worked examples accompany every phase, and a closing interpreter's guide explains what a net number actually means among the Polaroid and Sleekcraft factors. Survey design is fact-specific; engage a qualified survey expert and trademark counsel, such as Rightsy's virtual trademark attorneys, before you build.

Intellectual Property → Trademark Litigation | Published 28 June 2026 | rightsy.io

A consumer survey is the strangest witness in a trademark trial. Every other witness speaks for one person: what she saw on the shelf, what he meant to buy, what they remember from the ad. A survey claims to speak for thousands of people at once, compresses their collective state of mind into a single percentage, and delivers it with the quiet authority of arithmetic. Done right, it is the closest thing trademark law has to a window into the consumer's head, the one piece of evidence that can answer the decisive question in an infringement case directly rather than by inference: are people actually confused?

But the survey is also the only witness whose entire testimony can be thrown out before the jury hears a syllable, and thrown out not because the witness lied, but because of a choice somebody made in a conference room six months earlier. A leading question drafted on a Tuesday, a universe defined too broadly in a kickoff call, a missing control cell nobody budgeted for: any one of them can convert a six-figure scientific instrument into an exhibit the court strikes on a paper motion.

That paradox is what this checklist is built around. The phrase in the title is deliberate. We are not designing a survey that merely exists, or one that produces a headline number a press release can use. We are designing a survey that survives: survives the motion to exclude under Federal Rule of Evidence 702 and Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993); survives a cross-examiner who has read every word of the questionnaire backward; and survives a judge who has seen a hundred surveys and trusts none of them on faith.

Here is the organizing secret of survey design, and the reason the phases below run in the order they do: the sequence in which you make design decisions is the sequence in which the other side will attack them. A cross-examiner does not start with your coding protocol. She starts with whom you interviewed, then why you showed them what you showed them, then what you compared it against, then how you asked. Every checkbox in this document is therefore also a future deposition question. Tick it now, in the calm of the design phase, or answer it later under oath, in the worst possible light.

Two framing reminders before we build. First, a survey is one factor, not the whole case. Confusion is proven through the full multifactor inquiry, the Polaroid set in the Second Circuit and the Sleekcraft set in the Ninth, and a strong survey is powerful evidence of the "actual confusion" factor and a useful proxy for the "strength of mark" and "consumer sophistication" factors, but it is never the entire ballgame. For the larger picture, see Running the Likelihood-of-Confusion Analysis and Likelihood of Confusion: A Brand Owner's Field Map. Second, a survey is never mandatory. Plenty of cases are won with anecdotal actual confusion, the parties' own documents, and the other factors. The decision to commission one is itself a strategic choice, and the first phase of this checklist is about making it honestly.

A practical note on groundwork. Before you can define a universe or build a realistic stimulus, you have to know what the contested marks actually look like at the point of sale, who is really selling under them, and who owns the rights on each side. A clearance-grade trademark and logo search on Rightsy, paired with a chain-of-title check in Rightsy's assignment records, grounds the whole exercise in marketplace reality rather than in the lawyers' assumptions about it. Surveys fail far more often from a bad picture of the market than from bad statistics.


Phase 0 — Decide whether to commission a survey at all

This is the phase everyone wants to skip. Skip it at your peril.

WHY. Surveys are persuasive precisely because they are expensive and rigorous; that same rigor is what makes a half-built one dangerous. Courts and juries give weight to surveys because they look scientific, which means a sloppy survey borrows credibility it has not earned, and the borrowed credibility cuts against you the instant the cross-examiner exposes the flaw.

TRAP — the survey that finds nothing. A mediocre survey is frequently worse than no survey. Run a weak design, get a low confusion number, and you have manufactured your opponent's best exhibit at your own expense. The defense will wave your own 4% net confusion figure at the jury for the rest of the trial. If a credible expert tells you, in good conscience, that she cannot build a survey likely to find meaningful confusion on these facts, that is not a failure. That is data about your case. Listen to it.

TRAP — the late expert. Retaining the expert after you have filed a complaint that pleads a specific theory of confusion, or after you have taken positions in discovery that define the market a certain way, can box the survey into a universe that no longer fits. The expert should help shape the theory, not inherit a frozen one.

FIELD NOTE. Whether a survey is worth it often depends on the forum. A bench trial before a trademark-fluent judge, a jury trial, a preliminary-injunction sprint, and a Trademark Trial and Appeal Board proceeding all reward different things. Surveys are admissible and used at the TTAB but are comparatively rare there given cost and the Board's expertise; if your fight is an opposition or cancellation, weigh that before spending. See The TTAB Practice Toolkit and, on the forum's effect on every evidentiary choice, Judge or Jury: Choosing Your Factfinder.


Phase 1 — Define the universe (the load-bearing wall)

If there is one decision that holds up the entire structure, it is this one. Define the wrong population to interview and nothing built on top of it can be salvaged. You did not measure confusion badly; you measured the wrong people's confusion, and no clever cross-examination repair, no eloquent closing argument, no re-weighting can fix a foundation poured in the wrong place.

WHY. The universe is the population your sample is supposed to represent. Get it right and a clean sample of a few hundred people can speak for millions. Get it wrong and the most elegant questionnaire in the world is describing strangers. Courts treat a misdefined universe as a foundational, often fatal, defect rather than a quibble that goes to weight; the canonical statement is Bristol-Myers Squibb Co. v. McNeil-P.P.C., Inc., 973 F.2d 1033 (2d Cir. 1992), where the survey of the wrong population could not support the inferences drawn from it.

TRAP — over-inclusive and under-inclusive universes. Two named errors, both lethal. An over-inclusive universe sweeps in people who are not real prospective purchasers (interviewing the general public about a $40,000 industrial 3D printer), which dilutes and distorts the result. An under-inclusive universe excludes relevant buyers, so the result cannot be generalized to the market that matters. Either way, a sophisticated opponent will probe your screeners line by line in deposition, because loose or tight screeners are the back door through which a universe problem walks right back in after you thought you had closed the front.

TRAP — letting the format pick the universe. The format you choose (next phase) does not change the universe. A Squirt survey run on the wrong population is exactly as broken as an Eveready survey run on the wrong population. Universe first, always; format second.

Worked example — forward confusion (FATHOM). Suppose FATHOM is a well-known maker of premium dive watches, an arbitrary, commercially strong mark. A newcomer launches FATHOMLESS smartwatch bands and accessories. Under a forward-confusion theory, the universe is prospective purchasers of FATHOMLESS bands, not FATHOM watch buyers, because the alleged harm is that the newcomer's customers think the senior watchmaker stands behind the bands. Screeners should capture people who have bought or intend to buy smartwatch accessories in a realistic window and price range.

Worked example — reverse confusion (VAULT). Now suppose VAULT is a tiny regional scheduling app sold to a handful of local gyms. A national fitness conglomerate rolls out a VAULT workout platform with a Super Bowl ad and blanket sponsorships. The senior user's grievance is reverse confusion: its own prospective customers now assume little VAULT is an offshoot of, or rip-off of, the famous newcomer. The universe therefore flips to prospective purchasers of the senior (small) VAULT's services. Reverse-confusion theory traces to Big O Tires, Inc. v. Goodyear Tire & Rubber Co., 561 F.2d 1365 (10th Cir. 1977); the universe flip is the single most-missed move in this corner of survey design. If priority and ownership are contested, confirm who actually holds the senior rights with a chain-of-title check in Rightsy's assignment records before you build a universe around "the senior user."


Phase 2 — Choose the format that fits the marks

Two workhorse formats dominate confusion surveys, and choosing between them is choosing how closely your laboratory will mimic the way real consumers encounter the marks. The choice is not a matter of taste or habit. It is dictated by the strength of the senior mark and by whether shoppers genuinely meet the two marks together in the wild.

WHY — Eveready. The Eveready format, from Union Carbide Corp. v. Ever-Ready, Inc., 531 F.2d 366 (7th Cir. 1976), is widely regarded as the gold standard because it mirrors how confusion actually happens for a strong brand: a shopper sees the junior product alone, and the senior mark is so embedded in memory that it surfaces unprompted. Because it relies on the respondent's own recall and shows nothing suggestive, it minimizes the risk that the survey created the very association it claims to measure.

TRAP — Eveready on an obscure mark. Eveready has a built-in floor: it only counts confusion among people who already carry the senior mark in their heads. Run it on a senior mark few consumers have ever heard of and it will under-count confusion for the wrong reason, because respondents cannot name a brand they do not know, even if the marks are genuinely confusable side by side. That is a mismatch between tool and fact, not proof of no confusion.

WHY — Squirt. The Squirt format, from SquirtCo v. Seven-Up Co., 628 F.2d 1086 (8th Cir. 1980), shows both marks and is appropriate where consumers really do see them in proximity, on the same shelf, in the same search results, in the same category. It rescues cases where Eveready would unfairly under-count a weaker senior mark.

TRAP — Squirt's suggestion problem. Showing two marks side by side can manufacture a relationship no real shopper would ever perceive, simply because the human mind, handed two things at once, goes looking for how they connect. This is the demand effect at the level of the whole design. A bare two-way comparison is one of the most exploitable artifacts in trademark litigation, which is why the salted array exists and why the control cell (next phase) has to be especially muscular for Squirt. If a plainly famous senior mark's expert quietly chose Squirt over Eveready, the most common backstage reason is that Eveready did not show enough confusion. Expect the cross-examiner to ask, and expect the judge to wonder.

Worked example — Squirt done honestly (EMBERHOUSE vs. EMBER & ASH). Two regional hot-sauce makers, EMBERHOUSE (senior) and EMBER & ASH (junior), sit in the same refrigerated specialty-foods cooler in the same stores. Consumers really do encounter both at once, so Squirt fits the facts. The honest version shows both bottles inside an array of four or five other hot sauces of similar style, photographed as they appear on the shelf, and asks a neutral affiliation question. The dishonest version puts the two bottles alone on a white background and asks, in effect, "don't these go together?" Same format name, opposite credibility.


Phase 3 — Build the control and plan the net

Here is a number that means nothing by itself: "34% of respondents were confused." Here is a number that means something: "the test cell showed 34% confusion, the control cell showed 9%, for net confusion of 25%." The difference between those two sentences is the difference between a survey that survives and one that gets shredded, and it all lives in the control.

WHY. Every survey generates background noise: people guess, people yea-say to be agreeable, people carry pre-existing beliefs into the room, and the instrument itself nudges. The control measures that noise floor by running the same gauntlet without the infringing element. Subtract it and you isolate the confusion actually attributable to the defendant's mark, the way a clinical trial subtracts the placebo response to isolate the drug. The first question a sophisticated reader, or a sharp judge, asks about any confusion survey is not "what was the rate?" It is "what was the control, and what is the net?" If your answer is "there was no control," you have already lost the witness.

TRAP — the controlless Squirt. A Squirt survey with no control is close to uninterpretable, because the format's own suggestiveness pumps up the raw number and there is nothing to subtract it back down. Judges and good cross-examiners know this cold. The thorough opinion in Simon Property Group, L.P. v. mySimon, Inc., 104 F. Supp. 2d 1033 (S.D. Ind. 2000), is a master class in how a court walks through control design and net confusion, and it rewards a careful read for anyone building one of these.

TRAP — the strawman control. A control so obviously different from the test stimulus that it would never confuse anyone produces an artificially low control number and an artificially high net. The cross-examiner will argue you rigged the baseline. The control should be a fair, realistic non-infringing comparator, not a punching bag.

Worked example. Back to FATHOMLESS bands against famous FATHOM watches. The test cell shows the actual FATHOMLESS product. The control cell shows the same kind of product under a coined, plainly unrelated name (say, "OAKHELM") with the same look, price, and packaging style. If the test cell yields 31% "made or approved by the watch company" and the control yields 8% (the people who would attribute any upscale band to a watchmaker), your net is 23%, and you can defend every point of it.


Phase 4 — Draft the questions, build the stimuli, run it blind

This is the craftsmanship phase, where good intentions die in the wording. The governing principle is restraint: the questionnaire must extract the respondent's genuine, unguided belief without ever hinting at the answer you are hoping for.

WHY — neutral wording. A leading question generates confusion that exists nowhere but inside the instrument. Ask "Do you think the same company makes both of these?" and you have planted the idea of one company; some fraction of agreeable respondents will salute it regardless of the marks. The open-ended form lets confusion surface on its own or not at all, which is the only kind a court cares about.

WHY — the "don't know" valve. Without an easy escape hatch, respondents who genuinely have no idea will guess, and guesses contaminated by the test stimulus often land on the plaintiff. A real "don't know" option keeps non-opinions out of your numerator. (It also models honesty for the court: a survey unafraid of "don't know" looks like one built to find truth, not a number.)

WHY — marketplace realism. The legal question is whether confusion is likely in the marketplace, so the further your stimulus drifts from the real shopping experience, the less your result measures the thing in dispute. An artificial side-by-side that no shopper would ever see measures something other than real-world confusion. This has special bite in trade dress cases, where the look and feel is the mark, and stripping it down to a logo on a card guts the very thing you are testing. Surveys have been excluded for stimuli unmoored from how consumers actually meet the products; THOIP v. Walt Disney Co., 690 F. Supp. 2d 218 (S.D.N.Y. 2010), is a useful cautionary tale.

WHY — double-blind. Single-blind is not enough. An interviewer who knows the desired answer can signal it without meaning to, through tone, pacing, an encouraging "mm-hm" at the right moment. Double-blind closes that channel, and just as important, it lets you tell the court, truthfully, that no human in the chain could have steered the result. The Reference Manual on Scientific Evidence (Federal Judicial Center), and within it Shari Seidman Diamond's Reference Guide on Survey Research, is the standard articulation of these controls and the document your expert should be able to cite from memory.

TRAP — the tidy white background. The single most common realism failure is presenting word marks alone, decontextualized, because it is cheap and easy to program. It also strips away exactly the marketplace cues that drive or dispel confusion. If the defendant's house brand, dress, or disclaimer would be visible at the point of sale, hiding it in your stimulus is not neutrality; it is a thumb on the scale your opponent will expose.

TRAP — the leading "filter that isn't." Sometimes the leading happens up in the screeners, where a poorly worded qualifier primes the category or the brand before the test question is ever asked. Audit the whole instrument for priming, not just the headline question.


Phase 5 — Sample and quantify

Now the statistics. A beautiful questionnaire administered to too few of the wrong people, with no honest accounting of error, is still a broken survey. This phase is where the result earns the right to be called a measurement.

WHY. Sample size and margin of error are not decoration; they are the difference between a number and a guess. A net confusion figure of "20%" computed from 40 people per cell carries a margin of error so wide it could honestly be anything from trivial to overwhelming, and the cross-examiner will read your own confidence interval back to you. Reporting the error candidly is not a weakness to hide; it is the signature of a real measurement, and it inoculates you against the accusation of overstatement.

TRAP — the unjustified panel. Online panels are accepted, but "we used a panel" is not a methodology; how you sourced, screened, and validated respondents is. An undocumented panel of professional survey-takers who blow through the questionnaire for points is a gift to the defense. Build and keep the validation paper trail from the first respondent.

TRAP — the disappearing denominator. Be able to account for everyone: how many were invited, how many qualified, how many completed, how many were tossed and why. A survey that cannot reconstruct its own funnel invites the inference that inconvenient respondents quietly vanished.


Phase 6 — Code, document, and report

The last mile, and the place where mischief most easily hides, because coding open-ended answers is a human act of judgment performed on thousands of ambiguous sentences. "I think it's the watch people, but honestly I'm just guessing" is confused, uncertain, or both, depending entirely on who holds the pen and what rules they follow.

WHY. Blind, protocol-driven coding with a reliability statistic is how you convert subjective interpretation into a reproducible measurement. Without it, the coding step becomes an unfalsifiable judgment call, and an unfalsifiable judgment call is not science; it is advocacy in a lab coat. Producing the verbatims is the ultimate proof of good faith: it lets the court check your work. A report that resists producing its verbatims is signaling, loudly, that the verbatims would not survive the daylight.

TRAP — coding toward the conclusion. When coders know which cell they are scoring, ambiguous answers drift, ever so slightly, toward the hoped-for result. Blind coding removes both the temptation and the appearance of it. If your inter-coder reliability is poor, do not paper over it; fix the protocol and recode.

TRAP — the orphaned methodology. A survey whose design choices were never written down until the expert report was drafted looks reverse-engineered, because it often was. Memorialize the why contemporaneously, in working memos and the expert's notes, so the rationale predates the result.


Phase 7 — Build it to survive Rule 702 and Daubert from day one

Admissibility is not a hurdle you clear at the end; it is a design constraint you honor from the first decision. The discipline is simple to state and hard to live: make every choice as though the judge will read the questionnaire word for word, because in a serious case, she will.

WHY. The 2023 amendment to Rule 702 was a course correction. For years, courts waved methodological problems through as going to "weight, not admissibility," letting the jury sort it out. The amended rule makes explicit what was always implicit: the proponent carries the burden, by a preponderance, and the court must find the method was reliably applied, not merely that the expert is credentialed and the method exists in the abstract. The practical upshot for surveys is that fundamental defects increasingly draw exclusion rather than a footnote. A wrong universe, an absent control, leading questions, an unrealistic stimulus: these used to be cross-examination fodder; now they are exclusion arguments. Designing for survival is no longer optional polish. It is the price of admission.

WHY — hearsay is handled, if you do it right. A confusion survey is technically a mountain of out-of-court statements, yet surveys come in routinely. Courts treat properly conducted surveys as admissible either as non-hearsay evidence of the respondents' present states of mind under Federal Rule of Evidence 803(3), or as the reliable basis for expert opinion under Rule 703. The throughline is trustworthiness: the more rigorously you followed the methodology in this checklist, the more comfortably the hearsay objection dissolves. Sloppiness, by contrast, reopens the door.

TRAP — the Daubert mirror. Everything that makes your survey survive a challenge is also a blueprint for attacking the other side's. The defensive checklist and the offensive checklist are the same document read in two directions. For the attack version, see Keeping the Survey Out: Daubert Challenges to Trademark Survey Experts. For the affirmative build, our companion piece Building a Bulletproof Consumer Survey in Trademark Cases goes deeper on the construction side.


The cross-examiner's checklist (read your survey through enemy eyes)

Before you call a survey finished, run it through the exact sequence a hostile expert and a hostile lawyer will use. If you cannot give a confident, documented answer to each, you have found your homework.

Notice the order. It is Phases 1 through 7, run in sequence. The cross-examination is the design checklist. That is not a coincidence; it is the whole point.


The sister surveys: secondary meaning, fame, and genericness

Confusion is the headline use, but the same methodological discipline powers surveys aimed at other questions. The instrument is one discipline; only the target moves. Each of the following must still satisfy the universal commandments, right universe, neutral questions, blind administration, objective coding, before its specialized design matters at all.

Secondary meaning (acquired distinctiveness)

When a descriptive term or a product look can only be protected on a showing of acquired distinctiveness under Lanham Act § 2(f), 15 U.S.C. § 1052(f), a survey can test whether the relevant public has come to treat the term as identifying a single source. (For product design trade dress, secondary meaning is always required; Wal-Mart Stores, Inc. v. Samara Bros., Inc., 529 U.S. 205 (2000).)

Worked example (FRESHPRESS). A juice company claims FRESHPRESS for cold-pressed beverages, an obviously descriptive term, and must prove the market sees it as a brand rather than a description. A clean survey asks whether FRESHPRESS juice comes from one company or several, controls with a comparably descriptive non-claimed term, and reports the gap. For the doctrine this evidence serves, see From Descriptive to Distinctive: How Marks Acquire Secondary Meaning and The Abercrombie Spectrum.

Fame (for dilution)

Dilution under 15 U.S.C. § 1125(c) is reserved for marks that are famous, meaning "widely recognized by the general consuming public of the United States," a deliberately demanding standard the Trademark Dilution Revision Act set in 2006.

TRAP. Sampling category buyers for fame is a classic mismatch: you may prove a mark is well known to enthusiasts, which is precisely not the legal question. Niche fame is not fame for dilution. For where dilution sits among brand rights, see Trademark Overview: Infringement, Dilution, and Related Rights.

Genericness — the Teflon format

When the fight is whether a term has become the common name for the product itself (a generic), the primary-significance test asks what the term means to the relevant public. Congress codified that focus in the 1984 Trademark Clarification Act, directing courts to the primary significance of the mark to the relevant public and rejecting any "purchaser motivation" detour. The Teflon survey, named for E.I. DuPont de Nemours & Co. v. Yoshida International, Inc., 393 F. Supp. 502 (E.D.N.Y. 1975), is the preferred tool.

WHY courts prefer Teflon. The mini-tutorial plus forced classification produces cleaner, more interpretable data than open-ended approaches, which is why it has become the dominant genericness instrument. Its modern star turn came in USPTO v. Booking.com B.V., 591 U.S. 549 (2020), where survey evidence of consumer perception helped establish that even a "generic.com" term could function as a brand if the public so perceived it. The lesson cuts both ways: a well-built Teflon survey can save a mark from the genericide graveyard or push it in.

Worked example (GLIDEBOARD). A maker of self-balancing scooters wants to keep GLIDEBOARD from sliding into generic use. A Teflon survey tutors respondents, then asks them to sort a rotated list, KLEENEX-style known brands, plainly generic terms like "stapler," and the disputed GLIDEBOARD, into "brand" or "common name." The brand percentage, properly controlled, is the headline. On how a once-strong mark slides toward the public domain, see Use It or Lose It: How Trademarks Are Abandoned.

Genericness — the Thermos format

The Thermos format, from King-Seeley Thermos Co. v. Aladdin Industries, Inc., 321 F.2d 577 (2d Cir. 1963), comes at the same question from the other side: it asks what respondents would spontaneously call the product if they were shopping for one, capturing generic usage in the wild.

WHY. Thermos catches the consumer who reaches for the disputed word as the ordinary name of the thing, with no brand prompting at all. It is the linguistic snapshot of a mark in the act of going generic. Paired with Teflon's structured classification, it gives a court two independent angles on primary significance.


The interpreter's guide: what the number actually means

Suppose you did everything right and you have a clean net confusion figure. What does it prove? Less than the press release wants, and more than the defense will admit.

WHY the thresholds are soft. Those percentages are rules of thumb distilled from decades of cases (collected at length in McCarthy on Trademarks §§ 32:158 et seq.), not statutory lines. Courts have found likely confusion below 15% and rejected it above, depending on the marks, the goods, and the quality of the survey. A pristine 12% from a flawless design can outweigh a shaky 25% from a contested one. Methodology is destiny: a number is only worth the design that produced it. On how these factors play out at the dispositive-motion stage, see The Polaroid Factors at Summary Judgment in the Second Circuit. For where the survey sits in the broader campaign, from watching to verdict to appeal, see The Trademark Enforcement Toolkit.

FIELD NOTE — the modern survey moment. Surveys are not a relic. A confusion survey was squarely in the record in Jack Daniel's Properties, Inc. v. VIP Products LLC, 599 U.S. 140 (2023), the "Bad Spaniels" dog-toy dispute, and Booking.com turned in significant part on consumer-perception survey evidence. The instrument is alive, contested, and frequently decisive, which is exactly why building it to survive matters now more than ever.


Common fatal flaws, in one place

A field guide to the wreckage, in roughly the order an opponent will find it:


A modular toolkit: reusable components you can lift into any design

Think of these as the standard parts you bolt onto every survey, regardless of the legal question. Keep templates of each so you are not reinventing them under deadline.

Reusing vetted components does more than save time. It builds a consistent, defensible house style across your matters, so that when an opponent's expert tries to paint a particular choice as ad hoc or outcome-driven, you can show it is your standard practice, applied the same way every time. Consistency is credibility.


A short worked case study, end to end

To see the phases lock together, follow one hypothetical from intake to interpretation.

The dispute. EMBERHOUSE, a regional craft hot-sauce maker, sues EMBER & ASH, a newcomer whose bottles sit in the same specialty coolers. EMBERHOUSE has moderate but not household-name recognition. The theory is straightforward forward confusion: are EMBER & ASH's prospective buyers misled into thinking the senior brand makes or backs it?

Phase 0. Counsel and an independent expert agree the issue truly turns on perception, the budget supports roughly 200 completes per cell, and the expert believes a fair design can find real confusion if it exists. They retain her before the complaint locks in a theory.

Phase 1 (universe). Forward confusion sets the universe to prospective purchasers of EMBER & ASH (the junior) hot sauce: adults who have bought or intend to buy craft hot sauce in the next three months in the relevant price band, captured by audited screeners.

Phase 2 (format). Because the senior mark is only moderately strong and shoppers genuinely meet both bottles in the same cooler, the expert chooses Squirt, and salts the two contested bottles among four other craft sauces, photographed as they appear on the shelf, to avoid a suggestive two-way matchup.

Phase 3 (control). The test cell shows the real EMBER & ASH bottle in the array. The control cell is identical except that EMBER & ASH is replaced with a coined, plainly unrelated name (say "CINDERWICK") in the same style and price. The team commits in advance to reporting net confusion.

Phase 4 (questions and administration). Open-ended source and affiliation questions, a real "don't know" option, neutral probes, rotated order, full-dress stimuli, double-blind online panel.

Phase 5 (sampling). About 200 completes per cell, 95% confidence, margin of error reported, panelists validated, the funnel fully documented.

Phase 6 (coding and report). Blind coders work a written protocol, inter-coder reliability is measured, every verbatim is preserved, and the report explains each design choice with its contemporaneous rationale.

The result. Test cell: 28% attribute EMBER & ASH to, or as approved by, the EMBERHOUSE company. Control cell: 7%. Net confusion: 21%. Above the soft 15% line, generated by a design that answers every cross-examination question in advance, and offered as one strong factor within the full Polaroid analysis rather than as the whole case.

The afterlife. Because the survey was built to survive, the Rule 702 motion fails, the verbatims withstand audit, and the cross-examination becomes a tour of choices the expert can defend point by point. That is what "a survey that survives" means in practice: not a louder number, but a number nobody can take away from you.


When to bring in counsel and an expert

Survey design sits at the seam of three disciplines, trademark law, survey methodology, and statistics, and the costly mistakes usually happen where the seams meet: a lawyer who picks a format for litigation advantage without regard to fit, or a methodologist who designs an elegant instrument around the wrong legal theory. The fix is to pair an experienced, independent survey expert with trademark litigation counsel early, before the theory hardens and before anyone promises the client a number.

Rightsy's virtual trademark attorneys can help you make the threshold call (is a survey worth it here, and what is it likely to show?), align the universe to the right theory of confusion, and connect the survey strategy to the rest of the dispute, from clearance and TTAB proceedings through trial and the remedies that follow a win. Grounding the work in Rightsy's trademark and logo search, brand-watch, and assignment records keeps the design tethered to the real marketplace, which is where confusion is, in the end, the only place it legally counts.


Primary authority

Survey design and admissibility are intensely fact-specific. The benchmarks, formats, and thresholds above are starting points, not rules; consult a qualified survey expert and trademark counsel about any particular matter.


Related Resources

This checklist is general information, not legal advice, and does not create an attorney-client relationship. Consult qualified trademark litigation counsel and a qualified survey expert about any particular matter.

Read this article on Rightsy