What Is Sensory Difference Testing and How Do ISO 4120, 10399, and 8587 Work?

Sensory difference testing answers a simple but critical question in quality control: can panelists actually tell two products apart? Whether you are reformulating a food, switching a supplier, or verifying that a batch matches a standard, these standardized tests give you a statistically sound answer—not just an opinion.

What It Is

Sensory difference testing is a set of controlled procedures where trained or selected panelists evaluate products to determine whether a perceptible difference exists. The methods are governed by three key ISO standards:

  • ISO 4120 – Triangle test: Panelists receive three samples, two of which are identical, and must identify the odd one out.
  • ISO 10399 – Duo-trio test: Panelists receive a reference sample first, then two coded samples, and must pick which one matches the reference.
  • ISO 8587 – Ranking test: Panelists receive multiple samples and rank them by the intensity of a specified attribute (e.g., sweetness, bitterness, off-flavor).


These tests are used in product development, quality assurance, and shelf-life studies. They rely on binomial statistics to decide whether the number of correct answers exceeds what chance alone would predict.

How It Works: Steps and Statistics

### Triangle Test (ISO 4120)
  1. Prepare three samples per set: two from Product A, one from Product B (or vice versa). All possible serving orders are balanced across panelists.
  2. Present samples simultaneously, coded with random three-digit numbers.
  3. Ask each panelist: "Which sample is the odd one out" A forced-choice answer is required.
  4. Count the number of correct responses.


Statistical basis: Under the null hypothesis of no difference, the probability of a correct guess by chance is 1/3. The test uses the binomial distribution. For a given significance level (commonly α = 0.05) and panel size n, the minimum number of correct answers required is found from binomial tables or calculated as:

  • Critical value = smallest integer c such that P(X ≥ c) ≤ α, where X ~ Binomial(n, 1/3).


### Duo-Trio Test (ISO 10399)
  1. Present a labeled reference (R) first.
  2. Then present two coded samples: one is R, the other is the test product.
  3. Ask: "Which sample matches the reference" Again, forced choice.
  4. Count correct responses.


Statistical basis: The chance probability is 1/2, so the binomial model uses p = 0.5. The critical number of correct answers is determined similarly.

### Ranking Test (ISO 8587)
  1. Present three or more coded samples simultaneously.
  2. Ask panelists to rank them by the intensity of a defined attribute (e.g., from least to most sweet).
  3. Analyze ranks using Friedman's test or Page's test for ordered alternatives, depending on whether the expected order is known.


Statistical basis: Friedman's test compares rank sums across samples against a chi-square distribution. If a specific order is hypothesized, Page's test (an extension of Friedman's) provides greater power.

A Worked Illustrative Example

Example data (illustrative only) — Triangle test with 30 panelists:

  • You want to know if a new sweetener is perceptibly different from sugar in a beverage.
  • You run a triangle test per ISO 4120 with 30 panelists.
  • Results: 15 panelists correctly identified the odd sample.


Analysis: With n = 30 and chance probability p = 1/3, the expected number of correct answers by chance is 10. Using the binomial distribution at α = 0.05, the critical value for 30 panelists is 14 (i.e., at least 14 correct answers are needed to conclude a significant difference). Since 15 ≥ 14, you reject the null hypothesis and conclude that a perceptible difference exists.

If only 12 panelists had been correct, you would fail to reject the null hypothesis—the result is consistent with chance, and no significant difference is demonstrated.

Common Pitfalls

  • Insufficient panel size. Small panels rarely yield enough statistical power. Always calculate the required n before testing.
  • Inadequate sample balancing. In triangle tests, the position of the odd sample must be randomized equally across all six possible orders.
  • Ignoring forced-choice rules. Allowing "no difference" answers weakens the binomial model and biases results.
  • Using the wrong chance probability. Confusing 1/3 (triangle) with 1/2 (duo-trio) will produce incorrect critical values.
  • Over-testing with ranking. Ranking is less sensitive for detecting small differences; use triangle or duo-trio when the difference is subtle.


Closing

Sensory difference testing is a cornerstone of objective quality verification. By following ISO 4120, 10399, or 8587, you replace guesswork with a defensible, statistically valid decision. To simplify your calculations, use the free sensory difference testing tool at https://www.6sq.com/tools/sensory/—it applies the correct binomial and rank-based statistics for you.
Invited:

0 replies, guests cannot view replies. For more features, please log in or register