Beginner perfumer · 13
Everyone's nose is different — so does "smells good" mean anything? Blind test design, and when one person's evaluation can be trusted
· 23 min read
Two people differ functionally at over 30% of their odorant receptor alleles. Yet across 235 people in 9 non-western cultures, molecular identity explained 41% of the variance in pleasantness rankings and culture only 6% — leaving the largest share to neither. How much consensus actually exists, how to run a blind test, and which questions one evaluator can and cannot answer.
The previous article ended on a line: you are not composing for an average nose.
Which raises the obvious question. If everyone smells something different, is there any consensus about what smells good? When somebody says they like your work, is that information or politeness? And can an evaluation you ran alone at your desk be trusted at all?
All three have answers, and more precise ones than "it's subjective".
How different is the hardware?
Start with the size of the difference.
In 2014 Mainland and colleagues published work in Nature Neuroscience using a heterologous assay to measure directly whether genetic polymorphisms in odorant receptors change receptor function.
Humans have approximately 400 intact odorant receptors, but each individual has a unique set of genetic variations that lead to variation in olfactory perception. […] We identified agonists for 18 odorant receptors and found that 63% of the odorant receptors we examined had polymorphisms that altered in vitro function. On average, two individuals differ functionally at over 30% of their odorant receptor alleles.
Over thirty percent. "Some people are a bit more sensitive" describes a difference of degree; this is a different kind of thing. You and the person beside you have different parts fitted in a third of your sensors.
The OR5AN1 musk study from the previous article is one concrete instance of this.
Yet pleasantness really is shared
Given hardware that different, chaotic preference would be the natural prediction. Then someone measured it properly.
In 2022 Arshamian and colleagues published a study in Current Biology that was genuinely hard to run. They asked 235 individuals from 9 diverse non-western cultures (including hunter-gatherer and horticulturalist communities) to rank 10 monomolecular odorants by hedonic value.
The result:
We observed substantial global consistency, with molecular identity explaining 41% of the variance in individual pleasantness rankings, while culture explained only 6%. These rankings were predicted by the physicochemical properties of out-of-sample molecules and out-of-sample pleasantness ratings given by a separate group of industrialized western urbanites, indicating human olfactory perception is strongly constrained by universal principles.
Two figures are worth sitting with.
Molecule: 41%. "Smells good" has a footing in the molecule itself. Certain molecules are ranked highly everywhere, across every culture in the sample.
Culture: only 6%, far lower than most people expect. The folk belief that different cultures like completely different smells turns out to be largely wrong.
Putting the two together is where your answer is
Now do some arithmetic. Molecule 41%, culture 6%, together 47%.
So what claims the remaining 53%? Something beyond both the molecule and the culture.
This step is our arithmetic from those two figures, not a number the paper reports directly. The residual also contains measurement error and anything else the model did not capture, so treat it as an upper bound on individual difference rather than an estimate of it.
Even so, the direction is enough to work from: in a set of pleasantness ratings, the single largest source of variance is neither the odour nor the rater's culture. It is the rater.
Two conclusions follow, and they point in opposite directions.
-
Consensus exists. The molecule explains 41%, so you cannot wave away every criticism as taste. If eight people out of ten point at the same spot, something at that spot is probably wrong.
-
One person's opinion is a sample of size one, drawn from the largest variance component. Your own liking it, or one friend liking it, carries very little information.
The article on record keeping argued that "smells nice" is not data because a label can invert the judgement. This is the second reason, and it is quantified.
The materials most likely to cause disagreement
Which materials are most likely to make two people reach opposite verdicts? The descriptors give us a proxy.
We split descriptors into two classes: attractive (floral, sweet, fruity, fresh, rose, honey, vanilla, creamy, powdery…) and harsh (fecal, animal, sulfurous, sweaty, rancid, burnt, phenolic, medicinal, musty, metallic, pungent…), then found materials carrying both.
The answer: 22.3% (352 of 1,579).
That figure was corrected twice, and both corrections are traps this series has hit before.
First: the raw figure was 18.5%, but materials carrying both classes have a median of 7 descriptors against 4 for the rest — the more thoroughly something is described, the more likely it touches both lists. So we restricted the sample to materials with exactly 5 to 7 descriptors, giving a denominator of 1,579.
Second: after that control the
sulfurousfamily came out at 47.1%, which looked striking. Butsulfurousis itself on our harsh list, and materials in that family obviously carry the word — inflation by definition, the same problem as the descriptor-overlap calculation in the olfactory fatigue article. Excluding the primary family term itself drops sulfurous to 25.0% and gives the 22.3% above.
By family:
| Odour family | Share carrying both |
|---|---|
| Spicy | 31.0% |
| Green | 28.1% |
| Sulfurous | 25.0% |
| Herbal | 22.1% |
| Fruity | 21.3% |
| Woody | 20.2% |
| Floral | 14.3% |
| Balsamic | 12.5% |
| Citrus | 6.6% |
Citrus is two-sided only 6.6% of the time; spicy is five times as often.
The practical reading is direct: citrus is the highest-consensus family, spicy and green the most divisive. If your formula rests on spicy or green materials, expect opinion to split. A split of that kind says little about whether the formula failed; it says you need more people testing it. Sniffing it yourself a few more times adds nothing.
An inference that did not hold. We expected two-sided materials to carry lower recommended evaluation dilutions, since concentration decides which side dominates. Both groups have a median recommendation of 10% (two-sided n=290, rest n=467) — no difference. The hypothesis is not supported.
How to design a blind test
Here is a version you can actually run. No laboratory required, just discipline.
-
Cover the labels. Herz and von Clef, cited in the record-keeping article, showed that the same odour under a different label gets the opposite hedonic rating. Cover them even when testing yourself. Number samples randomly rather than A/B/C; the ordering itself hints at old version and new version. I learned this the annoying way, back when I habitually put my favourite last in the row. Friends would ask whether the final one was the new version before they had smelled anything.
-
Recruit five to eight people. This is the most important rule here, and the 53% above is the reason. Forget statistical significance; what you need to see is whether disagreement exists. Four out of five people pointing at the same problem is a different class of information from one person saying they don't like it.
-
Randomise order, and interleave families. The olfactory fatigue article made the point that what wears you out is a run of similar samples in a row, not the total count. Give every person a different order so order effects cannot pile up on one sample.
-
Ask for descriptors from a fixed list, and hold off on liking. That is exactly what the Dravnieks study cited in the record-keeping article did: a fixed list of 146 descriptors, about 150 subjects, with duplicate profiles correlating highly. Your scale is much smaller; the method copies directly.
-
For preference, ask one question: "would you want to smell it again?" It works far better than a 1–10 score. Scores get contaminated by expectation and politeness; "again" is a behavioural inclination and rather more honest.
-
Record the disagreements; that is where the information lives. When two people rate the same sample oppositely, check the descriptors first. If the material is spicy or green, you may be looking straight at that 31%.
When can one person's evaluation be trusted?
Probably the most useful section here. The answer depends on which question you are asking.
One person can reliably determine:
| Question | Why it holds |
|---|---|
| What is this? (descriptors) | The molecule explains 41%, and a fixed vocabulary is repeatable |
| Has it changed? (new batch vs old) | Same person, side by side — you are your own control (batch article) |
| How long does it hold? (on a strip, not a wrist) | An objective time point, provided a strip sidesteps your own adaptation |
| Where is the hole in the formula? (a break in the curve) | A structural question, not a hedonic one |
One person cannot determine:
| Question | Why not |
|---|---|
| Does it smell good? | The largest source of variance is you |
| Will other people like it? | Same, and a third of your receptors differ from theirs |
| Absolute intensity | Your threshold is not their threshold (the OR5AN1 study) |
| Is this musk prominent enough? | Musks are among the most common categories for specific anosmia |
In one line: one person can measure what something is; one person cannot measure whether it is good.
The evaluations you run at home still earn their keep. They handle identification and structure, which is the bulk of formulation work. But when the question becomes "is this ready to leave the house", you need other noses, and more than one.
Consensus exists, but it lives at the level of populations. The molecule accounts for 41%, and that is the part you can design. The largest remaining share sits inside each person who smells the thing, and that part you can only sample. Keeping those two apart is how you know when to change the formula and when to go find more people.
References
J. D. Mainland et al., The missense of smell: functional variability in the human odorant receptor repertoire, Nature Neuroscience, 17(1), 114–120 (2014). PMID 24316890
A. Arshamian et al., The perception of odor pleasantness is shared across cultures, Current Biology, 32(9), 2061–2066.e3 (2022). PMID 35381183
Next up is cost: how to work out the raw material cost of a formula, why the most expensive material is usually not the problem, and what the real threshold is for buying in small quantities.