Beginner perfumer · 8
What an evaluation note should contain — and why "smells nice" is not data
· 20 min read
The same odour, given a different label, gets the opposite hedonic rating — so "smells nice" records you on that day, not the material. Two studies plus a descriptor analysis of 5,824 materials: what makes a description stable, how many words it takes to identify a material, and what a note looks like that still reads six months later.
Six months later you open your evaluation notebook and find this line:
12 Mar — patchouli — quite nice, a bit sweet, I like it
That line is useless. What dilution? How far into the drydown? Compared against what? Is this morning's delivery even the same material? It answers none of it. All it actually preserves is your mood on the twelfth of March.
Writing more would not have saved it. The thing you wrote down was never stable in the first place.
Why "smells nice" is not data
In 2001 Herz and von Clef ran a very simple experiment in Perception. They chose five odours whose source is not strongly fixed and which can carry different hedonic readings: violet leaf, patchouli, pine oil, menthol, and a 1:1 mixture of isovaleric and butyric acids.
Subjects smelled each odour at two sessions a week apart. The same odour was given a different verbal label each time — one positive, one negative — and subjects rated it on several hedonic scales.
The result: the label significantly influenced perception of the odour. The authors propose that the cases where a label inverted perception are the first empirical demonstrations of olfactory illusions.
Sit with that for a second. The stimulus did not change. Same bottle, same concentration, same person. The only thing that moved was the words next to it, and the hedonic rating flipped.
So when you write "smells nice", you are not recording a property of the material. You are recording your reading of it that day. Six months on, your reading has moved, and the line is dead.
There is a practical rule hiding in this: cover the label when you evaluate. When you know what you are smelling, part of what you smell is what you expect it to be.
What kind of description is stable
So what does hold still?
Dravnieks answered that in Science in 1982. Ten compounds, roughly 150 subjects, each describing them from a fixed list of 146 descriptors. Duplicate profiles correlated highly (P < .001), consistently higher than profiles of different odours, and they agreed with profiles obtained previously.
The conclusion: profiles built from the combined responses of many subjects are stable constructs.
Stability comes from two things, a fixed vocabulary and repetition. You do not have 150 subjects. Even so, you can reproduce about one and a half of the two:
- The fixed vocabulary: fully reproducible. If you take one thing from this article, take this.
- Repetition: you will never muster 150 people, but you can smell the same material on different days and watch whether your own description drifts. That part costs nothing.
Free prose cannot do this. Is "a bit sweet, quite nice" the same perception as "warm, rounded" written three months later? There is no way to answer, because the two notes use different coordinate systems. A fixed vocabulary exists to nail the coordinate system down.
Descriptors in our data
We took the 5,824 materials that have both odour descriptions and supply information, and broke the descriptions into individual terms.
First, a figure we went and checked on ourselves. Across those 5,824 odour fields, how often do
good,nice,pleasant,unpleasant,beautiful,deliciousand similar evaluative words appear? Zero times. Not once.To be sure this was not an artifact of how we split the text into terms, we searched the raw unprocessed strings again. Still zero.
The trade's descriptive vocabulary has no field for "smells nice". Which makes sense: a professional vocabulary records properties of the odour and leaves the evaluator's reactions to the evaluator.
Those 5,824 materials use 557 distinct descriptors between them. Only 41 (7%) appear exactly once, and the top 50 terms account for 62% of all uses. That is a convergent, effectively controlled vocabulary, the same thing Dravnieks built by hand.
How many words does it take to identify a material?
Suppose you write down one descriptor. How far does it narrow 5,824 materials?
We took the 4,438 materials carrying at least three descriptors and queried using each material's own most common terms, i.e. the least favourable case:
| Descriptors written down | Materials still matching (median) |
|---|---|
| 1 | 1,333 |
| 2 | 223 |
| 3 | 23 |
| 4 | 4 |
One word filters almost nothing: green covers 23.1% of the database, sweet 22.9%, fruity 22.6%. But three words cut 5,824 down to 23, and four words to 4.
That is why a note needs three to five terms rather than one adjective. Diligence has nothing to do with it; the information only really starts arriving at the third word.
But descriptors alone are not enough
One finding cut the other way, and it is worth showing how we checked it.
Our first pass said that 26.5% of materials share their complete descriptor set with another material, which looked striking. Before using it we looked closer: the largest such group was 27 materials sharing one identical set, and that set was a single word, floral.
So the 26.5% could easily be an artifact of thin description rather than genuine indistinguishability. Broken down by how much was written:
| At least this many descriptors | Share their full set with another |
|---|---|
| 1 or more (n=5,824) | 26.5% |
| 2 or more (n=5,201) | 19.0% |
| 3 or more (n=4,438) | 13.4% |
| 4 or more (n=3,685) | 11.4% |
| 5 or more (n=2,904) | 11.1% |
Part of it was indeed an artifact; 26.5% falls to 11%. But it plateaus after the fourth descriptor (11.4% → 11.1%). That residual 11% comes from the ceiling of the descriptor tool itself, and no amount of extra writing removes it.
And within those identical-descriptor groups, only 29% can be separated by strength or substantivity. Since only about a third of materials carry strength and substantivity at all, 29% is a floor rather than a measurement.
The practical upshot: recording what it smells like is only half the job. You also have to record what you did.
A note that still reads six months later
Put the three findings together (labels distort perception, only a fixed vocabulary is stable, descriptors need conditions to be complete) and a useful record needs at least these fields:
What you did:
- Date
- Material, supplier and batch number (the same name from a different batch is a different thing; this is the field people skip most often and want most badly afterwards)
- Concentration and solvent (e.g. 10% in DPG)
- Carrier and amount (touch the strip, do not soak it)
What you smelled:
- Time points: 0 / 15 min / 1 h / 4 h / 24 h
- At each point, three to five descriptors, chosen only from your own fixed list
- Intensity 1–5
The conditions at the time (Herz showed context changes perception):
- What you smelled immediately before
- Which number this was in the session
- Whether the label was covered
And one line on use:
- ✗ "Lovely, really like it"
- ○ "Clean woody base, could hold up the tail of a citrus"
That last field is how "smells nice" becomes data: instead of writing whether you liked it, write what it could do. Liking changes over six months. Usefulness does not.
On building your own fixed list
Thirty to fifty terms is enough to start, covering the families you actually work in. Two rules:
- Only use terms on the list. When nothing on it fits, that failure is itself information. Record it and resist coining a word on the spot.
- Add new terms deliberately, and go back afterwards. Every addition changes the coordinate system, and if you leave the older notes unaligned, old and new stop being comparable.
This sounds rigid. It is exactly what Dravnieks' 146 descriptors were doing. Free prose is pleasant to write, and that same looseness is what keeps it from being comparable.
Three common mistakes
-
Only writing notes when something is remarkable. You end up with a notebook of outliers. My own early notebook reads that way, all raptures and disasters, and the merely-fine materials that would have made a baseline are simply missing.
-
Writing too much. A hundred-word paragraph is as unusable in six months as "quite nice", because you cannot tell which words were load-bearing. Three to five fixed terms plus a use-line wins.
-
Recording only the first impression. The previous article on olfactory fatigue made this point and it bears repeating: a note taken only at minute zero cannot answer "does it hold". Keep at least a one-hour and a twenty-four-hour point; that is what decides where a material sits in a formula.
A record exists so that you-in-six-months can be compared with you-today. That takes a coordinate system both of you recognise, which is why the fixed list exists, and why "smells nice" will never be on it.
References
A. Dravnieks, Odor quality: semantically generated multidimensional profiles are stable, Science, 218(4574), 799–801 (1982). PMID 7134974
R. S. Herz & J. von Clef, The influence of verbal labeling on the perception of odors: evidence for olfactory illusions?, Perception, 30(3), 381–391 (2001). PMID 11374206
Next up: batches. Why the same natural material from the same supplier differs every time, what the natural/synthetic distinction actually amounts to, and what that means for the record you have just started keeping.