Beginner perfumer · 34
Grading longevity from excellent to poor draws a line through a distribution that has no natural cut points
· 24 min read
A forum complaint blamed a fragrance database for rating longevity from excellent to poor instead of just showing the time with no value judgement. The data can be looked at. This site's database holds 2,027 substantivity records with only 210 distinct values, and twelve of those values account for 55.4% of all records. Worse, the values are round numbers — 4, 8, 12, 24, 400 — so a line at eight hours does not fall in a gap between groups. It cuts straight through the 92 records sitting on exactly 8.0. Survey methodology has long measured that labels change how people answer, though those studies measured respondents rather than readers.
One reply in the "what annoys you" thread drew 28 upvotes and named a target:
"That was Fragrantica's fault when they started rating longevity from excellent to poor instead of just showing the time with no value judgement. Annoying."
The agreement underneath:
"Exactly! Just that alone has influenced a whole new generation of people just discovering fragrances. They are incorrectly getting the impression that a perfume is not worth even smelling unless it's 'beast mode'."
That is a complaint about rating scale design, and it can be checked from two directions: what the data looks like, and whether labels change judgement.
The short version
- This site's database holds 2,027 substantivity records, with only 210 distinct values.
- Twelve values account for 55.4% of them. The biggest pile is 400 hours, at 334 records (16.5%).
- Those piles sit on round numbers: 4 hours (8.4%), 12 hours (4.9%), 8 hours (4.5%), 24 hours (4.1%).
- So any "under N hours is poor" line lands on a pile rather than in a gap. At eight hours, the line cuts straight through the 92 records sitting on exactly 8.0 — which side they fall on depends entirely on whether the rule says "≥8" or ">8".
- Overall: median 72 hours, interquartile range 12 to 296, and 20.1% at or under 8 hours.
- Survey methodology has measured label effects: verbal and numeric labels attached to response categories may substantially influence how respondents make their choices within a given response format (Blais and Grondin 2011). Those studies measured people filling in questionnaires, not people reading a database.
About 9 minutes.
What the distribution looks like
The data first. There are 2,027 materials with a substantivity record:
| Value | |
|---|---|
| Median | 72 hours |
| First quartile | 12 hours |
| Third quartile | 296 hours |
| Minimum | 1 hour |
| Maximum | 901 hours |
| Distinct values | 210 |
Two thousand records and 210 different numbers.
And twelve of those 210 take up more than half:
| Substantivity | Records | Share |
|---|---|---|
| 400 hours | 334 | 16.5% |
| 4 hours | 170 | 8.4% |
| 12 hours | 99 | 4.9% |
| 8 hours | 92 | 4.5% |
| 24 hours | 83 | 4.1% |
| 1 hour | 65 | 3.2% |
| 16 hours | 62 | 3.1% |
| 48 hours | 62 | 3.1% |
| 168 hours | 53 | 2.6% |
| 20 hours | 40 | 2.0% |
| 28 hours | 35 | 1.7% |
| 336 hours | 27 | 1.3% |
| Total | 1,122 | 55.4% |
This is not a smooth curve. It is a row of columns.
Why the columns' positions matter
Because they stand on round numbers.
4, 8, 12, 16, 20, 24, 28, 48, 168, 336, 400 — these are numbers a person chooses when recording, not numbers nature produced. A material lasting seven hours and forty minutes on a blotter gets written down as 8.
And whoever draws a label boundary chooses from the same set of round numbers.
"Under eight hours is poor" sounds like a neutral line. It runs straight through the 92 records sitting on 8.0.
Whether those 92 materials land in "poor" or out of it depends entirely on whether the rule reads "≤8" or "<8". Both are reasonable, and they produce different lists.
Put the line at 9 hours or 7 and the problem disappears — almost nothing is recorded there. The trouble is exactly that the convenient round boundaries and the data's piles are the same set of numbers.
The cumulative view: every line is steep
| Boundary | At or below | Share |
|---|---|---|
| ≤ 4 hours | 276 | 13.6% |
| ≤ 8 hours | 408 | 20.1% |
| ≤ 12 hours | 520 | 25.7% |
| ≤ 24 hours | 729 | 36.0% |
| ≤ 48 hours | 916 | 45.2% |
| ≤ 100 hours | 1,115 | 55.0% |
| ≤ 200 hours | 1,379 | 68.0% |
| ≤ 400 hours | 2,010 | 99.2% |
Moving the line from 8 hours to 12 pulls in another 112 materials, 5.6 percentage points. Four hours is a very small step in sensory terms.
The last row deserves particular attention. Between 200 and 400 hours the cumulative share jumps from 68.0% to 99.2% — because of that 400-hour column.
We've written about this: 400 hours is where many measurements were truncated rather than a measured value. 334 records sit there, and many of them mean "still detectable when we stopped looking".
So labels break at the top of the scale as well, in the opposite direction: a large group of "at least 400" materials all get the same grade, and how far apart they really are is unknown.
Do labels change judgement
Here I can offer only indirect evidence, and the limits come first.
Survey methodology has studied how labels affect answers. Blais and Grondin open their 2011 paper in the Journal of Applied Measurement with exactly this:
Survey questionnaires are among the most used data gathering techniques in the social sciences researchers' toolbox and many factors can influence respondents' answers on items and affect data validity. Among these factors, research has accumulated which demonstrates that verbal and numeric labels associated with item's response categories in such questionnaire may influence substantially the way in which respondents operate their choices within the proposed response format.
Their own work is narrower: using Andrich's rating scale model to show what influence the quantifier adverb "totally", used to label or emphasise extreme categories, could have on respondents' answers.
Krosnick's 1999 review of survey research in the Annual Review of Psychology lists the same thing among recent advances:
Surprising effects of the verbal labels put on rating scale points have been identified, suggesting optimal approaches to scale labeling.
Now the limitation, because it matters.
Those studies measured how people filling in a questionnaire pick a box, and the forum complaint is about what expectations readers form. Adjacent, and not the same.
I found no study measuring the second thing directly. The claim that a database's labels shaped a whole generation's expectations is something I can neither verify nor refute. It is plausible, and for now it is an observation rather than a finding.
What I can do is confirm that "labels affect how people use a scale" has been studied methodologically, and then do the data half properly.
The commenter's alternative
What they proposed was just showing the time with no value judgement.
Which is what this site's database does — it records strings like "400 hours at 10% in DPG", with no adjectives.
That approach has costs too, and this piece has already shown them:
- The raw numbers pile up, and a reader has no way to know that an 8-hour value means "about eight hours" rather than "measured at 8.0".
- Measurement conditions are inconsistent. Some neat, some at 10% dilution, some with a greater-than sign. We've written about the two ways longevity gets measured; sorting the numbers directly produces a wrong ordering.
- The 400-hour column is a truncation point, not a measurement.
So "no labels" is not the same as "neutral". An unlabelled numeric column still pushes readers toward "bigger is better" — it just hides the push.
The genuinely honest version shows the measurement conditions next to the number. And that is hard to make look good.
Collecting it up
- The line has no natural position. The data is columnar, the columns sit on round numbers, and round numbers are also what people use as boundaries.
- The eight-hour line cuts through 92 records.
- The top of the scale is worse: 334 records at 400 hours, which is a truncation point.
- Labels affect how people use a scale, and that has been studied; whether labels changed a generation's expectations is something I could not check.
- Leaving labels off is not neutral either. It hands the value judgement to a reader who does not have the measurement conditions.
And for anyone actually making something, the most practical line is still the one from last round: write down the time, on your own skin, come back in a few hours. Your own records need no labels, because you know how they were measured.
What this doesn't establish
I have not verified how Fragrantica displays longevity now or in the past. That is the commenter's account; I did not visit the site or check its revision history. This piece addresses the general problem of grading a continuous quantity, not any particular site's practice.
The 2,027 substantivity records come from inconsistent measurement conditions. Some neat, some at 10% or 20% in DPG or benzyl benzoate, as covered elsewhere. This distribution pools all of them, which does not affect the observation that values pile on round numbers, but does mean the 72-hour median is not a meaningful physical quantity.
The 210 distinct values include decimals. Values like 16.6 and 4.5 each count as one. The piling is worse than that number suggests.
I did not re-count how many of the 334 records at 400 hours carry a greater-than sign. An earlier piece counted 206 such records database-wide; the two numbers cannot be subtracted, because the scopes differ.
I read only the abstract of Blais and Grondin. Their own work concerns the adverb "totally", and what I quote is their introduction's summary of existing research rather than their own experimental result.
Krosnick's review is from 1999, and I read only the abstract.
Both papers are survey methodology and measure respondents answering. The forum is talking about readers forming expectations. Flagged in the body and repeated here: I have no direct evidence.
"No labels is not neutral either" is my argument rather than a citation. It rests on the data earlier in this piece, not on any study.
References
J. G. Blais, J. Grondin, The influence of labels associated with anchor points of Likert-type response scales in survey questionnaires, Journal of Applied Measurement, 12(4), 370-386 (2011). PMID 22357158
J. A. Krosnick, Survey research, Annual Review of Psychology, 50, 537-567 (1999). PMID 15012463. doi:10.1146/annurev.psych.50.1.537
Related: the 400-hour ceiling, two ways of measuring longevity, the eight-hour standard, three habits when evaluating.