Beginner perfumer · 117
Beginner perfumer #117: when both ends of a spec are round numbers, the range is three to five times wider
· 19 min read
Refractive index and specific gravity are the two fields in this database that look most like measurements. Sweeping four thousand-odd materials, the median range width is 0.0060 for refractive index and 0.0080 for specific gravity. Then a signal turned up that needs no chemistry at all: when both endpoints are round to two decimal places, as in 1.40000 to 1.60000, the median width is 0.0200, which is 3.3 times the rest. For specific gravity it is 0.0300 against 0.0064, a factor of 4.7. Round endpoints mean the interval was typed rather than measured. Following that line, 38% to 54% of the wide ranges belong to records whose names contain fragrance or specialty: compounded blends, which cannot have a single specification in the first place. I also tested one hypothesis, that unpopular records go unchecked. It failed. The correlation between supplier count and spec width is only -0.055.
About a 5 minute read.
The short version
- 4,094 refractive index values parse, with a median range width of 0.0060; 4,181 specific gravities, median 0.0080
- Records whose endpoints are both round to two decimals run visibly wider:
| Field | Round-endpoint records | Their median width | Everything else | Factor |
|---|---|---|---|---|
| Refractive index | 391 (9.6%) | 0.0200 | 0.0060 | 3.3× |
| Specific gravity | 496 (11.9%) | 0.0300 | 0.0064 | 4.7× |
- 38% to 54% of the wide-range records have
fragrance,specialtyorflavorin the name: blends - Endpoints written in descending order: 9 for refractive index, 18 for specific gravity
- I tested the hypothesis that unpopular records go unchecked. It failed (rho = −0.055 and −0.024)
The shape the last round left
Writing up folded sweet orange oil I hit a problem: one record has to cover every product from 5-fold to 50-fold, and its optical rotation interval is thirteen degrees wide, which cannot.
Ylang's four grades are the same shape.
After twice, the question is: can this "one record describing a whole spectrum" situation be seen directly from the width of a specification range?
A baseline first
Refractive index and specific gravity are the two fields here that look most like measurements. They are physical constants, unlike substantivity or odour, which are subjective.
| Field | Parseable | Median width | Quartiles |
|---|---|---|---|
| Refractive index | 4,094 | 0.0060 | 0.0040 – 0.0100 |
| Specific gravity | 4,181 | 0.0080 | 0.0060 – 0.0140 |
A pure compound's refractive index is one number. A real specification allows a little room for measurement and purity, and around 0.006 is what that room should look like.
Round endpoints are a signal
Look at these:
egyptian musk fragrance 1.40000 to 1.60000
rosemary fragrance 1.40000 to 1.58000
apple fragrance 1.34000 to 1.50000
longifolene 1.42000 to 1.60000
ginger root oil cochin 0.87000 to 1.50000
Both endpoints round to two decimals. No measurement lands that neatly.
I swept the database with that as a rule, and the result came out cleaner than expected:
| Field | Round endpoints | Median width | Everything else | Factor |
|---|---|---|---|---|
| Refractive index | 391 (9.6%) | 0.0200 | 0.0060 | 3.3× |
| Specific gravity | 496 (11.9%) | 0.0300 | 0.0064 | 4.7× |
You do not need to know what the material is. Look at whether the two endpoints are round, and you can tell whether the interval was measured or typed.
This is the same method I used to find
the default assay value 95.00 to 100.00:
not checking whether a value is right, but checking whether its shape looks like a measurement.
Half of the wide ranges are blends
Take the records more than four times the median width and look at their names:
| Field | Wide-range records | With fragrance / specialty / flavor in the name |
|---|---|---|
| Refractive index | 283 (6.9%) | 153 (54%) |
| Specific gravity | 486 (11.6%) | 183 (38%) |
egyptian musk fragrance, rosemary fragrance, apple fragrance,
jasmin specialty, herbal specialty, peach fragrance.
Those are not materials; they are blends. Every supplier's "apple fragrance" is a different formula, so that record cannot have one specification. Writing 1.34 to 1.50 says only "somewhere around here".
The problem is not that the numbers are wrong. It is that the field is meaningless for that kind of record.
One pure compound that does not belong on that list
Longifolene (#22800, 16 suppliers) is a single compound, C15H24, CAS 475-20-7.
Its specifications:
Refractive index 1.42000 to 1.60000 (width 0.18, thirty times the median)
Specific gravity 0.93000 to 1.11000 (width 0.18)
Optical rotation +14.00 to +54.00 (width 40 degrees; the median is 10)
Three fields, all with round endpoints, all anomalously wide.
The other sesquiterpenes with the same molecular formula look like this:
| Material | Suppliers | Refractive index | Specific gravity |
|---|---|---|---|
| beta-caryophyllene | 61 | 1.498 – 1.504 | 0.899 – 0.908 |
| valencene | 44 | 1.498 – 1.510 | 0.910 – 0.922 |
| alpha-cedrene | 19 | 1.500 – 1.503 | 0.926 – 0.937 |
| alpha-humulene | 16 | 1.499 – 1.505 | 0.889 – 0.895 |
| (−)-isolongifolene | 11 | 1.495 – 1.501 | 0.926 – 0.932 |
| longifolene | 16 | 1.420 – 1.600 | 0.930 – 1.110 |
Its own structural isomer, (−)-isolongifolene, carries a spec thirty times tighter.
And every well-behaved sesquiterpene's refractive index clusters between 1.495 and 1.510. Longifolene's true value must be in there too. The interval 1.42 to 1.60 contains it and tells you nothing.
I tested a hypothesis and it failed
Seeing longifolene with 16 suppliers and a spec this bad suggested an explanation: unpopular records go unused, so nobody notices they are broken.
That is testable. I grouped spec widths by supplier count:
| Suppliers | Median refractive index width | Median specific gravity width |
|---|---|---|
| 0 | 0.0078 | 0.0120 |
| 1–4 | 0.0060 | 0.0070 |
| 5–19 | 0.0060 | 0.0070 |
| 20–49 | 0.0060 | 0.0080 |
| ≥50 | 0.0060 | 0.0080 |
Apart from the zero-supplier group running slightly wide, it is completely flat. The correlations are rho = −0.055 (refractive index) and −0.024 (specific gravity), which is no relationship at all.
The hypothesis fails. How good a spec is has nothing to do with how many people sell the material. Longifolene's 16 suppliers match alpha-humulene's, and alpha-humulene's spec is perfectly normal.
(What survives is the zero-supplier group, and the fact that round endpoints appear on only 4% of the ≥50 group against 9% to 14% elsewhere. The most popular materials do use round endpoints less often, but not by much.)
This is the second hypothesis of mine I have refuted in public over these rounds. The last was the substantivity field's long end.
Three rules you can run yourself
Open any material's page:
- Are both ends of the refractive index or specific gravity round numbers?
(
1.40000 to 1.60000) → that was typed - Is the range wider than 0.01? The refractive index median is 0.006, and anything past 0.02 is in the top 11%
- Does the name contain
fragrance,specialtyorflavor? → it is a blend, and the spec was never meaningful
The first is the most useful because it needs no baseline. You do not have to know what a normal refractive index range is; you only look at whether the two numbers are suspiciously tidy.
And if you have sibling materials to compare against, as in the sesquiterpene table above, that is the strongest check of all: two materials with the same molecular formula should not have specifications thirty times apart.
What this doesn't establish
- I looked up no correct refractive index or specific gravity. All of this compares shapes and distributions.
- The "round endpoint" test is one I defined: exactly two significant decimal places. A different definition, three places say, moves the numbers.
- The median width of 0.0060 belongs to this database, not to instrument precision or any industry standard.
- The blend section judges by name. Some
specialtyrecords may be single materials; I did not check them one by one. - For longifolene I only showed it is inconsistent with its family, not which number is correct.
- I tested the supplier-count hypothesis on two fields only. Other fields, such as substantivity or boiling point, might behave differently; I did not try.
- Automated curation of chemical datasets has its own literature. Mansouri and colleagues' 2016 procedure (PMID 27885862) addresses this class of problem; what I ran here is three rules you can check by eye.