Beginner perfumer · 27
His own control test lied to him for a year: a case of the wrong suspect
· 31 min read
Someone spent a year reformulating a perfume around salicylates, because he had run an isolation test and it confirmed his suspicion. Three strangers said it sounded like veramoss. He retested: the earlier test had been contaminated, and the salicylates ran 16 hours with no bitterness. The database explains why contamination runs one way — veramoss is rated high strength and lasts 400 hours at 10% dilution, while benzyl salicylate is rated low and lasts 384 hours neat. Across 954 formulas, high-strength materials sit at a median dose of 0.80% and low-strength at 5.78%. Also: no published formula uses Calone above 2.50%, and he used 25%.
Someone posted a request for help on r/DIYfragrance, and the first line reads:
"See end of post for update."
That line was added later. What happened in between is worth more than the original question.
He had been building a "90s nostalgia pool party" aquatic for a year. The opening pleased him, but six to eight hours in, the drydown turned into what he described as bitter, dry, papery. He identified the culprit as the salicylates in his formula.
And he did not guess. He pulled the suspects out and tested them alone on skin. The result confirmed his suspicion: bitter in isolation too.
So he spent a year reformulating around that conclusion — scrapping the benzyl salicylate entirely, cutting the salicylate-carrying Solafleur by forty percent.
Then three people who did not know each other said the same thing in the comments: that sounds like veramoss.
He retested. The update at the bottom of the post reads:
"It absolutely had [been compromised]. I am not certain what may have happened, but the last time I tested solafleur and benzyl sal alone on my skin, I must have been playing with veramoss as well and had cross contam. After the new test, both performed gorgeously, lasting over 16 hours on my skin with no bitter drydown."
A year of reformulating, aimed at the wrong material. And what aimed him wrong was the act of running a control test.
The short version
- Contamination is one-directional in effect. The database rates veramoss high strength with substantivity of 400 hours at 10% dilution in DPG; it rates benzyl salicylate low, with 384 hours neat. The same trace residue lets the strong one bury the weak one, and never the reverse.
- The pattern is continuous across the corpus: in 954 formulas, materials labelled low have a median dose of 5.78%, medium 2.00%, high 0.80%. The stronger the label, the less goes in.
- A result that confirms what you already believed is the most dangerous kind of evidence, because contamination favours whatever you have been handling — and what you have been handling is usually the thing you suspect.
- He used 25% Calone. Of the 954 public formulas, 22 use it, at a median of 0.955% and a maximum of 2.50%. He is at ten times the highest published figure.
- He suspected himself of going nose-blind to the Calone. Olfactory adaptation has been measured (Chen et al. 2020, n = 4,120).
- "Do I just dislike salicylates" is a real question too: the relationship between pleasantness and individual detection threshold is odour-specific (Bontempi et al. 2022).
About 9 minutes.
Why contamination only runs one way
Start with the two materials.
| veramoss | benzyl salicylate | |
|---|---|---|
| Strength | high, recommended smelling at 10% or less | low |
| Substantivity | 400 hours at 10% in DPG | 384 hours at 100% |
| Suppliers | 40 | 58 |
| Median corpus dose | 0.80% | 9.47% |
The two substantivity figures look similar, and the concentrations behind them differ by a factor of ten.
Veramoss holds for 400 hours at a tenfold dilution. Benzyl salicylate holds for 384 hours neat. To reach comparable persistence, the first needs an order of magnitude less material than the second.
That is why cross-contamination travelled in the direction it did. Leave the same tiny residue of each on skin: the veramoss residue is still above threshold, and the benzyl salicylate residue stopped being perceptible long ago.
Contamination is not two materials interfering with each other. It is the strong one burying the weak one.
Nor is this specific to these two. Matching the corpus against the database strength labels, restricted to materials appearing in at least 5 formulas:
| Strength label | Materials | Median dose | Interquartile range |
|---|---|---|---|
| low | 6 | 5.78% | 3.50-9.47% |
| medium | 178 | 2.00% | 1.00-4.60% |
| high | 54 | 0.80% | 0.56-1.00% |
A clean ladder. The stronger the label, the smaller the share that goes into a formula, by roughly a factor of seven end to end.
That ladder is your contamination risk table. The materials you dose at 0.8% are exactly the ones liable to show up uninvited in your next test.
(We've written that strength and diffusion are different things. This uses the strength half: the strength label predicts how much you can use, not how loud it smells.)
Why the test confirmed him
What he did looks procedurally correct: suspect a material, pull it out, test it alone.
The problem is in the ordering, not the method.
Three separate commenters pointed at veramoss. One put it most directly:
"'Bitter, dry, papery' sounds like what a Veramoss overdose will do."
And his own reply gives away the mechanism: veramoss is 0.5% of his formula, and worried about that dry bitterness he had already been cutting it back — meaning he had been handling the veramoss bottle around the time of that isolation test.
He was testing the salicylates, and the thing he had touched most that day was veramoss.
A result that confirms your existing hypothesis needs stricter checking than one that refutes it, because the source of contamination is usually the subject of the investigation. It is the bottle on your bench, the one recently opened, the pipette just used.
I walked into this one myself last round. Computing corpus statistics, I first used a deduplication that differed from the one the earlier articles used, and it made phenethyl alcohol come out "second most common" — a better-sounding result than the correct one, and I nearly put it in a title. Checking the old script showed the dedup conditions differed and the right answer was third. That was the same shape of error: an expectation first, then a method that produced it. My cost was a recalculation, and his was a year.
About that 25% Calone
The thread has a second strand worth separating out, because it has numbers attached.
He wrote that he uses "upwards of 25%" Calone. A commenter replied:
"Calone at 25% of your formula is extremely overdosed. That stuff is quietly powerful. If you go much above 0.5% of a formula, you are courting disaster... 25% is just crazy."
The corpus agrees with them. Of the 954 public formulas, 22 use Calone or watermelon ketone:
- median dose 0.955%
- highest single formula 2.50%
No published formula goes above 2.5%. He is at ten times the maximum.
(We covered Calone last round: substantivity > 600 hours at 10% dilution, in the top 0.84% of the whole database. Like veramoss, it belongs to the "high, recommend smelling at 10% or less" class.)
And he raised the right suspicion about himself in a reply:
"Strangely, with my other calone-forward perfumes, I can smell the calone forever - after showering and sleeping even, but with this one i can barely smell it at all. Perhaps I need to spray a piece of clothing with this and keep it in a different room and go back to smell it every couple of hours to see if I'm blinding myself to the calone."
That instinct points the right way.
Chen and colleagues analysed odour threshold testing data from 4,120 subjects in Rhinology in 2020, using the staircase technique of the Sniffin' Sticks and looking at the trajectory of turning points.
Their conclusion: people with poor olfaction seem to adapt faster to olfactory stimuli, and the trajectory of turning points in a threshold test may serve as an indicator of olfactory adaptation and receptor function.
Mind the population in that paper. Of the 4,120, some 1,684 had hyposmia and 1,742 anosmia, with only 694 normosmic controls. It is an ENT clinical study, not a study of perfumers. I cite it for the established fact that olfactory adaptation is measurable and trackable, not to transfer patient adaptation rates onto you.
His practical remedy is sound, though: another room, a few hours apart, come back and smell it. That is a washout period for his own nose.
"Do I just dislike salicylates"
That was the second question on his list. The retest turned the answer into no, and the question still deserves an answer, because many people ask it.
Bontempi and colleagues measured the relationship between hedonic ratings and individual detection thresholds in Scientific Reports in 2022. For two pleasant odorants (apple, jasmine) and two unpleasant ones (durian, trimethylamine), they first established each participant's own threshold, then presented concentrations above it in randomised order for hedonic rating.
The result: apart from trimethylamine, pleasantness and unpleasantness ratings rose with concentration at low and middle levels, then plateaued at high concentrations. When individual thresholds were taken into account, the correlation coefficient was significantly higher for trimethylamine.
Their summary: the relationship between hedonic ratings and individual detection threshold is odour specific, and this helps explain the large variability of hedonic tone at a specific concentration across the general population.
Two things follow for formulation:
- "This concentration smells good" is not a property of the material but of your relationship with it. A recommended dose from someone else is a dose recommended from their threshold.
- And this varies by material. Individual variation is not uniform across materials; only one of the four odorants in that study showed a difference once thresholds were accounted for.
So "do I dislike salicylates" is a reasonable question. It just requires first establishing that what you smelled was in fact the salicylates, and his case shows that step is harder than it looks.
A control procedure that won't lie to you
Here is the whole thing as something you can run. Every line blocks one route for contamination.
1. Test in ascending order of strength. Low first, then medium, then high. Once a strong material is on your hands, nothing else you test that day is clean. Use the ladder above to order your list.
2. Do not open the strong bottle at all that day. Not "not while testing" — not at all, for the whole session. Pipettes, gloves, bench, caps all count.
3. Carry a blank control. Take an untouched blotter and put it through the same hands, the same bench, the same elapsed time as the others. If you can smell anything on the blank, the whole round is void. This costs thirty seconds and it is the only step that detects contamination directly.
4. Put skin tests last, and switch arms. Skin retains material, and soap does not remove residues of high-strength materials. The residue from your last test on that arm is the contamination source for the next one.
5. Retest any result that confirms what you already thought. This is the most important line on the list and the easiest to skip. A result that refutes you gets checked automatically; a result that agrees with you does not.
6. Have somebody else label them, if you can. When you do not know which blotter is which, the contamination is still there and the expectation is gone.
7. And the thing he got right in the end: post the result where strangers can see it. Three of them gave the same answer, and that answer contradicted a year of his own work. He went and retested, which is the single most instructive move in the whole thread.
What this doesn't establish
I have smelled neither his formula nor the two blotters he tested. All of this is written against text he posted. Veramoss being the culprit is his conclusion from his own retest, not something I verified.
Strength labels are a database classification, not a measurement. Those strings ("high, recommend smelling in a 10.00% solution or less") are editorial judgements rather than threshold data. The ladder shows that labels agree with real dose levels, not that labels equal thresholds.
The substantivity comparison has a trap in it. The 400 hours for veramoss was measured at 10% dilution and the 384 hours for benzyl salicylate at 100% — two numbers produced under different conditions, so you cannot subtract one from the other. I use them to show a tenfold difference in concentration, not a 16-hour difference in longevity.
"Contamination is one-directional" is a claim about effect, not physics. Molecules travel both ways. The point is the asymmetry in perceptible intensity from an equal residue.
The 2.50% Calone ceiling is the ceiling of the public corpus, not a safety or recommended limit. 74% of the corpus comes from patents and published sources, so it describes what the industry has published, not what must not be exceeded. His 25% is his own business; I am only noting how far it sits from known practice.
The Chen study population is mostly patients with olfactory impairment, and I use it only for the general fact that olfactory adaptation can be measured. That paper tested neither perfumers nor Calone.
The Bontempi study covers four odorants, none of them a salicylate. I cite its conclusion that the hedonic-threshold relationship is odour-specific rather than applying it to salicylates.
I have not run a controlled trial of that seven-step procedure. It is an operational suggestion derived from the contamination mechanism, not a validated method.
References
B. Chen, A. Haehner, M. K. Mahmut, T. Hummel, Faster olfactory adaptation in patients with olfactory deficits: an analysis of results from odor threshold testing, Rhinology, 58(5), 489-494 (2020). PMID 32478337. doi:10.4193/Rhin19.465
C. Bontempi, L. Jacquot, G. Brand, A study on the relationship between odor hedonic ratings and individual odor detection threshold, Scientific Reports, 12(1), 18482 (2022). PMID 36323760. doi:10.1038/s41598-022-23068-1
Related: strong versus diffusive, Calone, which trials to keep, white particles in the bottle.