Beginner perfumer · 17
You bought thirty materials and froze — stop buying and go look at which two turn up together
· 25 min read
Someone ordered thirty-odd materials in one go and got home with no idea what to do. I computed pairwise co-occurrence across 953 public formulas: after merging synonyms, eight clusters fall out of the 132 common materials. Two things surfaced along the way — my highest-lift pair turned out to be a spelling artifact, and 69% of the pairs that really co-occur are absent from the database's official blend suggestions, oakmoss with patchouli and bergamot with lavender among them.
Someone posted this to r/DIYfragrance: I wanted to make my own perfume, so I ordered thirty to fifty essences and oils along with bottles, alcohol and DPG. And then I had no idea what to do with them.
He started by mixing at random out of pure excitement, realised after the first attempt that this was a bad idea, and switched to smelling materials one at a time. Most either smelled faint or smelled the same. Diluting to 10% made them readable, but then how do you calculate anything? And DPG — people seem sharply split on it. He signed off: I got overwhelmed. What steps do you actually follow?
Forty-nine replies. This is that question answered with data.
The short version
- Stop buying. The thirty you own are enough. The problem is that you are smelling them as thirty singles rather than as pairs.
- Computing pairwise co-occurrence across 953 public formulas, and merging synonyms, eight clusters fall out of the 132 common materials. Find the cluster where you already own three members. That is your first chord.
- Two papers explain why pairs are the right starting unit: the adult nose processes binary mixtures as a whole, and once a mixture passes seven or eight components only a third to a half of people can pick a given one out.
About a 9 minute read.
Why pairs rather than singles
This comes first because it decides everything downstream.
Romagny and colleagues, in Developmental Psychobiology in 2026, compared children, teenagers and adults (mean ages 10, 15 and 39) on odour mixture perception, using two binary mixtures and two senary mixtures across three tasks: typicality rating, target detection, and complexity evaluation.
The senary mixtures behaved consistently across all three age groups. The difference showed up in the binary mixtures: adults rated AB as more typical of the configural odour than children did, while showing reduced access to the individual components within binary mixtures, with teenagers in between. The authors read this as a progressive shift from more elemental toward more configural processing, most visible in mixtures of lower chemical complexity.
In plain terms: you are an adult, so when two materials meet, your nose hears a new chord rather than two notes. That is not a defect. That is the adult default.
And with more components? Drnovsek and colleagues, in the European Journal of Neuroscience in 2025, tested 90 healthy participants and 40 patients with olfactory dysfunction on picking a named target odorant out of a mixture. The established background: detection gets harder as background odorants increase, and falls below chance in mixtures of 16 components.
Their own result: even with seven or eight odorants in the mixture, around 50% of healthy participants still found eugenol and around 30% to 40% found phenylethanol. Both distributions differed significantly from chance (p < 0.001), so this is not guessing — but read it the other way and half the room has already lost the thread at seven or eight materials.
Put the two together and the conclusion for our overwhelmed poster is direct: two is a learnable unit. Thirty is not.
Who turns up with whom across 953 formulas
The method is simple. Take the same public demo formula corpus, treat each formula as a set of materials, count how often any two appear together, and divide by how often they would co-occur if they were independent. That ratio is usually called lift; above 1 means they turn up together more than coincidence explains.
First, the pairs with the highest raw counts.
| Formulas | Lift | Pair |
|---|---|---|
| 153 | 1.36 | Benzyl acetate + linalool |
| 138 | 1.51 | Benzyl acetate + phenethyl alcohol |
| 125 | 1.70 | Linalool + linalyl acetate |
| 99 | 1.87 | Benzyl acetate + benzyl salicylate |
| 91 | 2.05 | Citronellol + geraniol |
| 88 | 1.90 | Hydroxycitronellal + phenethyl alcohol |
| 88 | 1.94 | Alpha-terpineol + linalool |
These are backbone pairings, frequent because both halves are frequent. To find what actually sticks together, look at lift.
| Lift | Together | Pair |
|---|---|---|
| 7.69 | 24 | Ethyl 2-methyl butyrate + gamma-undecalactone |
| 6.03 | 27 | Cis-3-hexenyl acetate + hexyl acetate |
| 5.26 | 23 | Gamma-undecalactone + Verdox |
| 4.90 | 31 | Allyl amyl glycolate + Iso E Super |
| 4.70 | 39 | Cis-3-hexenol + cis-3-hexenyl acetate |
| 4.49 | 23 | Oakmoss absolute + patchouli oil |
| 4.34 | 37 | Allyl amyl glycolate + dihydromyrcenol |
| 4.08 | 36 | Heliotropin + para-anisaldehyde |
The second table's pairs are not especially common, but when one appears the other nearly always does. Those are chords.
I got this wrong first. Before merging synonyms, the top lift pair was
linalol+phenylethyl alcoholat 10.65. Those are alternate spellings of linalool and phenethyl alcohol. The real reason they "co-occur unusually often" is that one author writing one formula uses their own spelling for both names. That is a signal about documents, not about smell. Merging made the pair disappear.
Eight clusters
Chain the high-lift pairs together and groups emerge. I named them by hand, then went back and measured how much of the corpus each covers.
| Cluster | Example members | Formulas with ≥2 | With ≥3 |
|---|---|---|---|
| Jasmine / white floral | Benzyl acetate, alpha-hexyl cinnamaldehyde, hedione, indole, benzyl salicylate | 31% | 12% |
| Lavender / fougère | Linalool, linalyl acetate, coumarin, bergamot, lavender oil | 28% | 9% |
| Muguet | Hydroxycitronellal, linalool, citronellol, Lilial, Lyral | 27% | 8% |
| Rose | Geraniol, citronellol, phenethyl alcohol, rhodinol, nerol | 24% | 9% |
| Modern fresh woody | Iso E Super, dihydromyrcenol, allyl amyl glycolate | 10% | 4% |
| Fruity esters | Gamma-undecalactone, hexyl acetate, cis-3-hexenyl acetate, gamma-decalactone | 8% | 4% |
| Chypre / moss-wood | Oakmoss, patchouli, vetiveryl acetate, musk ketone, sandalwood | 8% | 3% |
| Aldehydic | Aldehydes C-8, C-9, C-10, C-11, C-12 | 5% | 2% |
The first four are general-purpose backbone, each appearing in a quarter to a third of formulas. The last four are much narrower and internally tighter — signature sets for particular styles rather than a general chassis.
What to do with it: hold your thirty bottles against this table and find the cluster where you already own three members. That cluster is tonight's exercise. Put the other twenty-seven away.
The official "blends well with" list doesn't match practice
This site's database carries an official blend-suggestion dataset covering 7,129 materials across 673,662 records. With measured co-occurrence in hand, the two can be compared.
Of the 601 measured pairs, 326 have both halves matched to a database record. The result:
- Present in the official suggestions: 102 pairs, 31%
- Absent: 224 pairs, 69%
Among that 69%:
| Lift | Pair |
|---|---|
| 4.49 | Oakmoss absolute + patchouli oil |
| 4.25 | Musk ambrette replacer + musk ketone |
| 3.78 | Musk ketone + sandalwood oil |
| 2.89 | Patchouli oil + sandalwood oil |
| 2.66 | Bergamot + lavender oil |
Oakmoss with patchouli is the floor of the chypre. Bergamot with lavender is the floor of the fougère. Neither appears in the official suggestion list.
The likely reason: those suggestions are built per material on odour adjacency, answering "what smells like it belongs beside this". Co-occurrence in formulas reflects structural division of labour — oakmoss and patchouli appear together not because they smell alike but because one supplies moss and the other supplies earth, and the pair is what the chypre base is made of.
So the two datasets answer different questions. For odour-adjacent neighbours, use the official list. For what actually gets written into the same formula, use co-occurrence.
The poster's other questions
He asked two more things worth answering.
"I can only smell them at 10%, so how do I calculate anything?" Work by weight; the stock concentration is only a conversion factor. Weighing and dilution walks the full workflow. The key idea is that the formula sheet always states grams of neat material, and a 10% stock is just how you hit that number.
"DPG or not?" Both camps are half right; it depends on what you are making. Solvents has the comparison. The short answer for his situation: DPG is excellent for stocks you intend to smell, and alcohol is what you want for something sprayable.
What this doesn't establish
74% of the corpus carries a patent or source row. A patent example is not a commercial formula, and co-occurrence reflects the writing habits of this set of public documents.
Synonyms were merged by hand. I handled the ones I could recognise — spelling variants, trade versus systematic names, some of the origin variants of naturals. Unmerged synonyms keep producing the kind of fake lift described above, so read the cluster table as a starting point rather than a settled result.
That 31% carries a bias risk. Of the 601 pairs, 275 could not be compared because a name failed to match a database record or the material had no suggestion data. If those failures skew systematically toward one side, 31% moves. What the figure supports is "the two datasets overlap less than you would assume", not a precise overlap rate.
Lift is not causation. Two materials appearing together may be complementary, or may simply both be standard issue for one era or one genre. This data cannot separate those.
References
E. Drnovsek et al., Detection of odorants in odour mixtures among healthy people and patients with olfactory dysfunction, European Journal of Neuroscience, 61(1), e16633 (2025). PMID 39803925. doi:10.1111/ejn.16633
S. Romagny, G. Coureaud, T. Thomas-Danguin, Developmental Switch From Elemental to Configural Processing of Odor Mixtures in Humans, Developmental Psychobiology, 68(5), e70188 (2026). PMID 42552896. doi:10.1002/dev.70188
Related: The first ten materials, Weighing and dilution, The first accord, The starter palette, mapped.