Beginner perfumer · 39
What thirty materials buy you — worked out against 954 published formulas
· 28 min read
A student in Japan asked r/DIYfragrance for the cheapest way into the hobby, since shipping from overseas alone runs to $50. The most concrete answer was buy about 30 materials and think hard about which ones. That advice can be checked. Against 954 published formulas, owning the 30 most-used materials covers 35.2% of the median formula's weight; in a 16-material formula you would hold 5 of them. But the same corpus says something else: of 2,381 distinct materials, 1,065 appear in exactly one formula, and all the materials appearing in two formulas or fewer add up to a median 3.3% of a formula's weight. The tail is long and light.
Someone on r/DIYfragrance asked a question with the budget written into it:
"I've always wanted to create my own scent, but I can't find anywhere in my country that sells local aroma chemicals (I live in Japan). When I try to order from overseas, the shipping costs can reach $50. Aroma chemicals are already not cheap, and shipping ends up making the total more than I can afford. I know this isn't the cheapest hobby, but as a student, I'm wondering what the most affordable way to get started would be."
Seven replies came in. The two most concrete:
"Buy maybe 30 materials and think well about what you want/need, and what's less essential at this time. The seduction is to buy a lot of materials — too many in fact — because they look interesting." (kriebelrui)
"One 50-100g × 0.001g jeweler's scale. 100 various 1mL-4mL glass vials. 200 disposable plastic droppers. One litre of 95% or higher ethanol. 250 smell strips. Fewer than 50 materials. Labels, notebook, isopropyl alcohol for cleaning." (_nate69)
Thirty. Under fifty. Both numbers can be checked against the corpus: what does owning 30 materials actually let you do?
The short version
- Against 954 deduplicated published formulas, owning the 30 most frequently occurring materials covers 35.2% of the median formula's weight. A typical formula has 16 materials; you would hold 5 of them.
- Only 292 formulas (30.6%) are half covered. Only 40 (4.2%) reach 80%. None can be made completely.
- You have to reach 100 materials before median coverage climbs to 60.4%. At 200 materials, still only 14 formulas (1.5%) are ones you could make in full.
- The same corpus says something else. Of 2,381 distinct materials, 1,065 appear in exactly one formula (45%). And all the rare materials — those appearing in two formulas or fewer — together account for a median of 3.3% of a formula's weight.
- Which 30 matters more than how many. Picking 30 by how often they appear gives 35.2% median coverage; picking 30 by how much weight they contribute gives 39.8%, for the same 30 slots. The two lists overlap on only 21 materials.
- And on the perception side there is a ceiling. In a 1998 study, 41 subjects tried to name the components of mixtures; regardless of which odour set was used, identification stalled at around four components.
The coverage curve
| Materials owned | Median formula weight covered | Formulas ≥50% covered | ≥80% | Complete |
|---|---|---|---|---|
| 10 | 15.5% | 87 (9.1%) | 10 | 0 |
| 20 | 28.8% | 214 (22.4%) | 21 | 0 |
| 30 | 35.2% | 292 (30.6%) | 40 | 0 |
| 50 | 40.5% | 374 (39.2%) | 75 | 0 |
| 100 | 60.4% | 615 (64.5%) | 191 | 3 |
| 200 | 72.6% | 758 (79.5%) | 340 | 14 |
The column worth staring at is the last one. At 200 materials, 1.5% of formulas become buildable. Published formulas are not written to be reproduced from a small palette; each one carries a thing or two that you don't have.
If the goal is "build a published formula as written", then 30 is not enough, and neither is 200. That goal does not sit anywhere near 30 materials.
But if the goal is "make something", the table reads in the other direction.
The tail is long and light
Of 2,381 materials, 1,065 appear once. These are the names that look enticing online — they are exactly the ones kriebelrui meant by "they look interesting".
Computed: within a single formula, all the rare materials — appearing in two formulas or fewer — sum to a median 3.3% of the weight.
Half the material names you are tempted by account for 3.3% of a formula's weight between them.
One honest caveat, because weight and perception are different things. Something at 0.3% can dominate the impression of an entire perfume — β-damascenone and dimethyl sulfide, both covered in this series, are exactly that kind of material. So light is not the same as unimportant. The number says one thing only: if you are buying by weight, buy the tail last.
Which 30 beats how many
The first time I ran this, I ranked the top 30 by "appears in the most formulas". It is the intuitive move, and it was wrong.
Frequency does not distinguish one thing: plenty of materials appear constantly and are used at 0.5% every time.
Reranking by total weight contributed across all formulas, still taking 30, moves median coverage from 35.2% to 39.8%. Same number of bottles, 4.6 points more coverage, at zero cost — just a different sort order.
The difference between the lists is instructive.
In the top 30 by frequency, not by weight (buy these and you get a shelf of "one drop is plenty"):
| Material | Formulas | Median share |
|---|---|---|
| cis-3-hexenol | 100 | 0.54% |
| Indole | 88 | 0.77% |
| Vanillin | 137 | 1.02% |
| γ-undecalactone | 85 | 1.56% |
In the top 30 by weight, not by frequency (buy these and you get the structure):
| Material | Formulas | Median share |
|---|---|---|
| Phenylethyl alcohol (the entries recorded under that spelling) | 49 | 15.38% |
| Hexylcinnamic aldehyde | 33 | 16.16% |
| Benzyl alcohol | 43 | 10.00% |
| Ethylene brassylate | 59 | 9.57% |
| Hexyl salicylate | 57 | 7.05% |
| Lavender oil | 55 | 6.06% |
The starter materials recommended on forums are almost all from the first table: strong, distinctive, effective in a single drop. That is reasonable, since those are the most interesting things to smell on their own. But buy only the first table and you end up with a shelf of characters and nothing to put them in.
Hexyl salicylate, from the second table, is the type case. The database rates its strength as low and lists its odour as fresh, herbal, orchid, green — it reads as useless. It appears in 57 formulas at a median 7.05%, above 5% in 40 of them, and once at 84.3%. You do not buy it to be smelled; you buy it to hold everything else. It is this round's material guide.
The ceiling on the perception side
Suppose money is not the constraint and you buy 200. Can you tell them apart?
In 1998, Livermore and Laing ran an experiment that measures exactly this. 41 subjects, two odour sets: one chosen by an expert panel as good blenders in mixtures, the other as poor blenders. A computer-controlled air dilution olfactometer delivered a single odorant or a mixture of up to eight, and subjects had to name what was in it.
The result: the poor blenders were more easily discriminated, but that advantage held within a narrow range, and with either set, identifying mixture components was limited to approximately four.
Their conclusion is blunt: odour type changes which odorants get perceived in a mixture, but the limited capacity to discriminate mixture components is independent of the type of odorants.
Four keeps coming up in olfactory research. When Laing's group tested taste and taste-odour mixtures in 2002, they also stopped at three to four. This looks less like insufficient training than like the capacity of working memory.
So 200 materials will not let you hear 200 layers. They give you 200 ways to fill those four seats.
What the cheapest input actually is
Salt-Stone's reply was the only one of the seven that costs nothing:
"Watching YouTube videos from professional perfumers or educators. Reading books about perfumery. Acquiring samples of mass market scents and training your nose. Spending time on this subreddit."
"Train your nose" reads like the kind of advice that sounds good and contains nothing. It has been measured.
In 2023, Li, Anne and Hummel published a study in Chemical Senses: 100 healthy participants in four groups. All but the controls (n = 26) performed olfactory training four times a day for three months. Sniffin' Sticks tests (threshold, discrimination, identification) at baseline and after.
The training odours were randomly assigned as either single molecules or complex mixtures.
Two of the results bear on this question.
- The groups that improved were "video" and "counter" — the first watched congruent video while training (multisensory integration), the second additionally counted the number of odours one day a week (forced attention). The training-only group, with no additional measure, did not show the same improvement against controls.
- Odour complexity had limited effects. Training on single molecules and training on complex mixtures did not separate.
The second point matters to a student who cannot buy fifty materials: practising on single molecules measures out about the same as practising on complex ones. You do not need expensive naturals to train a nose.
The first point matters more, and it is not cheap. What it asks for is not money but attention. Over the same three months of sniffing, the group that added "count how many you smelled" improved; the group that only put things under its nose did not.
A limit on this paper: it measures olfactory function (threshold, discrimination, identification), not perfumery skill. Nobody has run a randomised trial on "formulas improved after three months of training", and I don't expect one. What it supports is the first half only: the sensitivity of a nose changes with attentive repeated exposure, and that costs nothing.
What I got wrong this round
I ranked the top 30 by frequency of appearance. Coverage came out at 35.2%, the number looked reasonable, and I started writing.
Halfway through I went back to check whether the sort order was right, reranked by weight, and 35.2% became 39.8%. Those 4.6 points did not come from new data; they were what my original method had been dropping. I had been treating "appears often" as "matters", and a large share of the materials that appear most often are used at half a percent each time.
The mistake is worth writing down, because it is the same mistake the forum recommendation lists make. People recommend the materials that made an impression, and impressive usually means strong and used in small amounts. The layer that carries a formula's weight never shows up on anyone's starter list, because those materials are boring on their own.
The actual answer for that student
I can't solve the shipping. I don't know which suppliers inside Japan are reliable, and this series doesn't make recommendations it hasn't checked. That part I can't answer.
Here is the part that is usable:
- Thirty materials will not reproduce published formulas. Neither will 200 (1.5%). If that is the expectation, adjusting it matters more than buying anything.
- Buy by weight, not by odour strength. You need the things that hold — phenylethyl alcohol, benzyl alcohol, hexyl salicylate, lavender oil — not a full set of the things where one drop is plenty. Across the same 30 slots, that ordering is worth 4.6 points of coverage.
- Buy the rare materials last. They are 45% of the names and 3.3% of the weight. (Remembering that something at 0.3% can dominate a perfume; this ordering is about starting, not about forever.)
- Component identification tops out around four. Two hundred materials give you choices, not layers.
- The item that costs nothing has measurements behind it: attentive repeated exposure improves olfactory function, and single molecules work about as well as complex mixtures.
The person who said "buy maybe 30" had the direction right; the number just sounds like a restriction. From the corpus, 30 is not a compromise — it was already on the same side of a line that 200 doesn't cross either. The only difference is how well those 30 are chosen.
What this doesn't establish
- The corpus is not the population of formulas. These 954 published formulas skew towards what gets published: classic reconstructions, teaching examples, thread exercises. Real commercial formulas need not distribute this way.
- Weight is not perception. Every coverage figure here is weight coverage. Holding 35% of the weight does not mean it smells 35% right. It might smell closer, or nothing like it.
- The greedy pick is not optimal. I sorted by total weight contributed, which is an approximation, not the strictly optimal set of 30. Finding that needs combinatorial optimisation, which I did not do.
- Cost does not enter the calculation. The corpus has no prices. High weight contribution does not imply cheap, although this particular set (phenylethyl alcohol, benzyl alcohol, hexyl salicylate) is broadly commodity chemistry.
- The 1998 study measured identification, not preference or creation. Failing to name components does not mean the mixture has no effect, or that more materials are pointless.
- The 2023 study's participants were normosmic members of the public, not perfumery apprentices. It supports "sensitivity changes", not "you will become a perfumer".