Reading the data · 14
Can jamming GC-MS protect a formula? I measured how formula weight is distributed
· 29 min read
Someone posted an anti-dupe technology: a colourless, odourless base added to a fragrance to interfere with GC-MS, which across 133 tested ingredients caused 93 to be incorrectly quantified and 25 incorrectly identified. A GC analyst who reverse-engineers fragrances at a Givaudan-owned company replied that they are only expected to get 70 to 90% of the way, with the perfumer's nose doing the rest. I measured how big that rest is across 953 formulas: the top three materials carry 54.7% of the weight, and knowing only the 344 most common materials covers 83% of a median formula's weight.
Someone posted a technology to r/DIYfragrance.
"They developed a colorless and odorless base that can be added to a fragrance to interfere with gas chromatography (GC-MS) analysis. It can cause certain materials to be misidentified or incorrectly quantified, making it significantly harder to reproduce the original formula. They tested it on 133 natural and synthetic ingredients. Apparently, 93 were incorrectly quantified and another 25 were incorrectly identified during analysis."
The poster asked: does this mean the dupe industry is on its way out?
Those three figures sound decisive, and the most valuable reply in the thread put them in context. Someone describing themselves as a GC analyst at a Givaudan-owned company that recreates fragrances for clients wrote:
"As an analyst, we're only really expected to get the fragrance about 70-90% of the way and the perfumers will do the rest of the work just using their nose usually so something that makes it slightly harder to get a good match won't be the end of the world, just means us analysts will take longer doing the ID and/or formulation steps."
How big is that rest? That question is computable, so I computed it.
The short version
- Across 953 public formulas the median is 16 materials, and 9 of them carry 90% of the weight, with the top three carrying 54.7%.
- Materials dosed under 1% make up 14.3% of the count but only 1.01% of the weight. Formula weight is heavily concentrated in a few materials.
- The corpus contains 2,380 distinct materials. But knowing only the 344 most common covers 83.0% of a median formula's weight — and that step needs no instrument at all.
- The 200 most common cover 72.6%, the 100 most common 60.4%, the 20 most common 28.9%.
- The analyst says GC gets them 70 to 90%. A pure common-material prior lands in the same range. So jamming GC defends a channel that may carry only the last ten or twenty percent.
- A 2017 study demonstrates the other half: 80 odorants were identified in a class of liquor, and 27 of them were enough to recreate the aroma.
About a 9 minute read.
What formula weight actually looks like
First, why "the remaining 10 to 30%" cannot be read as "10 to 30% of the work left".
Taking 953 deduplicated public formulas with solvents removed:
| Median material count | 16 |
| Materials carrying 90% of weight (median) | 9 |
| Weight in the top three materials (median) | 54.7% |
| Share of material count dosed under 1% | 14.3% |
| Share of weight dosed under 1% | 1.01% |
Look at the last two rows. Those low-dose materials are a seventh of the list and a hundredth of the weight.
That cuts both ways for an interference technology, and both directions belong in the same paragraph.
One way: if those 93 misquantified materials land at the "top three carry 54.7%" end, the effect is large — reading a 10% material as 12% is a two-percentage-point absolute error, and a nose has a chance at that.
The other way: if they land at the "under 1%" end, reading 0.5% as 0.6% is a tenth of a percentage point. Inside a sixteen-material formula, almost nobody smells that.
So "93 of 133 incorrectly quantified" cannot on its own answer "is the formula protected". Answering that requires knowing which end got hit, and that information is not in those three numbers.
How much you can guess without an instrument
This is the calculation I found more worth doing.
Those 953 formulas draw on 2,380 distinct materials. At first glance, picking the right 16 out of 2,380 looks hard.
But usage frequency is extremely uneven. Ranking materials by how many formulas they appear in and taking the top N, here is how much of a formula's weight those N cover:
| Knowing only the top | Weight covered (median) | Materials hit (median) |
|---|---|---|
| 20 | 28.9% | 3 |
| 50 | 40.6% | 6 |
| 79 | 53.3% | 7 |
| 100 | 60.4% | 7 |
| 200 | 72.6% | 10 |
| 344 | 83.0% | 12 |
Knowing the 344 most common materials gives you 83% of a median formula's weight.
That 344 is not a number I picked at random. The Cai 2024 automatic formula generation study we cited before extracted composition from 210 working formulas and built its reference library from 344 common materials.
Put the two together: the analyst says GC gets them 70 to 90%. And a prior using no instrument at all — just knowing which materials the industry commonly uses — already lands between 72.6% (200 materials) and 83.0% (344).
That does not make GC useless. GC gives you specifics about this formula; a prior gives you what the average formula looks like, and they do not substitute for each other. The prior says "there is probably linalool in here"; GC says "there is 5.93% of it".
But it does change the shape of the problem. If protecting a formula is approached only through making GC read wrong, what is being defended is that one channel — while most of a formula's weight sits, as a matter of course, in a set of materials everybody knows.
Twenty-seven were enough
The analyst said the perfumers do the rest by nose. That step has been measured too.
Niu and colleagues, in Food Chemistry in 2017, investigated the aroma profiles of five Chinese light aroma-type liquors. Using gas chromatography-olfactometry (GC-O) and GC with flame photometric detection, they identified 80 odorants, including ten sulfur compounds.
They then used quantitative study and odor activity values to determine which were aroma-active (dilution factor FD ≥ 16).
The outcome: 27 key aroma compounds — mainly fruity and floral — dissolved in 53% hydroalcoholic solution at their natural concentrations, successfully simulated the aromas of those liquors. They also related the key compounds, seven sensory attributes and the samples through partial least-squares regression.
Twenty-seven out of eighty.
This is a different product class — liquor is not perfume, and the concentration scales and matrix differ — so the ratio does not transfer. But it demonstrates a principle: the number of components an analysis finds and the number needed to rebuild the smell are not the same number. The other 53 compounds were detected and identified, and then went unused in the recombination.
That is the same shape as what we measured on our own formula corpus: a median of 16 materials, of which 9 carry 90% of the weight.
The chemist's rebuttal
One reply, from someone describing themselves as a chemist and hobby perfumer, is rude in tone and worth reading technically:
"Oh, no! They'll have to run a liquid chromatography column before sending it through the GC/MS machine. Whatever shall they do? /s … It's either too reactive and will destroy the scent, or it can be separated by another means. I think somebody is trying to sell something and overstating its effectiveness."
"Run another separation first" is standard practice in analytical chemistry.
Tranchida and colleagues, in the Journal of Chromatography A in 2016, developed a comprehensive two-dimensional gas chromatography-quadrupole mass spectrometry (GC×GC-qMS) method for determining 54 fragrance allergens in cosmetics — the set recently highlighted by the EU Scientific Committee on Consumer Safety. Flow modulation conditions were tuned to roughly 7 mL/min for compatibility with the mass spectrometer, and six-point calibration curves from 1 to 100 mg/L gave satisfactory linearity for all 54. Quantification used extracted ions; identification used three things together — ion ratios (qualifier/quantifier), full-scan MS database matching, and linear retention indices.
Note those three identification criteria. Co-elution in one-dimensional GC — different substances emerging together with overlapping signals — is an old problem in this field, and two-dimensional separation with multiple identification criteria exists precisely to address it.
None of that means an interference technology has no effect. It means the effect belongs in a frame where the analytical side is also advancing. Another reply put it more plainly: "It is a bit of an arms race but I suppose you can only make a GC-MS analysis more difficult."
Other things in the thread worth keeping
This is not new. One user said they first heard about the technology in 2021, that it was released in 2022, and that "this is a 5 years old tech, yet today clones are sprouting like mushrooms and on a daily basis".
Masking has been done for a long time. The GC analyst added: "Things like ISO E Super, IBCH and Vertofix Coeur already were used to try and hide things within a formulation and it's not really stopped anyone trying to recreate fragrances so far."
And there is a completely different approach. Someone describing themselves as a food and flavour technologist wrote: "Givaudan does something similar with their sulphur compounds where they let them react in very specific condition so they end up getting a product that you cannot make from the single ingredients. Also most houses use some sort of fingerprinting in their formulation."
That approach differs from interfering with the reading: it does not make you read wrong, it makes the thing unbuyable even when you read right.
GC-O in practice. The same analyst described how they use the olfactometry port: the perfumery team uses it more than the analysts do, usually when a peak of decent concentration turns up (above 0.1%). The analysts note rough retention times, then the perfumers arrive about five minutes before the peak is expected. They mentioned one case where a material has multiple peaks and the smell is mostly in the first one, while the larger second peak was hard for the analytical team to detect at all.
Peak size and odour importance are not the same quantity. Which is exactly why a nose is still needed after the GC.
What this doesn't establish
I did not check the Cryptosym source data. The 133, 93 and 25 were relayed by the original poster from an article; I did not obtain the underlying test report or manufacturer documentation. This article analyses what those numbers would answer if true; it does not verify them.
My corpus is not commercial perfume. Those 954 are public demo formulas, 74% carrying a patent or source field. Commercial fine fragrance may differ in material count, weight distribution and material list — especially in proprietary bases that never appear in public documents. So "the top three carry 54.7%" describes these documents.
"Common materials cover 83% of the weight" is not "you can guess the formula". That figure measures weight coverage, not identification accuracy. Knowing a formula probably contains linalool is a long way from knowing it contains 5.93% of it and that its neighbour is benzyl acetate rather than phenethyl acetate. I measured how strong the prior is, not that the prior replaces GC.
The analyst's 70–90% is self-reported. I cannot verify their occupation or that figure, and practice may vary a great deal by company and product class.
Niu is about liquor. The 80-to-27 ratio comes from another product class, another concentration scale and another matrix. I used the principle — components detected is not components required — rather than the ratio.
Tranchida is about measuring allergens, not reverse engineering. I cite it to show that two-dimensional separation and multiple identification criteria are routine in the field, not to claim anyone uses that method to break formulas.
I do not know whether Cryptosym is actually deployed, or where. One user said clones are everywhere five years on, which is their observation, not a market survey.
References
Y. Niu et al., Characterization of the key aroma compounds in different light aroma type Chinese liquors by GC-olfactometry, GC-FPD, quantitative measurements, and aroma recombination, Food Chemistry, 233, 204–215 (2017). PMID 28530568. doi:10.1016/j.foodchem.2017.04.103
P. Q. Tranchida et al., Four-stage (low-)flow modulation comprehensive gas chromatography-quadrupole mass spectrometry for the determination of recently-highlighted cosmetic allergens, Journal of Chromatography A, 1439, 144–151 (2016). PMID 26718184. doi:10.1016/j.chroma.2015.12.002
Related: Notes are not materials, Look up how many formulas it appears in, What pairs with what, How to read a data sheet.