Beginner perfumer · 122
Beginner perfumer #122: values marked (est) agree with each other better, and that is exactly why their agreement proves nothing
· 17 min read
This database marks many numbers with est. The rates vary widely: xlogp3 is 100%, logP 99.7%, appearance 99.8%, vapour pressure 94%, boiling point 57.7%, and melting point only 4.2%. I used the boiling-point-to-vapour-pressure regression to test whether est values are less accurate. The answer runs against intuition. Restricted to fragrance materials with an odour field, the 518 records where both fields are estimates have a residual standard deviation of 18.7 °C with 1.0% exceeding 50 degrees, while the 186 where both are measured have 27.1 °C and 7.5%. The estimates are more consistent on every measure. The reason is not hard: two estimates come from one model and their agreement is by construction, while two measurements come from different laboratories and decades, where disagreement is normal. And that is the point. Consistency is not accuracy, so two estimates agreeing cannot be evidence that either is right. This also constrains a lot of the cross-field checking I have done myself.
About a 5 minute read.
The short version
- The
(est)rate forms a clear ladder:
| Field | Has value | Marked (est) |
Rate |
|---|---|---|---|
| xlogp3 | 18,450 | 18,450 | 100.0% |
| appearance | 10,270 | 10,248 | 99.8% |
| logp | 19,667 | 19,604 | 99.7% |
| vapor_pressure | 9,020 | 8,479 | 94.0% |
| solubility | 23,941 | 20,158 | 84.2% |
| bp | 13,009 | 7,511 | 57.7% |
| flash_point | 14,316 | 8,109 | 56.6% |
| mp | 3,080 | 130 | 4.2% |
- I tested whether estimates are less accurate. Restricted to fragrance materials, the answer runs the other way:
| Boiling point / vapour pressure | Records | Residual SD | Deviation >50 °C |
|---|---|---|---|
| Both measured | 186 | 27.1 °C | 7.5% |
| Both estimated | 518 | 18.7 °C | 1.0% |
- The reason is that two estimates come from one model, so their agreement is by construction
- And consistency is not accuracy, which constrains a lot of my own checking
Where this little marker appears
Last round I found that (est) resolves boiling point conflicts:
of the 208 materials with two boiling points at one pressure, nine in ten have one marked and one not,
and you keep the unmarked one.
That made me want to look at the marker itself. It appears on eight fields at very different rates.
Melting point is only 4.2% estimated; xlogp3 is 100%.
The ordering makes sense. Melting points are easy to measure: put a fragment on a hot stage and watch for the temperature at which it goes clear. The instrument is cheap, the procedure is quick, the result is unambiguous. Partition coefficients are hard to measure: you need a two-phase distribution and concentrations on both sides, and for materials that barely dissolve in water the measurement does not work at all. So people compute them.
A field's estimate rate reflects how inconvenient it is to measure.
I expected estimates to be worse
That is testable. Using the boiling-point-to-vapour-pressure regression I have leaned on for
several rounds — 8,048 materials,
boiling point ≈ 181.5 − 36.5 × log10(vapour pressure) — I compared residuals by combination.
Across the whole database:
| Combination | Records | Median residual | SD | >50 °C | >100 °C |
|---|---|---|---|---|---|
| Both measured | 328 | 13.2° | 24.8 | 5.8% | 0.9% |
| Boiling point measured, vapour pressure estimated | 3,791 | 10.8° | 26.6 | 4.0% | 1.3% |
| Both estimated | 3,913 | 10.5° | 40.4 | 2.7% | 1.9% |
The medians are nearly identical and the standard deviations are not. Estimates do as well as measurements in the middle and have much fatter tails.
What is in those tails? I pulled the 74 records with residuals over 100 °C: 26 are herbicides, 13 are pharmaceuticals, 10 are natural extracts, and 96% have an empty odour field.
The estimator breaks in chemical space it never saw, and that space is most of this database.
Restricted to fragrance materials, it inverts
Narrowing to the 2,774 records with an odour field:
| Combination | Records | Median residual | SD | >50 °C | >100 °C |
|---|---|---|---|---|---|
| Both measured | 186 | 10.9° | 27.1 | 7.5% | 1.6% |
| Boiling point measured, vapour pressure estimated | 2,066 | 10.5° | 26.3 | 4.9% | 1.3% |
| Both estimated | 518 | 10.2° | 18.7 | 1.0% | 0.6% |
The estimates win on every measure. Deviations over 50 degrees run 1.0% against 7.5%, a factor of seven.
Why
Not because estimates are accurate, but because they come from one model.
That model computes a boiling point from molecular structure and computes a vapour pressure from the same structure. Both outputs use the same parameters and the same assumptions. Their agreement is designed in, not demonstrated.
Two measurements, meanwhile, come from different laboratories, decades, purities and instruments. Their disagreeing by 50 degrees means at least one is wrong, and that is information.
Estimates agreeing tells you nothing. Measurements disagreeing tells you there is something to check.
This constrains a lot of my own articles
This round's conclusion sends me back over my own checks.
My method for dozens of rounds has been to check two fields of one record against each other: boiling point against melting point, boiling point against vapour pressure, optical rotation against deterpenation state.
When both of those fields carry (est), their agreement carries no evidential weight.
Concretely:
- I wrote that Verdox and Vertenex have "identical" vapour pressures.
Both are
(est). One model reading two near-identical structures will of course return the same answer. I noted this in that article's closing section without quantifying how badly it bites. - I argued from vapour pressure that 2-methyl undecanal's boiling point should be 233. That argument rests mainly on the homologous series, with vapour pressure as corroboration, so it still holds.
- Conversely, melting points are only 4.2% estimated, which makes the "boiling point field equals melting point field" check unusually trustworthy: the melting point side is nearly all measured.
How to use the marker
- Do not read
(est)as "unreliable". Inside perfumery's chemical space it performs well - But do not use two
(est)values agreeing as evidence. They are two outputs of one model - Two measurements disagreeing is what deserves a look. It means something is wrong
- The melting point field is the most measured field in this database (95.8% unmarked), which makes it the safest baseline
- xlogp3 contains no measured values at all. All 18,450 carry
(est): a purely computed column
What this doesn't establish
- I compared against no external ground truth. What I measured is how well two fields agree, which is consistency rather than accuracy. Estimates can be consistently wrong.
- The regression's own residual standard deviation is 34 °C across the database. It is not a precision instrument.
- "One model" is my inference. The database does not say how
(est)values were computed; I only know they behave as though they share a source. - The measured group's higher variance may partly come from my parsing. Where a field holds
several values I take the first
@ 760.00 mm Hgsegment, which may be picking the worse one. - The fragrance / non-fragrance split uses "does the odour field have a value", the rough criterion I set last round.
- How measurement error affects statistical inference has its own literature (regression dilution, PMID 20573762, is one case), as does chemical dataset curation (PMID 27885862). What I ran here is a comparison inside one dataset, not work at that level.