
What a Perfume Match Score Actually Measures
What a perfume match score actually measures: the four data layers behind the percentage, why shared notes mislead, and where the score is honestly unreliable.
The number is the last thing our engine produces, and it is the part people read first. That inversion causes most of the confusion around match scores, so this is a walk through what sits behind the percentage, in the order the engine actually builds it.
I wrote the matching algorithm, so the description below is of our system specifically rather than of AI recommendation in general. The details that follow are the ones visible in the app.

What the engine reads before it scores anything
A fragrance in the catalogue is not a name and a note list. It is four separate layers, and the match is computed against all four.
Notes. The composition broken into individually indexed atomic notes, positioned as top, heart, base, or distributed across the wear.
Accords. The perceptual shape of the fragrance, with a strength value on each character. This is the layer people underestimate.
Vibe. Gender lean, age range, occasion, and season, each recorded as a range rather than a single label.
Performance. Longevity and sillage as measured behaviour, not marketing claims. Montale Roses Musk, for instance, records 6 hours or more of longevity at medium sillage, and reads across spring, summer, and autumn.

On your side of the comparison there is a personal scent map, built from the profile quiz, from every fragrance you have explicitly liked or disliked, and optionally from a photo you upload. It is redrawn after each of those interactions, so the field you get in month three is not the field you got on day one. The score is the comparison of those two structures. Nothing more mystical than that is happening.
Why two perfumes that share a note smell nothing alike
This is the single most common objection we get, usually phrased as "but they both have oud, so why is one a 91 and the other a 52?"
Open the cinnamon note in the app and you get 111 fragrances. The four at the top of that list are Lattafa Khamrah, By Kilian Angels' Share, Xerjoff XJ 1861 Naxos, and Emporio Armani Stronger With You Intensely. Four houses, four price tiers, four completely different compositions, and all four are correctly tagged with the same note.

A note tells you an ingredient is present. It tells you nothing about how much of it there is, where it sits in the wear, or what it is standing next to. Cinnamon at two percent behind vanilla and tonka reads as warmth. Cinnamon in front of a dry woody base reads as heat and spice. Same note, opposite experience.
That is the whole reason the accord radar exists. Here is Penhaligon's The Omniscient Mr Thompson: sweet at 100 percent, vanilla at 82, woody at 77, powdery at 59, fresh spicy and floral both at 45, drawn from the 14 accords detected in it.

Those percentages are the part that carries the smell. A note list is a set of ingredients. The radar is the proportions, and proportions are what your nose actually receives. When the engine tells you a fragrance from a family you have never explored is a strong match, it is almost always because the radars line up even though the note lists do not overlap at all. Family membership in our data is weighted rather than exclusive, precisely so those crossings can be found.
What "82.2 percent" is a percentage of
It is the compatibility between one fragrance profile and one taste vector. It is not a rating of the perfume, not a probability that you will buy it, and not a claim about quality.
The consequence matters when you shop. A 62 percent match can be an extraordinary fragrance that is simply far from your centre. A 91 percent match can be a competent, unexciting bottle that happens to sit exactly on it. The score answers "how close is this to what you have told me you like", which is a narrower and more useful question than "is this good".
This is also why a session returns 40 to 50 ranked matches rather than one answer. A single recommendation implies a confidence the data does not support. A ranked field lets you see the shape of your own taste, including the places where two very different fragrances score the same.
Where the score is honestly unreliable
Three failure modes, none of which we can engineer away.
Skin chemistry. pH, moisture, and body heat change projection and dry-down. No dataset contains your skin. This is the hard limit on every fragrance recommendation system, ours included.
Reformulation. A fragrance keeps its name across reformulations that materially change it. When a house thins a top note or substitutes a restricted material, the bottle on the shelf drifts from the record. We correct these as we find them, but there is always a lag.
Application. Two sprays and six sprays of the same fragrance are different products in a room. Sillage data describes the fragrance, not your habits.
Because of those three, the written reason attached to each match matters more than the number. If a match says it chose a fragrance for a leather and tobacco base, and you dislike leather, the reason has just told you more than the percentage did. The reason is falsifiable. The number is not.
How to read your own result
Read in this order.
- The reason first. Check whether the notes and accords it names are ones you actually want. If they are not, the profile behind your score needs correcting more than the fragrance does.
- The radar second. Look at the two or three tallest axes. If the fragrance is 100 percent sweet and you do not wear sweet, no percentage should move you.
- The vibe and performance third. Occasion and season are where a technically correct match becomes a practical mistake. A high-scoring fragrance with heavy sillage is still wrong for an open-plan office.
- The number last. Use it to rank, not to decide.
Then go and wear it. The engine's job is to reduce a whole catalogue to a shortlist worth a weekend of your skin. It is not to replace the skin. There is a full walkthrough of that sampling process in How to Find Your Signature Scent in 7 Honest Steps.
What we do when the score is wrong
It is wrong regularly, and we treat that as data rather than noise.
Three signals correct the graph. Internal consistency, where structurally similar fragrances must receive consistent note, accord, and family assignments, and disagreements get inspected by hand. Perceptual ground truth, where algorithm output is checked against blind sniff panels and published reviews from recognised noses. And behaviour in the app, where a match that consistently scores high and consistently fails to convert across many users is a signal that the weighting is wrong, not that the users are.
That third one is the uncomfortable one, because it means the fragrances we are most confident about are the ones most likely to expose an error. It is also the only signal in the set that scales.
The full data model, including the eight families and the sources behind them, is documented on the methodology page. If you want to understand the note structure underneath all of this first, start with Top, Middle, Base: How the Notes Pyramid Actually Works, then The Fragrance Families, Decoded.
Frequently asked
- Is a higher match score a better perfume?
- No. The score measures compatibility between a fragrance profile and your profile. A 90 percent match is a fragrance whose structure lines up closely with the taste you have expressed. A 60 percent match may be a more celebrated perfume, better made and better reviewed, that simply sits further from your centre. The score is a distance, not a verdict.
- Why do two perfumes with the same notes get different scores?
- Because notes are a small part of the record. The engine also reads the accord radar, which is the perceptual shape of the fragrance and carries how loud each character is rather than just whether it is present, plus the vibe layer and measured longevity and sillage. Two fragrances can share a note list and still land far apart once the proportions are read.
- What is the accord radar?
- A per-fragrance fingerprint of its dominant characters with a strength value on each. Penhaligon's The Omniscient Mr Thompson, for example, reads sweet at 100 percent, vanilla at 82, woody at 77, powdery at 59, and fresh spicy and floral both at 45. It is a perceptual profile, not a chemical inventory, and it is what makes cross-family recommendations possible.
- Does the score account for my skin?
- It cannot, and no engine can. Skin pH, moisture, and body heat move projection and dry-down in ways that are not in any dataset. This is why the app returns a ranked shortlist with a written reason attached rather than picking a winner, and why the last step of finding a fragrance is always a full day of wear on your own skin.
Explore Fragnatique


