Friday, October 9, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Chemistry

AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us

October 9, 2026
in Chemistry
Bethany Barker
By Bethany Barker Scienmag Editorial Profile - Catalysis
Reading Time: 5 mins read
0
AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us

AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Predicting how a molecule will smell from its chemical structure alone remains one of the most stubborn unsolved problems in sensory science. Unlike vision, where wavelength maps neatly onto color, or hearing, where frequency determines pitch, the link between molecular architecture and perceived odor is wildly nonlinear. A molecule activates not one receptor but a combinatorial pattern across hundreds of human olfactory receptors, and tiny structural changes can transform a pleasant floral note into something foul, while structurally unrelated compounds can smell nearly identical. Now, a new study from researchers at the Osaka Research Institute of Industrial Science and Technology suggests an unexpected ally in cracking this problem: generative large language models, the same class of artificial intelligence that powers modern chatbots.

In research published in Discover Chemistry, Takuya Ehiro and Reiko Yamashita systematically evaluated GPT-family models as zero-shot generators of odor descriptors, meaning the models received no task-specific training whatsoever. Working with a curated dataset of 3,493 odorant molecules drawn from the Leffingwell PMP 2001 database and distributed via the Pyrfume repository, the team asked the models to describe odors, rate pleasantness on a scale from minus four to plus four, and even estimate their own reliability. The results reveal that these language systems, trained on vast human-generated text spanning chemistry, perfumery, and food science, have absorbed a remarkable amount of collective olfactory knowledge.

The most striking finding concerns concentration. In human perception, the same compound can smell entirely different depending on how much of it reaches the nose. Indole, for example, is floral at trace levels and fecal at high concentrations. When the researchers prompted the models with low- and high-concentration conditions for four such odorants, including indole, diacetyl, skatole, and cis-3-hexen-1-ol, the newer GPT-5 models reproduced these hedonic shifts with impressive fidelity. GPT-5-mini and GPT-5 assigned positive pleasantness scores to indole and skatole at low concentrations and collapsed to the bottom of the scale at high concentrations, exactly mirroring the documented perceptual transitions. The older GPT-4o-mini failed these tests, suggesting that advances in model capability translate directly into better encoding of concentration-percept relationships.

The team also probed whether language models could capture cultural dimensions of smell, something no structure-based approach can attempt. They assigned the models distinct cultural personas through role-play prompting, simulating sensory scientists from East Asia, North America, and Western Europe, then asked them to evaluate methional, a key aroma compound in soy sauce and kimchi, and methyl salicylate, the wintergreen compound beloved in American candy but associated with medicine in Europe. The models produced trends in the predicted directions, with East Asian personas rating methional more favorably, but the differences were not statistically significant. The authors frame this as a proof of concept rather than evidence of culturally grounded odor modeling, noting that it echoes cross-cultural research showing molecular identity matters far more than cultural background in odor pleasantness.

To benchmark descriptor quality, the researchers compared LLM-generated odor descriptions against expert-assigned ground-truth labels using two vectorization schemes: embedding-based similarity in a continuous semantic space and count-based matching against 180 canonical descriptor columns. A message-passing neural network pretrained on data including the evaluation set achieved the highest scores, with a mean cosine similarity of 0.814 under embedding-based vectorization, but this comparison was inherently asymmetric, since the neural network had seen the test data during training while the language models operated entirely zero-shot. All LLM conditions substantially exceeded a frequency-informed random baseline of 0.677, demonstrating genuine chemical-olfactory knowledge rather than simple label-frequency matching. Notably, under top-10 sampling, the neural network’s performance dropped to a level comparable to or below GPT-5-mini and GPT-5.

Perhaps the most revealing experiments involved scrambling the inputs. When the researchers paired each compound’s correct chemical name with a SMILES string from a different molecule, 76.3 percent of the generated descriptors were semantically closer to the name-source compound than to the SMILES-source compound, establishing that the chemical name is the primary carrier of odor-relevant information. Yet the models were not ignoring structure entirely: in the name-masked condition, they still produced chemically plausible descriptors from SMILES strings alone, and in 5.6 percent of scrambled cases they explicitly flagged the mismatch between name and structure, evidence of cross-referencing between the two representations. This aligns with earlier findings that IUPAC nomenclature, which systematically encodes functional groups and substitution patterns, outperforms other molecular string formats for language model prediction tasks.

Can a chatbot’s self-reported confidence be trusted? Partially, the study suggests. The models’ expressed reliability scores correlated positively with prediction accuracy across name-containing prompt conditions, with correlations reaching 0.45 against the minimum cosine similarity at low concentration, and negatively with prediction variability. The correlation vanished for SMILES-only prompts, indicating that confidence is only informative when the model has sufficient semantic context. Meanwhile, principal component analysis of the embedding space revealed that hedonic valence dominates the first principal component, accounting for over 42 percent of variance, meaning pleasantness is the single largest organizing axis of the models’ internal odor representations.

To verify that the models’ pleasantness judgments rest on genuine chemistry rather than linguistic coincidence, the team trained an XGBoost surrogate model to predict LLM-assigned pleasantness scores from molecular fingerprints alone. It achieved an R-squared of 0.810 at high concentration, and even under a demanding scaffold-based split that excluded all unseen molecular cores from training, performance remained at 0.754. SHAP analysis surfaced chemically interpretable patterns: acyclic ether and aromatic fragments contributed positively to predicted pleasantness, while nitrogen-containing and carbonyl substructures contributed negatively. An independent effect-size analysis confirmed the picture, with thiol groups showing the strongest negative association and ethers and esters the strongest positive ones, directions fully consistent with established olfactory literature.

The practical payoff came in predicting instrumental odor sensor measurements. Using a Shimadzu FF-2020 electronic odor identification system, the researchers measured similarity scores across nine gas categories for 44 compounds, then tested whether features derived from LLM outputs could improve prediction of these readings. For aliphatic hydrocarbons, LLM-derived features achieved an R-squared of 0.567 with XGBoost, compared with just 0.222 for conventional molecular fingerprints; for organic acids, the LLM features reached 0.429 against 0.189 for fingerprints. The study is candid about limitations, including the small sensor sample, cross-contamination in the nine-dimensional axis projection, and the models’ heavy dependence on chemical names, but the direction is clear.

What emerges is a portrait of language models as complementary players in computational olfaction rather than replacements for structure-based methods. They cannot yet beat a supervised graph neural network on its own benchmark, but they offer something no fingerprint ever could: sensitivity to concentration, hedonic framing, and potentially cultural context, all accessible through nothing more than prompt text. As generative models continue to improve, the boundary between reading about smell and predicting it grows thinner, and the perfume lab, the flavor house, and the environmental sensor industry may all find a new kind of colleague waiting inside their chat windows.

Subject of Research: Zero-shot evaluation of generative large language models for predicting odor quality and odor sensor measurements from molecular information

Article Title: Evaluating generative large language models as zero-shot semantic odor descriptor generators with applications to odor sensor prediction

Article References: Ehiro, T., & Yamashita, R. (2026). Evaluating generative large language models as zero-shot semantic odor descriptor generators with applications to odor sensor prediction. Discover Chemistry, 3(1), Article 570. https://doi.org/10.1007/s44371-026-01015-7

Image Credits: AI Generated

DOI: 10.1007/s44371-026-01015-7

Keywords: large language models, odor prediction, olfactory perception, zero-shot learning, GPT-5, chemical nomenclature, molecular fingerprints, pleasantness, odor sensors, machine learning, quantitative structure-odor relationship, sensory science

Cite Scienmag News

Bethany Barker. (October 9, 2026). AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us. Scienmag. https://scienmag.com/ai-language-models-learn-to-smell-gpt-systems-predict-how-molecules-will-scent-us/

Bethany Barker. "AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us." Scienmag, 9 October 2026, https://scienmag.com/ai-language-models-learn-to-smell-gpt-systems-predict-how-molecules-will-scent-us/. Accessed 9 October 2026.

Bethany Barker. "AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us." Scienmag. October 9, 2026. https://scienmag.com/ai-language-models-learn-to-smell-gpt-systems-predict-how-molecules-will-scent-us/

Tags: AI in sensory descriptor estimationAI-driven odor predictionchemical nomenclaturechemical structure to smell mappingcomputational approaches to olfactory sciencegenerative language models in sensory scienceGPT-5large language modelslarge language models in chemistryMachine learningmachine learning for scent predictionmolecular architecture and scent perceptionmolecular fingerprintsmolecular structure and odor similarityodor pleasantness rating predictionodor predictionodor sensorsolfactory perceptionolfactory receptor activation patternspleasantnessquantitative structure-odor relationshipsensory sciencezero-shot learningzero-shot odor descriptor generation
Share26Tweet16
Previous Post

Fasting Starves Tumors of Taurine, Triggering Immune Cell Death in Colorectal Cancer

Next Post

Chatbots Speak English Even When They Don’t: AI’s Hidden Western Values

Related Posts

New Look-Up Tables Sharpen the Sizing of Airborne Particles Measured by Optical Counters
Athmospheric

New Look-Up Tables Sharpen the Sizing of Airborne Particles Measured by Optical Counters

October 8, 2026
New tellurium-rich silver mineral fengruiite reveals how hot fluids trap silver in China’s Qinling ore belt
Chemistry

New tellurium-rich silver mineral fengruiite reveals how hot fluids trap silver in China’s Qinling ore belt

October 8, 2026
Sinking Air Shapes Arctic Clouds: New Study Tracks Cold Air Outbreaks Over the Fram Strait
Athmospheric

Sinking Air Shapes Arctic Clouds: New Study Tracks Cold Air Outbreaks Over the Fram Strait

October 8, 2026
Hidden Atomic Spiral in Uranium Crystal Reveals Rare Dual Magnetism
Chemistry

Hidden Atomic Spiral in Uranium Crystal Reveals Rare Dual Magnetism

October 8, 2026
NMR Pulse Trick Now Rotates Both Coupled Spins at Once
Chemistry

NMR Pulse Trick Now Rotates Both Coupled Spins at Once

October 8, 2026
Soot’s Light-Bending Secret: Why the Optical Model You Choose Changes the Answer
Athmospheric

Soot’s Light-Bending Secret: Why the Optical Model You Choose Changes the Answer

October 8, 2026
Next Post
Chatbots Speak English Even When They Don’t: AI’s Hidden Western Values

Chatbots Speak English Even When They Don't: AI's Hidden Western Values

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Rapid HIV Blood Bank Tests in Mozambique Miss a Concerning Share of Infected Donations
  • Chatbots Speak English Even When They Don’t: AI’s Hidden Western Values
  • AI Language Models Learn to Smell: GPT Systems Predict How Molecules Will Scent Us
  • Fasting Starves Tumors of Taurine, Triggering Immune Cell Death in Colorectal Cancer

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading