Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

Can AI Discover New Drugs? A High-Level Overview

1Why Drug Discovery Is Hard, and Where AI Fits2How AI Learns From Molecules and Proteins3Finding and Validating a Biological Target4Designing Molecules: Generative AI and Virtual Screening5From Hit to Lead: Optimizing Properties With AI6What AI Still Cannot Do7Judging the Claims: Real Successes, Failures, and Open Questions
How AI Learns From Molecules and Proteins

Turning a Molecule Into Something a Model Can Read

1 / 4
Look at the left panel first. Every labeled circle is an atom, and every line between them is a bond. That is the whole chemistry — which atom touches which, and how. Now look right. The same molecule is written as a short string of letters. The letters are element symbols, and the order tells you how the atoms are chained together. Nothing was added or removed; only the format changed. Notice what is missing from both panels: the actual three-dimensional shape. Neither representation knows how the molecule folds in space, and that gap will matter later when we ask whether a molecule fits into a protein pocket.
0:00 / 0:00

A molecule is a set of atoms connected by bonds, and that connectivity is what determines its chemistry. To feed a molecule to a model, we keep the connectivity but change the format. Two representations dominate.

The first is a graph: atoms become nodes, bonds become edges. Each node carries a small list of features — element type, charge, whether it sits in a ring — and each edge carries bond order and type. A graph is a natural fit because most molecular properties depend on which atoms are near which, not on any absolute position in space.

The second is a line notation, most commonly SMILES. Here the same structure is written as a string of characters, for example \(\text{CCO}\) for ethanol: a carbon, bonded to another carbon, bonded to an oxygen. Branching uses parentheses and rings use paired digits, so a fairly complex molecule can be written as one line of text.

Both forms are lossy in the same way: they describe connectivity but not the molecule's actual three-dimensional shape. That is usually acceptable for property prediction, and it is a real limitation for anything that depends on shape.

References

  1. [1]Simplified molecular-input line-entry system (SMILES) — Wikipediaen.wikipedia.org
Previous1 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion