Fifty design heuristics that hold up across large compound sets — each with the number it actually claims, the study it came from, and the caveat that usually gets dropped in retelling.
——Compiled Sep 2026
How to read these
Every rule here is a population-level trend, usually a shift in the odds across 102–105 compounds. None of them predicts your next analogue. A series that sits outside the envelope and still works is a real series, not a violation — the rules describe where the base rate is favourable, not where chemistry is permitted.
Where the original paper is softer than the slogan it spawned, that gap is marked. Numbers are quoted from the primary source; two entries carry figures that could only be confirmed through citing reviews, and those say so.
Lipophilicity is the master variable
Every consequence of raising logP is a cost except one — and that one is the property measured first and most often. The diagram is the mechanism behind “molecular obesity”: the incentive arrow points one way, all five cost arrows point the other. Sources for each number are in the corresponding rule below.
The property envelope
Filters for where to start, not gates for what to keep. Each was fitted to a particular assay and a particular company’s chemistry.
Rule of Five MW · logP · HBD · HBA
Poor absorption or permeation becomes more likely once two or more of four limits are breached.
HBD > 5, HBA > 10, MW > 500, ClogP > 5 (or MlogP > 4.15). The rule was derived for passive oral absorption and explicitly excludes substrates of biological transporters — along with antibiotics, antifungals, vitamins and cardiac glycosides.
NoteIt predicts liability, not failure, and says nothing about potency, toxicity or metabolic stability.
Screen below the drug-like envelope, because optimisation reliably drifts upward in both size and lipophilicity.
Original cut-offs: MW < 350, ClogP < 3.5. Oprea later relaxed them to MW < 450 / ClogP < 4.5 as “too restrictive.” The underlying observation is that affinity enhancement is nearly always bought with mass and lipophilicity, so a drug-like starting point leaves no headroom.
Above MW 400 and logP 4, essentially every ADMET endpoint gets worse at once.
Solubility, permeability, bioavailability, volume of distribution, plasma protein binding, CNS penetration, brain tissue binding, P-gp efflux, hERG inhibition and CYP 1A2/2C9/2C19/2D6/3A4 inhibition all deteriorate with rising MW, logP, or both. Ionisation state is the third variable and is usually dropped from the shorthand — acids and bases behave oppositely for protein binding versus hERG.
Note“4/400” is community shorthand. The paper presents binned property trends; it never names a rule.
The window of acceptable lipophilicity narrows as molecular weight rises — permeability and clearance pull in opposite directions.
Plot MW on y against logD7.4 on x: the tolerable region is a triangle with its apex near MW 450. At the centre (logD 1.5, MW 350), 29% of compounds passed both the permeability and the metabolic-stability criterion; at MW 450 with logD 0 or 3, ≤4% did.
Check the figureSecondary sources disagree on the width of the base (logD −2 to 5 versus 0 to 3). The MW 450 apex is consistent. Some tools implement it as a crude bounding box, which is not the paper’s triangle.
Percentages are the share of compounds passing both the permeability and the metabolic-stability criterion. The apex at MW 450 is consistent across sources; the width of the base is not — secondary accounts give logD −2 to 5 and 0 to 3. Check Figure 3 of the paper before quoting vertices. Shaded region: where both criteria can be met — both MW-450 markers fall outside it.
Beyond Rule of Five MW > 500
Orally bioavailable drugs exist well outside Ro5 space, but they get there by specific mechanisms rather than by luck.
Macrocyclisation, N-methylation, intramolecular hydrogen bonding, and high dose or enabling formulation are what make bRo5 oral exposure achievable. Natural products and structure-based design from peptidic leads are the two main sources of this chemotype.
No fixed envelopeCiting reviews attribute several conflicting numeric ranges to this paper (MW ≤ 700 vs ≤ 1000, and so on). Treat bRo5 as a mechanism argument, not a new set of cut-offs.
Over two decades of approvals, the ClogP and HBD ceilings held while the MW and HBA ceilings moved substantially.
A property whose apparent limit tracks the calendar is a property of medicinal-chemistry practice, not of drugs. Shultz argues this “calls into question the entire hypothesis that ‘drug-like’ properties exist” — and at minimum means MW deserves less weight than lipophilicity and HBD count.
Two independent terms set aqueous solubility: how much the molecule likes octanol, and how much it likes its own crystal. Most design effort attacks only the first.
General Solubility Equation
Solubility is predicted from just two measured quantities: melting point and logP.
log S = 0.5 − 0.01 (MP − 25) − log Kow
S is intrinsic solubility in mol/L, MP in °C; for liquids (MP < 25 °C) the melting-point term goes to zero. Validated on 580 pharmaceutically, environmentally and industrially relevant nonelectrolytes.
Iso-solubility lines from the General Solubility Equation
Drawn directly from log S = 0.5 − 0.01(MP − 25) − log Kow; nothing here is fitted. The lines are parallel because melting point and logP enter the equation with fixed coefficients, which is the whole argument for treating crystal packing as a design variable: 100 °C of melting point costs exactly as much solubility as one logP unit. Shaded region: solubility better than 100 µM. Every line has slope −100 °C per logP unit — the two terms trade one for one.
↑ logP ↓ solubility
One unit of logP costs one log unit of aqueous solubility.
This is not an empirical correlation but the coefficient of −1.0 on log Kow in the GSE. The practical consequence: in a 711-compound set, roughly 50% of compounds with logP < 3 were soluble, against about 1% of those with logP > 3.
Jain & Yalkowsky, J. Pharm. Sci.2001 (as above). Empirical figures: Waring. Expert Opin. Drug Discov.2010, 5, 235–248. 10.1517/17460441003605098
melting point 100 °C = 1 log S
Crystal packing costs as much solubility as lipophilicity does, and nobody optimises against it.
The GSE’s −0.01 coefficient on melting point means each 100 °C of MP removes one log unit of solubility. Flat, rigid, symmetric molecules pack well and melt high — so planarisation is charged twice: once through logP, once through the lattice. A compound melting at 150 °C needs logP < 3.25 to reach 100 µM.
Jain & Yalkowsky, J. Pharm. Sci.2001; worked example from Waring, Expert Opin. Drug Discov.2010, 5, 235–248.
↑ Fsp³ escape from flatland
Raising the fraction of sp³ carbons raises solubility, lowers melting point, and tracks clinical progression.
Mean Fsp³ rises from 0.36 in discovery compounds to 0.47 in marketed drugs — a 31% increase, with a steady climb through Phase I → II → III. Chirality moves with it: compounds carrying at least one stereocentre go from 46% of discovery compounds to 61% of drugs. Lovering frames solubility as the mechanism behind the survival advantage.
Lovering, Bikker & Humblet. J. Med. Chem.2009, 52, 6752–6756. 10.1021/jm901241e
Saturation rises with clinical stage
Two independent measures of three-dimensionality, each higher in marketed drugs than in the discovery compounds they came from. Lovering reports a steady climb through Phase I→II→III rather than a jump at approval, which is what makes the trend hard to explain as survivorship alone.
aromatic rings ceiling of 3
More than three aromatic rings correlates with poorer developability across five independent parameters.
Solubility, lipophilicity, serum albumin binding, CYP450 inhibition and hERG inhibition all worsen past three aromatic rings. Critically, the solubility penalty persists within a fixed lipophilicity range — aromatic ring count is not merely a logP proxy.
Not all rings are equally expensive: carboaromatics are the problem, heteroaliphatics are an asset.
Rank order of harm is carboaromatic ≫ heteroaromatic > carboaliphatic > heteroaliphatic. For solubility, Spearman R is 0.39 (carboaromatic, detrimental), 0.10 (heteroaromatic), ≈0.01 (carboaliphatic, inert) and 0.16 in the beneficial direction for heteroaliphatic. Fused aromatic systems fare better than non-fused ones. Swapping a heteroaromatic for a carboaromatic buys lower solubility, higher protein binding and more CYP inhibition.
All ten correlations are in the detrimental direction, but the carbocycle is worse on every endpoint — nearly four-fold for solubility and nine-fold for hERG. Two findings do not fit on this axis: heteroaliphatic rings improve solubility (R = 0.16 beneficial), and CYP1A2 inhibition falls as carboaromatic count rises (R = −0.07).
Solubility Forecast Index
A two-term index combining lipophilicity and aromatic count forecasts solubility better than either alone.
SFI = ClogD7.4 + (number of aromatic rings)
Derived from roughly 100,000 kinetic solubility measurements. Low SFI is the target; the index makes explicit that an aromatic ring costs about as much solubility as a full logD unit.
Permeability is a desolvation problem. What matters is not how many polar atoms a molecule has but how much energy it costs to strip the water off them.
logD floor scales with MW
The minimum lipophilicity needed for permeability rises steeply with molecular weight.
For a >50% chance of Caco-2 A→B Papp > 100 nm/s: logD > 1.7 at MW 350–400; > 3.1 at 400–450; > 3.4 at 450–500; > 4.5 above MW 500. This is the quantitative core of the Golden Triangle — a big molecule must be greasy to cross, and greasy big molecules are cleared fast.
The lipophilicity floor rises faster than molecular weight
Minimum logD for a better-than-even chance of acceptable passive permeability. Read it alongside the clearance and hERG rules: by MW 450–500 the logD needed for permeability has already passed the ceiling a basic compound can afford for hERG, which is what closes the Golden Triangle at its apex. Shaded: below the floor, fewer than half of compounds reach Caco-2 A→B Papp > 100 nm/s. The jump from the 350–400 band to 400–450 is 1.4 logD units — the steepest step on the chart.
optimal window logD 1–3
The space where solubility, permeability and clearance are simultaneously acceptable is narrow.
Waring puts it at logD between roughly 1 and 3. Below it, permeability fails; above it, solubility and metabolic stability fail. Most optimisation campaigns drift upward through this window rather than settling in it.
Hydrogen-bond donors are the expensive polarity; acceptors are comparatively cheap.
Hitchcock & Pennington’s envelope for brain exposure is MW < 500, PSA < 90 Ų, ClogP 2–5, ClogD7.42–5, HBD < 3. Capping or masking a single NH is often the cheapest permeability fix available.
AttributionThe widely quoted TPSA < 60–70 Ų CNS figure is Pajouhesh & Lenz’s, not this paper’s. Hitchcock & Pennington use PSA < 90 Ų.
Hitchcock & Pennington. J. Med. Chem.2006, 49, 7559–7583. 10.1021/jm060642i
intramolecular H-bonds
An internal hydrogen bond can hide polarity from the membrane — but it is not a free win.
Kuhn mined the CSD and PDB for the propensity of five- to eight-membered pseudo-rings to form IMHBs, then measured matched sets. The result is explicitly conditional: the effect on permeability, solubility and lipophilicity “depend[s] on a subtle balance between the strength of the hydrogen bond interaction, geometry of the newly formed ring system, and the relative energies of the open and closed conformations in polar and unpolar environments.”
Not a ruleDesign the pseudo-ring geometry deliberately; an IMHB installed by accident can cost solubility without buying permeability.
Kuhn, Mohr & Stahl. J. Med. Chem.2010, 53, 2601–2611. 10.1021/jm100087s
chameleonicity above MW 600
Past MW 600 the solubility and permeability requirements become mutually exclusive for a rigid molecule.
Good solubility needs TPSA ≥ 0.2 × MW; good passive permeability needs 3D PSA in a nonpolar medium ≤ 140 Ų. Above roughly MW 600 one of these is almost always violated — so a large oral molecule must be a conformational chameleon, presenting polarity in water and burying it in membrane.
Accumulation in E. coli requires an ionisable nitrogen, low three-dimensionality and rigidity.
From >180 diverse compounds assayed for intracellular accumulation: a non-sterically-encumbered ionisable Nitrogen (a primary amine works best), low Three-dimensionality (globularity ≤ 0.25), and Rigidity (≤ 5 rotatable bonds). Applying the rules converted deoxynybomycin, a Gram-positive-only agent, into a broad-spectrum antibiotic.
NoteThe “eNTRy” acronym was coined in follow-up work, not in the 2017 paper.
Lipophilicity is the single property that drives clearance, CYP inhibition and protein binding in the same direction at once.
↑ logP ↑ clearance
Human clearance rises roughly four-fold across the practical lipophilicity range.
Median human CL is 2.2 mL/min/kg for compounds with logP < 0, against 7.8 mL/min/kg for logP > 4. CYP active sites are hydrophobic; partitioning into them is the rate-limiting recognition step.
Increasing saturation and stereochemical complexity reduces both promiscuity and CYP450 inhibition.
The follow-up to Escape from Flatland tested the selectivity half of the hypothesis against cross-screening panel data and CYP450 inhibition data. Higher Fsp³ and higher chiral-carbon count both track with fewer off-target hits — a three-dimensional molecule fits fewer pockets, including the promiscuous ones.
Carbocyclic aromatic rings drive CYP inhibition roughly twice as hard as heteroaromatic ones — with one reversal.
Spearman R for carboaromatic ring count, all detrimental: CYP3A4 0.13, 2C9 0.26, 2C19 0.19. Heteroaromatic count is weaker: 3A4 0.08, 2C9 0.12, and no effect on 2C19. The exception worth knowing: CYP1A2 inhibition falls as carboaromatic count rises (R = −0.07).
Fluorine at a site of oxidation blocks that metabolic pathway without adding steric bulk.
The C–F bond resists CYP-mediated hydroxylation, and fluorine’s van der Waals radius is close enough to hydrogen’s to be tolerated in most pockets. Beyond metabolic blocking, fluorination productively influences conformation, pKa, intrinsic potency and membrane permeability — and supplies the 18F handle for PET.
Trade-offBlocking one soft spot commonly reveals the next one. Metabolite ID before fluorinating, not after.
Efficacy tracks free drug concentration at the target, not free fraction — so optimising against plasma protein binding is wasted effort.
Designing to lower PPB “yield[s] no enhancement of the in vivo free drug concentration” and “could result in the wrong compounds being advanced.” A 99%-bound compound at high total exposure can deliver more free drug than a 90%-bound one at low exposure. Compare unbound concentrations against unbound potency, and stop reporting %PPB as a liability.
Smith, Di & Kerns. Nat. Rev. Drug Discov.2010, 9, 929–939. 10.1038/nrd3287
Potency & efficiency
Affinity per atom, not affinity. Every efficiency metric exists to catch the same failure mode: potency bought with mass and grease.
Kuntz ceiling 1.5 kcal/atom
Binding free energy rises about 1.5 kcal/mol per heavy atom for the first fifteen atoms, then essentially stops.
Across a large survey of ligand–receptor pairs, ΔGbinding has an initial slope near −1.5 kcal/mol per non-hydrogen atom; past 15 heavy atoms it “increases very little with relative molecular mass,” and the tightest binders plateau around −15 kcal/mol — picomolar. Everything you add after that first fifteen atoms is, on the average, paying for something other than affinity.
The maximal-affinity ceiling, and what it does to ligand efficiency
The envelope flattens; the LE = 0.3 line does not. That divergence is the whole reason ligand efficiency falls during optimisation even when potency improves — atoms keep being added on a straight line while the affinity available to buy with them runs out. Kuntz’s stated values are the initial slope, the 15-atom knee and the −15 kcal/mol plateau; the connecting curve is illustrative. Envelope drawn as ∆G = −15(1 − e−0.1N), the simplest curve reproducing all three quantities the paper states; the dashed tangent shows its initial slope.
Ligand efficiency LE ≥ 0.3
Normalise potency by heavy-atom count, and hold the line through optimisation.
LE = −ΔG / Nheavy ≈ 1.37 × pIC50 / HAC
The floor is derived rather than asserted: a 10 nM compound at a 500 Da ceiling (about 38 heavy atoms) requires LE ≥ 0.29 kcal·mol−1·HA−1, which is where the familiar “0.3” comes from.
AttributionThe 1.37 form is a later restatement (−2.303RT at 300 K, 1 M standard state). Hopkins et al. wrote LE = −ΔG/N.
Subtract lipophilicity from potency; what remains is the potency you actually earned.
LLE = pIC50 (or pKi) − cLogP
An optimised candidate should reach 5–7 or greater; LLE = 6 means a millionfold preference for the target over 1-octanol. Because raising logP raises potency in almost any assay, LLE is the metric that distinguishes a better molecule from a greasier one.
The median patented compound is a full logP unit greasier than the median oral drug.
Median patented compound: cLogP 4.1, MW 450. Median oral drug discovered since 1990: cLogP 3.1, MW 432. The mass difference is small; the lipophilicity difference is not. This single comparison is the empirical case for tracking LLE.
Leeson & Springthorpe. Nat. Rev. Drug Discov.2007, 6, 881–890. 10.1038/nrd2445
BEI · SEI size & polarity
Plot binding efficiency against surface efficiency to see whether potency is coming from bulk or from polar contacts.
BEI = pKi / MW(kDa) · SEI = pKi / (PSA / 100 Ų)
The BEI/SEI plane separates series that gain potency by adding mass from those gaining it through polar interactions, which LE alone cannot distinguish.
A single C–H → C–CH₃ substitution can improve potency more than a hundredfold — occasionally.
Methyl appears in more than 67% of the top-selling drugs of 2011, and the review catalogues cases where one methyl delivers >100-fold improvement, typically by enforcing a bound conformation or filling a small hydrophobic subpocket.
Base rateProfound methyl effects are the tail, not the distribution. Most methyl additions do very little; treat >100-fold as a possibility worth testing, never an expectation.
Optimise the dissociative half-life of the complex, not only its equilibrium constant.
Two compounds with identical Kd can behave completely differently in vivo if their koff differs. Long residence time prolongs pharmacodynamic effect beyond the pharmacokinetic exposure window and confers kinetic selectivity — selectivity that an equilibrium assay simply cannot see.
The same two properties — lipophilicity and basicity — account for most of the off-target liability that kills programmes.
Pfizer 3/75 ClogP · TPSA
ClogP above 3 combined with TPSA below 75 Ų raises the odds of in vivo toxicity roughly six-fold.
From animal in vivo toleration studies on 245 preclinical Pfizer compounds. Compounds with ClogP > 3 and TPSA ≤ 75 Ų were about 2.5× more likely to show a finding; those with ClogP ≤ 3 and TPSA > 75 were about 2.5× more likely not to — an overall odds ratio above 6. The trend held across toxicity types and chemotypes.
ContestedMuthas et al. analysed 150 preclinical and Phase I candidates and found the reverse — ClogP < 3 with TPSA > 75 more often associated with toxicity. 3/75 is a Pfizer-portfolio observation, not a law.
The six-fold figure everyone quotes is the ratio between the two diagonal corners, not the effect of either property alone. Note what the diagram does not say: the two off-diagonal quadrants are unremarkable, so neither lipophilicity nor polarity is doing this by itself. Muthas et al. found the opposite pattern in a different 150-compound set — see the caveat in the rule below. Based on animal in vivo toleration studies of 245 preclinical Pfizer compounds.
hERG basicity costs 2 logD
Being basic costs roughly two log units of allowed lipophilicity before hERG becomes a problem.
For a >70% chance of hERG IC50 > 10 µM: neutral compounds need logD < 3.3; basic compounds need logD < 1.4. At logD 3, about 76% of neutrals clear the bar against only 24% of bases; at logD 4, 55% versus 7%. The paper’s own conclusion is that “lipophilicity is a stronger driver for hERG potency than might have been expected.”
Check the PDFThe four probability figures come from a citing review quoting this paper; the abstract states only that target lipophilicity values are derived.
Two logistic curves fitted to the four quoted probabilities; both reproduce the stated thresholds to within rounding. The practical reading is the shaded band: a basic amine moves the lipophilicity ceiling down by roughly two logD units, which is usually the difference between a permeable compound and an impermeable one. Flagged in the rule below — these four figures come from a citing review, not the primary text. The dashed line marks a 70% chance of hERG IC50 > 10 µM; markers are the values quoted from the paper, curves are logistic fits through them.
hERG tactics what actually works
Five structural moves reliably attenuate hERG binding.
Lower the basic amine’s pKa; lower logP; introduce an acidic or zwitterionic group; add steric bulk adjacent to the basic centre; reduce aromatic ring count. hERG’s inner cavity is large and lined with aromatic residues, which is why lipophilic bases with a protonated nitrogen for cation–π contact are the canonical binders.
QualitativeThis is a case-based survey. Don’t attribute a per-logP-unit coefficient to it — use Waring & Johnstone for numbers.
Carboaromatic rings raise hERG risk; heteroaromatic rings do not.
Spearman R = 0.18 for carboaromatic ring count (detrimental) against 0.02 for heteroaromatic — effectively no effect. Heteroaliphatic count gives a slight increase (R = 0.08), driven by the basic amines those rings usually carry. Swapping phenyl for pyridyl is one of the cheapest hERG mitigations available.
A small set of substructures accounts for a large share of apparent screening hits.
From 93,212 compounds across six AlphaScreen campaigns at 10–30 µM, 480 substructures were required to filter out all assay-interference compounds. Family A (16 filters) alone removed 4,703 compounds; Family B (55) removed 2,196; Family C (409) removed 1,186.
Use as a promptThe filters were derived from one assay technology. A PAINS flag means “prove this one is real,” not “discard” — several approved drugs match PAINS substructures.
Baell & Holloway. J. Med. Chem.2010, 53, 2719–2740. 10.1021/jm901137j
colloidal aggregation
Most promiscuous inhibitors from HTS and docking are not inhibitors at all — they are colloids that sequester enzyme.
35 of 45 screening hits inhibited several unrelated model enzymes. Inhibition was reversible, attenuated by albumin, guanidinium or urea, and a 10-fold rise in enzyme concentration largely abolished it despite a 1000-fold excess of compound. Light scattering and EM showed particles 30–400 nm across.
Counter-screen:0.01% Triton X-100 restores activity if the compound is an aggregator. Four tells: detergent sensitivity; Hill slopes > 1.5 (often ≥ 5); IC50 that rises roughly linearly with enzyme concentration; and potency that increases on pre-incubation with protein, sometimes >50-fold.
Certain functional groups are bioactivated to reactive electrophiles, but the alert is a trigger for study, not a veto.
The comprehensive listing covers bioactivation pathways for anilines, thiophenes, furans, nitroaromatics, hydrazines, alkynes, quinones and the rest. The authors are explicit that the intent is “not to ‘black list’ functional groups” — what matters is the total body burden of reactive metabolite, which depends on dose, the fraction of clearance through the bioactivation route, and the availability of competing pathways.
Kalgutkar, Gardner, Obach, Shaffer, Callegari et al. Curr. Drug Metab.2005, 6, 161–225. 10.2174/1389200054021799
CNS penetration
The blood–brain barrier is two filters in series: passive permeability, which polarity governs, and active efflux, which basicity and polarity govern together.
CNS MPO score ≥ 4 of 6
Six desirability functions, summed, separate CNS drugs from CNS candidates better than any single cut-off.
Each parameter scores 0–1. Full marks at: ClogP ≤ 3; ClogD7.4 ≤ 2; MW ≤ 360; TPSA 40–90 Ų; HBD ≤ 0.5; most basic pKa ≤ 8. Zero at ClogP > 5, ClogD > 4, MW > 500, TPSA ≤ 20 or > 120, HBD > 3.5, pKa > 10. TPSA is the only two-sided function — a plateau with ramps on both sides.
Outcome:74% of 119 marketed CNS drugs scored ≥ 4, against 60% of 108 Pfizer CNS candidates. The value of the approach is that it grades rather than gates — a single breach can be paid for elsewhere.
Drawn from the published breakpoints. Five are one-sided ramps; TPSA is the only two-sided function, penalised for being too polar as well as not polar enough. Because the functions are continuous, a compound can fail one outright and still clear the ≥ 4 threshold — which is the point of scoring rather than gating. Each parameter scores 0–1 and the six sum to a score out of 6; the filled area is full desirability.
CNS attributes Pajouhesh & Lenz
A full profile for a successful CNS drug, physicochemical and in vitro together.
MW ≤ 450; ClogP ≤ 5; HBD ≤ 3; HBA ≤ 7; rotatable bonds ≤ 8; PSA ≤ 60–70 Ų; pKa7.5–10.5, neutral or basic — avoid acids; logD7.4 between 0 and 3; N + O count ≤ 5.
And the non-physicochemical half, usually omitted: a 30-fold margin between hERG IC50 and the efficacious unbound plasma concentration; ≥ 80% remaining after 1 h of metabolic stability incubation; CYP inhibition ≤ 50% at 30 µM; aqueous solubility ≥ 60 µg/mL.
Two properties predict P-glycoprotein substrate status with startling cleanliness.
Across >2000 compounds: fewer than 10% of those with TPSA < 60 Ųand most basic cpKa < 8 were P-gp substrates; more than 75% were substrates when TPSA > 60 and cpKa > 8. Either property alone is a much weaker predictor than the pair.
Not a donor countThe rule is framed in TPSA and basicity, not raw HBD. Desai’s companion paper argues H-bonds are not equivalent — relative H-bond strength and intramolecular H-bonding govern efflux, which undercuts naive per-donor counting.
Specific moves with predictable physicochemical consequences — the vocabulary you reach for once a property has to change without losing potency.
β-fluorination −1.7 pKa per F
Each fluorine beta to an amine lowers its basicity by roughly 1.7 pKa units.
The canonical series: ethylamine 10.6 → 2-fluoroethylamine 9.0 → 2,2-difluoroethylamine 7.5 → 2,2,2-trifluoroethylamine 5.9. This is the standard route out of a hERG or P-gp problem caused by a basic centre you cannot remove.
Position mattersThe effect is β (one carbon from N). An α-fluoroamine is not a design option — it is chemically unstable. γ-Fluorine effects are much smaller.
A methyl group adds about half a logP unit — and therefore costs about half a log unit of solubility.
This is the Hansch π substituent constant, π(CH₃) ≈ 0.5. Pair it with the magic-methyl entry above: the potency gain has to beat the lipophilicity cost, which is what an LLE calculation before and after the change will tell you.
Most property problems have a documented replacement that preserves the binding interaction.
Carboxylic acid to tetrazole or acylsulfonamide; phenyl to bicyclo[1.1.1]pentane; amide to oxadiazole or fluoroalkene; tert-butyl to trifluoromethylcyclopropyl. Meanwell’s two reviews are the standard working references — the 2018 one covers fluorinated motifs specifically, including CF₂H as a lipophilic hydrogen-bond donor and CF₃ as a metabolically robust methyl surrogate.
Local, transformation-specific evidence from your own series outperforms any global heuristic on this page.
Matched molecular pair analysis measures what a specific transformation did to a specific property, across every pair in a dataset where only that one change occurred. It is the correct answer whenever you have enough internal data — the rules here exist to cover the case where you do not.
The literature on property-based design contains its own critique. It is worth reading before leaning on any of the above.
molecular obesity the failure mode
Property inflation during optimisation is a behavioural problem, not a chemical one.
Potency is the property that gets measured first, weekly, and by everyone — so it is the property that gets optimised, and adding mass and lipophilicity is the reliable way to improve it. Hann’s argument is that the discipline’s incentive structure, not its chemistry, produces overweight candidates.
Some property distributions come from the target class; others come from how chemists behave.
Separating the two is the whole game. A property range that reflects what a binding site requires is a genuine constraint; one that reflects what was convenient to make is not, and following it forecloses chemical space for no reason.
Hann & Keserű. Nat. Rev. Drug Discov.2012, 11, 355–365. 10.1038/nrd3701
property profiles differ by route
“Drug-like” is not one distribution — drug-like, lead-like, CNS-drug-like, non-oral-drug-like and tool-like each have their own.
Lipinski’s own retrospective on the rule’s reception makes the point that applying an oral-absorption filter to an inhaled, injected or topical programme is a category error, and that tool compounds have no business being filtered at all.