Reagent · schema & worked example

The two needs that surfaced from the state-of-the-field review, written up; a proposed schema from functional groups through descriptors, properties and molecule–group graphs; a worked example dataset; and one correlation fitted end to end.

Papers/DFT-CECs · 2026-08-05 · fixture: example-data.json · chain step 3 (schema), DRAFTED

▸The two needs

The literature review produced two gaps that are not the same gap. One is a data asset nobody has built; the other is a method nobody has applied in this domain. They are separable — either could be pursued alone — but each makes the other stronger, and the schema below has to serve both.

Need 1 · A functional-group descriptor database for CECs

A curated store in which each record is one functional group at one named level of theory, carrying its computed descriptor set — frontier orbitals, ESP charges, Fukui indices, reactivity indices — with provenance sufficient to reproduce it.

Why it does not exist yet. QM9 caps near nine heavy atoms, so most CECs are out of range. QMugs holds drug-like whole molecules, not groups. CompTox holds QSAR predictions rather than computed descriptors. The state of the art for group-level descriptors is a Hammett-compatible calculator, not a database. Nothing published is indexed the way a screening campaign needs.

Why it is worth building. Groups are transferable: characterise a carboxyl once and the result informs every contaminant carrying one and every sorbent decorated with one. The database is therefore reusable across campaigns in a way a molecule-by-molecule calculation never is. And the cheap semi-empirical tier makes a full sweep affordable, with full DFT reserved for a shortlist.

Where the difficulty actually sits. Not in computing descriptors — in curating the measured properties to correlate them against. A few hundred trustworthy ozone rate constants, each with its pH, temperature, matrix and source, scattered across decades of papers. That curation is the scarce asset and a publishable contribution on its own.

Need 2 · Molecule–group graph metrics as correlating features

A molecule represented as a reduced graph whose nodes are functional groups and whose edges are the linkages between them, with graph-level statistics computed over it and offered to the model alongside the per-group descriptors.

Why this is not already done. Atom-level graph theory is saturated — Wiener 1947 through thousands of topological indices, now largely superseded by learned representations. Reduced graphs exist, but the literature took them to drug-discovery similarity searching, not to environmental-fate property prediction. The nearest property-side relative, group-contribution second-order groups, encodes neighbours as extra fitted terms rather than as genuine graph statistics.

Why it fits here specifically. Two reasons. The data regime: with a few hundred labels a graph neural network overfits, while a handful of interpretable graph statistics with quantum-chemical node features is the correct tool rather than a compromise. And the structural one: if groups are records, a molecule already is a graph over those records — so the representation costs no extra layer, it is the schema read another way.

The discipline it demands. Most topological indices track molecular size, so any correlation must survive controlling for size before it counts. The schema carries that guard as data (below), rather than leaving it to reviewer vigilance.

How the two needs meet. Need 1 supplies the node features; Need 2 supplies the topology and the group-to-molecule bridge. A model with only Need 1 cannot say where in a molecule a group sits; a model with only Need 2 has topology with no chemistry in it. Together they produce a feature matrix that is interpretable at both levels — which group, and where.

▸Proposed schema

Service prefix rg_. Read in dependency order: groups (themselves graphs) → methods → descriptor sets → descriptors → molecules → the reduced graph → graph metrics → properties → models. FKs are selectors, every column's intention declared at authoring. The _type tables are controlled vocabularies, not free text, so a new descriptor or metric is a row rather than a migration.

A · The functional group, as a graph

rg_group

One functional group. The reusable unit the whole system is built on.

  • group_id PK
  • code — short slug, COOH. The human handle used everywhere.
  • name — "Carboxyl".
  • smarts — the machine-readable definition. What makes a group detectable in a molecule rather than a label someone typed.
  • smiles_fragment — depiction and round-trip.
  • group_class — acidic · anionic · electron_donating · electron_withdrawing · aromatic · heteroaromatic · perfluoro · neutral_polar.
  • n_heavy_atoms · attachment_points — attachment count decides whether the node can be terminal or must be a linker in the reduced graph.
  • notes — the chemistry a reader needs; not a dumping ground.

rg_group_atom  ·  rg_group_bond

The group's own internal graph. Not the graph pursued for research, but the schema holds it — so a group is a structure the system can reason over, not an opaque label.

  • group_id FK → rg_group
  • idx — group-local atom index. Per-atom descriptors reference this.
  • element · aromatic · formal_charge · is_attachment_point
  • atom_a, atom_b, order — single · double · triple · aromatic (bonds table).
Carboxyl and carboxylate are two group records, not one with a flag. Their descriptor values differ materially, the anion needs diffuse basis functions, and at treatment pH it is the carboxylate that is actually present. Collapsing them would make the database quietly wrong at exactly the pH that matters.

B · Methods and descriptor sets

rg_method

A computational method as data — level of theory, engine, cost. Executed through the AnalystGateway compute envelope; never hardcoded.

  • method_id PK · code · label
  • engine — xtb · psi4 · orca-class adapter.
  • tier — screening · production · production_anion. Drives campaign budgeting.
  • level_of_theory · basis_set · solvent_model — together the comparability key.
  • typical_seconds_per_molecule — the cost profile a campaign budget is planned against.
  • reference_doi — the method's own citation, carried into published findings.

rg_descriptor_set

One group characterised by one method under one protonation state. THE record of the database — what "a functional-group record" actually means.

  • set_id PK
  • group_id FK → rg_group · method_id FK → rg_method
  • protonation_state — neutral · anion · cation.
  • converged · n_imaginary_freq — proof it is a true minimum. A non-converged set is visibly unusable, not silently averaged in.
  • provenance — computed_here · imported_qmugs · literature.
  • source_doi — required when provenance is not computed_here.
  • status — PROPOSED → APPROVED. Imports land PROPOSED.
  • computed_by FK → member · computed_at

C · Descriptors — one table, two scopes

rg_descriptor_type

Controlled vocabulary. A new descriptor family is a row, not a schema change.

  • type_id PK · code · label · unit
  • scope — molecular (HOMO) or atomic (ESP charge, Fukui). Declares which rows may carry an atom index.
  • family — frontier_orbital · electrostatic · reactivity_index · site_reactivity · thermo.
  • derived_from — array of codes this descriptor is computed from.
  • definition

rg_descriptor

  • descriptor_id PK · set_id FK → rg_descriptor_set · type_id FK → rg_descriptor_type
  • atom_idx — NULL for molecular descriptors; the group-local atom index for atomic ones. One table serves both scopes.
  • value
derived_from makes the colinearity rule machine-checkable. Hardness, chemical potential and electrophilicity are algebraic functions of HOMO and LUMO — two degrees of freedom wearing five names. Because each declares what it derives from, a model quoting a descriptor together with its parents is flagged before it is fitted, not caught in review. The methodological rule stops being vigilance and becomes a constraint.

D · Measured properties — the scarce half

rg_property_type  ·  rg_property_value

Properties are measured on MOLECULES, not groups — the asymmetry the whole modelling problem turns on.

  • type: code — k_ozone · k_oh · log_kow · pka · log_kd · unit · log_scale
  • value_id PK · molecule_id FK → rg_molecule · type_id FK
  • value · value_sd
  • ph · temperature_k · matrix · ionic_strength — conditions, because a rate constant without them is not a measurement.
  • quality — measured · estimated · read_across.
  • source_doi — required. A value without one cannot be APPROVED.
  • entered_by · entered_at · status

E · Molecules and the reduced graph

rg_molecule

  • molecule_id PK · code · name · smiles · inchikey · cas
  • formula · mw — mw is not decoration: it is the covariate every graph metric must survive.
  • cec_class — pharmaceutical_* · pfas · pesticide_* · industrial_edc.

rg_molecule_group  — the NODES

M:N between molecules and group records. This join IS the reduced graph's node set.

  • mg_id PK · molecule_id FK · group_id FK → rg_group
  • node_idx — position within this molecule's graph.
  • atom_map — which molecule atoms this node covers; makes the decomposition auditable and reversible.

rg_molecule_edge  — the EDGES

  • edge_id PK · molecule_id FK
  • node_a, node_b FK → rg_molecule_group
  • linker — direct · CH2 · C(CH3)2 · … the connecting fragment that is not itself a functional group.
  • path_length — bonds between the two groups. Groups need not be directly bonded; electronic communication falls off with distance, so the model needs this.

rg_graph_metric_type  ·  rg_graph_metric

  • code — n_nodes · wiener_index · randic_index · graph_diameter · mean_degree · cyclomatic · donor_acceptor_separation
  • size_correlated — boolean guard: true means this metric may not enter a model unless the fit demonstrably survives controlling for molecular size.
  • metric_id PK · molecule_id FK · type_id FK · value
The size trap is visible in the fixture itself. PFOA has the most group nodes (8) and the lowest reactivity; bisphenol A has among the fewest (4) and the highest. A size metric would correlate negatively with rate constant across these seven molecules and look predictive — purely as an artefact of which molecules were chosen. This is exactly what size_correlated forces a model to control for.

F · Models — fitted, provenanced, domain-bounded

rg_model  ·  rg_model_term  ·  rg_model_training_row  ·  rg_model_prediction

  • model_id PK · campaign_id FK · target_type_id FK → rg_property_type
  • form — lfer · ols · group_contribution · pca_regression · response_transform · equation
  • n_train · r2 · rmse
  • size_controlled · size_control_note — did the fit survive the covariate test, and what happened when it was run.
  • applicability_domain — the group space and descriptor range fitted over. Predictions outside it refuse loudly.
  • term: feature_kind (descriptor · graph_metric · intercept) · feature_code · aggregation · coefficient · std_error · t_stat · p_value · vif
  • training_row: model_id + molecule_id + property_value_id — exactly what it learned from; the provenance a published claim cites.
aggregation is the group-to-molecule bridge, and it is part of the model definition rather than hidden preprocessing. max_over_nodes takes the highest HOMO among a molecule's group nodes — chemically right for ozone, because the most electron-rich site governs the rate. sum_over_nodes gives group-contribution behaviour. Recording which was used, and which node governed each prediction, is what lets a result say why.

▸Worked example data

Read this before using any number below. Molecule identities, functional group decompositions and SMARTS patterns are real chemistry. Every numeric value is illustrative — ozone rate constants are order-of-magnitude figures and DFT descriptors are plausible values, not curated measurements or actual converged runs. They exist to prove the schema holds a real case end to end. Each must be replaced by a curated value with its own source_doi, or an actual run, before entering any real world. The fixture (example-data.json) is a design artefact and is never seeded into a database.

Seven CECs, decomposed into group nodes

MoleculeClassReduced-graph nodesNodesMW
SulfamethoxazoleantibioticArNH₂ — PhR — SO₂NH — ISOX4253.3
DiclofenacNSAIDCOOH — PhR — ArNH — PhR — 2 × ArCl6296.2
Bisphenol Aindustrial EDCArOH — PhR — PhR — ArOH4228.3
CarbamazepineanticonvulsantCONH₂ — PhR — olefin — PhR4236.3
IbuprofenNSAIDCOOH — PhR — alkyl3206.3
Atrazinetriazine pesticideAlkNH — triazine — AlkNH — ArCl4215.7
PFOAPFASCOOH — 6 × CF₂ — CF₃8414.1

Diclofenac's graph is a chain with one branch point: the carboxyl hangs off ring A through a CH₂ linker (path_length 2, not 1), ring A joins ring B through the secondary amine, and both chlorines attach to ring B. That branching is what mean_degree and cyclomatic pick up, and what a flat group count throws away.

▸A worked correlation

One model fitted end to end on the fixture, to show what the schema produces: log₁₀ of the ozone rate constant against the highest HOMO among a molecule's functional group nodes. The feature comes from Need 1 (group descriptors); the aggregation over nodes comes from Need 2 (the reduced graph). Arithmetic computed, not asserted.

Ozone reactivity tracks the most electron-rich functional group
log₁₀ kO₃ vs. highest group HOMO · 7 CECs · illustrative values
7 5 3 1 −1 −7.6 −7.0 −6.4 −5.8 −5.2 Highest group HOMO (eV) log₁₀ k(O₃) / M⁻¹s⁻¹ fitted LFER slope 3.60 PFOA · HOMO −7.42 eV · k 0.05 M⁻¹s⁻¹ · governing node COOH · residual −0.71 PFOA Atrazine · HOMO −6.95 eV · k 6.0 M⁻¹s⁻¹ · governing node alkyl amine · residual −0.32 Atrazine Ibuprofen · HOMO −6.90 eV · k 9.6 M⁻¹s⁻¹ · governing node benzene ring · residual −0.30 Ibuprofen Carbamazepine · HOMO −6.20 eV · k 3.0×10⁵ M⁻¹s⁻¹ · governing node olefinic bridge · residual +1.67 (largest) Carbamazepine Bisphenol A · HOMO −5.94 eV · k 1.7×10⁶ M⁻¹s⁻¹ · governing node phenolic OH · residual +1.49 Bisphenol A Diclofenac · HOMO −5.35 eV · k 1.0×10⁶ M⁻¹s⁻¹ · governing node secondary aryl amine · residual −0.87 Diclofenac Sulfamethoxazole · HOMO −5.21 eV · k 2.5×10⁶ M⁻¹s⁻¹ · governing node aniline amine · residual −0.97 Sulfamethoxazole
Dashed drops show each residual against the fit. Hover any point for its governing functional group. Seven orders of magnitude in rate constant across a 2.2 eV span in HOMO — the dynamic range that makes ozone the discriminating target and hydroxyl radical the misleading one.

The fit, and the data behind it

MoleculeGoverning nodeHOMO (eV)k(O₃)log₁₀ obspredresidual
Sulfamethoxazoleaniline amine−5.212.5×10⁶6.407.37−0.97
Diclofenacsecondary aryl amine−5.351.0×10⁶6.006.87−0.87
Bisphenol Aphenolic OH−5.941.7×10⁶6.234.74+1.49
Carbamazepineolefinic bridge−6.203.0×10⁵5.483.80+1.67
Ibuprofenbenzene ring−6.909.60.981.28−0.30
Atrazinealkyl amine−6.956.00.781.10−0.32
PFOAcarboxyl (nothing else reacts)−7.420.05−1.30−0.59−0.71
log₁₀ k(O₃) = 3.604 × HOMOmax + 26.15
R² = 0.882 · RMSE = 1.22 log units · slope SE 0.590, t = 6.11, p = 0.0017 (n = 7). Each electron-volt of HOMO destabilisation buys about 3.6 orders of magnitude in ozone rate constant — a slope with chemical meaning, which is what makes an LFER worth preferring over a black box when it fits.

The survive-controls test — run, not assumed

Because a correlation this clean invites the size objection, molecular weight was added as a covariate and the model refitted:

Termb (simple)b (MW controlled)tpVerdict
HOMOmax3.6043.4525.530.005Sign, magnitude and significance retained — survives
Molecular weight—−0.0066−0.900.42Not significant — size is not driving this
R² rises only 0.882 → 0.902 on adding MW. r(HOMO, MW) = −0.270, VIF 1.08 — the two features are near-independent, so the electronic effect is not size in disguise.
What this worked example does not establish. Seven points on illustrative numbers demonstrates that the schema, the aggregation bridge and the control test all function together. It is not a defensible model and must never be quoted as one. The largest residual is carbamazepine at +1.67 log units, and the reason is chemically legible: its ozone chemistry is governed by the olefinic bridge rather than by any group's HOMO, which a single-descriptor LFER cannot express. That failure is visible rather than hidden — which is the argument for interpretable models in the small-data regime, and it points at the next model rung.