Reagent · schema & worked example
The two needs that surfaced from the state-of-the-field review, written up; a proposed schema from functional groups through descriptors, properties and molecule–group graphs; a worked example dataset; and one correlation fitted end to end.
Papers/DFT-CECs · 2026-08-05 · fixture: example-data.json · chain step 3 (schema), DRAFTED
▸The two needs
The literature review produced two gaps that are not the same gap. One is a data asset nobody has built; the other is a method nobody has applied in this domain. They are separable — either could be pursued alone — but each makes the other stronger, and the schema below has to serve both.
Need 1 · A functional-group descriptor database for CECs
A curated store in which each record is one functional group at one named level of theory, carrying its computed descriptor set — frontier orbitals, ESP charges, Fukui indices, reactivity indices — with provenance sufficient to reproduce it.
Why it does not exist yet. QM9 caps near nine heavy atoms, so most CECs are out of range. QMugs holds drug-like whole molecules, not groups. CompTox holds QSAR predictions rather than computed descriptors. The state of the art for group-level descriptors is a Hammett-compatible calculator, not a database. Nothing published is indexed the way a screening campaign needs.
Why it is worth building. Groups are transferable: characterise a carboxyl once and the result informs every contaminant carrying one and every sorbent decorated with one. The database is therefore reusable across campaigns in a way a molecule-by-molecule calculation never is. And the cheap semi-empirical tier makes a full sweep affordable, with full DFT reserved for a shortlist.
Where the difficulty actually sits. Not in computing descriptors — in curating the measured properties to correlate them against. A few hundred trustworthy ozone rate constants, each with its pH, temperature, matrix and source, scattered across decades of papers. That curation is the scarce asset and a publishable contribution on its own.
Need 2 · Molecule–group graph metrics as correlating features
A molecule represented as a reduced graph whose nodes are functional groups and whose edges are the linkages between them, with graph-level statistics computed over it and offered to the model alongside the per-group descriptors.
Why this is not already done. Atom-level graph theory is saturated — Wiener 1947 through thousands of topological indices, now largely superseded by learned representations. Reduced graphs exist, but the literature took them to drug-discovery similarity searching, not to environmental-fate property prediction. The nearest property-side relative, group-contribution second-order groups, encodes neighbours as extra fitted terms rather than as genuine graph statistics.
Why it fits here specifically. Two reasons. The data regime: with a few hundred labels a graph neural network overfits, while a handful of interpretable graph statistics with quantum-chemical node features is the correct tool rather than a compromise. And the structural one: if groups are records, a molecule already is a graph over those records — so the representation costs no extra layer, it is the schema read another way.
The discipline it demands. Most topological indices track molecular size, so any correlation must survive controlling for size before it counts. The schema carries that guard as data (below), rather than leaving it to reviewer vigilance.
▸Proposed schema
Service prefix rg_. Read in dependency order: groups (themselves
graphs) → methods → descriptor sets → descriptors → molecules → the reduced graph → graph
metrics → properties → models. FKs are selectors, every column's intention declared at
authoring. The _type tables are controlled vocabularies, not free text, so a new
descriptor or metric is a row rather than a migration.
A · The functional group, as a graph
rg_group
One functional group. The reusable unit the whole system is built on.
- group_id PK
- code — short slug,
COOH. The human handle used everywhere. - name — "Carboxyl".
- smarts — the machine-readable definition. What makes a group detectable in a molecule rather than a label someone typed.
- smiles_fragment — depiction and round-trip.
- group_class — acidic · anionic · electron_donating · electron_withdrawing · aromatic · heteroaromatic · perfluoro · neutral_polar.
- n_heavy_atoms · attachment_points — attachment count decides whether the node can be terminal or must be a linker in the reduced graph.
- notes — the chemistry a reader needs; not a dumping ground.
rg_group_atom · rg_group_bond
The group's own internal graph. Not the graph pursued for research, but the schema holds it — so a group is a structure the system can reason over, not an opaque label.
- group_id FK → rg_group
- idx — group-local atom index. Per-atom descriptors reference this.
- element · aromatic · formal_charge · is_attachment_point
- atom_a, atom_b, order — single · double · triple · aromatic (bonds table).
B · Methods and descriptor sets
rg_method
A computational method as data — level of theory, engine, cost. Executed through the AnalystGateway compute envelope; never hardcoded.
- method_id PK · code · label
- engine — xtb · psi4 · orca-class adapter.
- tier — screening · production · production_anion. Drives campaign budgeting.
- level_of_theory · basis_set · solvent_model — together the comparability key.
- typical_seconds_per_molecule — the cost profile a campaign budget is planned against.
- reference_doi — the method's own citation, carried into published findings.
rg_descriptor_set
One group characterised by one method under one protonation state. THE record of the database — what "a functional-group record" actually means.
- set_id PK
- group_id FK → rg_group · method_id FK → rg_method
- protonation_state — neutral · anion · cation.
- converged · n_imaginary_freq — proof it is a true minimum. A non-converged set is visibly unusable, not silently averaged in.
- provenance — computed_here · imported_qmugs · literature.
- source_doi — required when provenance is not computed_here.
- status — PROPOSED → APPROVED. Imports land PROPOSED.
- computed_by FK → member · computed_at
C · Descriptors — one table, two scopes
rg_descriptor_type
Controlled vocabulary. A new descriptor family is a row, not a schema change.
- type_id PK · code · label · unit
- scope — molecular (HOMO) or atomic (ESP charge, Fukui). Declares which rows may carry an atom index.
- family — frontier_orbital · electrostatic · reactivity_index · site_reactivity · thermo.
- derived_from — array of codes this descriptor is computed from.
- definition
rg_descriptor
- descriptor_id PK · set_id FK → rg_descriptor_set · type_id FK → rg_descriptor_type
- atom_idx — NULL for molecular descriptors; the group-local atom index for atomic ones. One table serves both scopes.
- value
derived_from makes the colinearity rule machine-checkable.
Hardness, chemical potential and electrophilicity are algebraic functions of HOMO and LUMO —
two degrees of freedom wearing five names. Because each declares what it derives from, a model
quoting a descriptor together with its parents is flagged before it is fitted, not
caught in review. The methodological rule stops being vigilance and becomes a constraint.D · Measured properties — the scarce half
rg_property_type · rg_property_value
Properties are measured on MOLECULES, not groups — the asymmetry the whole modelling problem turns on.
- type: code — k_ozone · k_oh · log_kow · pka · log_kd · unit · log_scale
- value_id PK · molecule_id FK → rg_molecule · type_id FK
- value · value_sd
- ph · temperature_k · matrix · ionic_strength — conditions, because a rate constant without them is not a measurement.
- quality — measured · estimated · read_across.
- source_doi — required. A value without one cannot be APPROVED.
- entered_by · entered_at · status
E · Molecules and the reduced graph
rg_molecule
- molecule_id PK · code · name · smiles · inchikey · cas
- formula · mw — mw is not decoration: it is the covariate every graph metric must survive.
- cec_class — pharmaceutical_* · pfas · pesticide_* · industrial_edc.
rg_molecule_group — the NODES
M:N between molecules and group records. This join IS the reduced graph's node set.
- mg_id PK · molecule_id FK · group_id FK → rg_group
- node_idx — position within this molecule's graph.
- atom_map — which molecule atoms this node covers; makes the decomposition auditable and reversible.
rg_molecule_edge — the EDGES
- edge_id PK · molecule_id FK
- node_a, node_b FK → rg_molecule_group
- linker — direct · CH2 · C(CH3)2 · … the connecting fragment that is not itself a functional group.
- path_length — bonds between the two groups. Groups need not be directly bonded; electronic communication falls off with distance, so the model needs this.
rg_graph_metric_type · rg_graph_metric
- code — n_nodes · wiener_index · randic_index · graph_diameter · mean_degree · cyclomatic · donor_acceptor_separation
- size_correlated — boolean guard: true means this metric may not enter a model unless the fit demonstrably survives controlling for molecular size.
- metric_id PK · molecule_id FK · type_id FK · value
size_correlated forces a model to control for.F · Models — fitted, provenanced, domain-bounded
rg_model · rg_model_term · rg_model_training_row · rg_model_prediction
- model_id PK · campaign_id FK · target_type_id FK → rg_property_type
- form — lfer · ols · group_contribution · pca_regression · response_transform · equation
- n_train · r2 · rmse
- size_controlled · size_control_note — did the fit survive the covariate test, and what happened when it was run.
- applicability_domain — the group space and descriptor range fitted over. Predictions outside it refuse loudly.
- term: feature_kind (descriptor · graph_metric · intercept) · feature_code · aggregation · coefficient · std_error · t_stat · p_value · vif
- training_row: model_id + molecule_id + property_value_id — exactly what it learned from; the provenance a published claim cites.
aggregation is the group-to-molecule bridge, and it is part of
the model definition rather than hidden preprocessing. max_over_nodes takes the
highest HOMO among a molecule's group nodes — chemically right for ozone, because the most
electron-rich site governs the rate. sum_over_nodes gives group-contribution
behaviour. Recording which was used, and which node governed each prediction, is what lets a
result say why.▸Worked example data
source_doi, or an actual run, before entering any real world. The fixture
(example-data.json) is a design artefact and is never seeded into a database.Seven CECs, decomposed into group nodes
| Molecule | Class | Reduced-graph nodes | Nodes | MW |
|---|---|---|---|---|
| Sulfamethoxazole | antibiotic | ArNH₂ — PhR — SO₂NH — ISOX | 4 | 253.3 |
| Diclofenac | NSAID | COOH — PhR — ArNH — PhR — 2 × ArCl | 6 | 296.2 |
| Bisphenol A | industrial EDC | ArOH — PhR — PhR — ArOH | 4 | 228.3 |
| Carbamazepine | anticonvulsant | CONH₂ — PhR — olefin — PhR | 4 | 236.3 |
| Ibuprofen | NSAID | COOH — PhR — alkyl | 3 | 206.3 |
| Atrazine | triazine pesticide | AlkNH — triazine — AlkNH — ArCl | 4 | 215.7 |
| PFOA | PFAS | COOH — 6 × CF₂ — CF₃ | 8 | 414.1 |
Diclofenac's graph is a chain with one branch point: the
carboxyl hangs off ring A through a CH₂ linker (path_length 2, not 1), ring A joins
ring B through the secondary amine, and both chlorines attach to ring B. That branching is what
mean_degree and cyclomatic pick up, and what a flat group count throws
away.
▸A worked correlation
One model fitted end to end on the fixture, to show what the schema produces: log₁₀ of the ozone rate constant against the highest HOMO among a molecule's functional group nodes. The feature comes from Need 1 (group descriptors); the aggregation over nodes comes from Need 2 (the reduced graph). Arithmetic computed, not asserted.
The fit, and the data behind it
| Molecule | Governing node | HOMO (eV) | k(O₃) | log₁₀ obs | pred | residual |
|---|---|---|---|---|---|---|
| Sulfamethoxazole | aniline amine | −5.21 | 2.5×10⁶ | 6.40 | 7.37 | −0.97 |
| Diclofenac | secondary aryl amine | −5.35 | 1.0×10⁶ | 6.00 | 6.87 | −0.87 |
| Bisphenol A | phenolic OH | −5.94 | 1.7×10⁶ | 6.23 | 4.74 | +1.49 |
| Carbamazepine | olefinic bridge | −6.20 | 3.0×10⁵ | 5.48 | 3.80 | +1.67 |
| Ibuprofen | benzene ring | −6.90 | 9.6 | 0.98 | 1.28 | −0.30 |
| Atrazine | alkyl amine | −6.95 | 6.0 | 0.78 | 1.10 | −0.32 |
| PFOA | carboxyl (nothing else reacts) | −7.42 | 0.05 | −1.30 | −0.59 | −0.71 |
R² = 0.882 · RMSE = 1.22 log units · slope SE 0.590, t = 6.11, p = 0.0017 (n = 7). Each electron-volt of HOMO destabilisation buys about 3.6 orders of magnitude in ozone rate constant — a slope with chemical meaning, which is what makes an LFER worth preferring over a black box when it fits.
The survive-controls test — run, not assumed
Because a correlation this clean invites the size objection, molecular weight was added as a covariate and the model refitted:
| Term | b (simple) | b (MW controlled) | t | p | Verdict |
|---|---|---|---|---|---|
| HOMOmax | 3.604 | 3.452 | 5.53 | 0.005 | Sign, magnitude and significance retained — survives |
| Molecular weight | — | −0.0066 | −0.90 | 0.42 | Not significant — size is not driving this |
| R² rises only 0.882 → 0.902 on adding MW. r(HOMO, MW) = −0.270, VIF 1.08 — the two features are near-independent, so the electronic effect is not size in disguise. | |||||