diff --git a/Biochemistry/REACTIONS.md b/Biochemistry/REACTIONS.md index 225f539f..20bc97fa 100644 --- a/Biochemistry/REACTIONS.md +++ b/Biochemistry/REACTIONS.md @@ -28,6 +28,25 @@ files. They are described here: ### Format of stoichiometry field +### Compartment indices + +The compartment index `m` that appears in the formats below is **relative, not +absolute**. `0` is the inside of the cell relative to the reaction and `1` is +the outside; neither names a particular compartment. The reaction +`(1) cpd00002[0] + (1) cpd00009[1] => ...` therefore describes "hydrolyse ATP +on the inside and move phosphate in from the outside" without committing to +which membrane that is, and the same record can be instantiated against a +bacterial plasma membrane, a mitochondrial inner membrane or a plastid +envelope. A model template resolves the indices to named compartments (`c0`, +`e0`, and organelles) when it builds a model. + +This is what Rhea expresses as `in`/`out` and maps onto the two indices +directly. It also means a reaction record carries no membrane potential and no +pH gradient, since both are properties of an organism and a condition rather +than of the reaction; thermodynamic values distributed with the database are +computed from stoichiometry and energy alone, and transport reactions are +flagged so that this can be accounted for downstream. + ### Format of reaction definition using compound IDs Each compound participating in the reaction is in this format: diff --git a/Papers/NAR_Update_2026/latex/sections/M05_methods_thermodynamics.tex b/Papers/NAR_Update_2026/latex/sections/M05_methods_thermodynamics.tex index 55cee6c4..bbd19235 100644 --- a/Papers/NAR_Update_2026/latex/sections/M05_methods_thermodynamics.tex +++ b/Papers/NAR_Update_2026/latex/sections/M05_methods_thermodynamics.tex @@ -4,12 +4,14 @@ %% Supplementary Methods S2. \subsection{Multi-source thermodynamics}\label{sec:methods-thermo} -The 2020 release combined two sources: the historical group contribution approach~\cite{jankowski2008} and eQuilibrator~\cite{flamholz2012}. This update carries four sources, each stored as its own record: group contribution, eQuilibrator~3.0~\cite{beber2022}, dGPredictor~\cite{wang2021}, and experimental values where they exist. Only the experimental values carry a published direction. The three predictors were released as energy estimators, so every direction reported here is an inference from their energies. All values are reported at pH~7.0, ionic strength 0.25\,M, pMg~3.0 and 298.15\,K; the basis for each condition and its effect on the released energies are given in Supplementary Methods~S2. A key caveat: as we publish a generalized database that can be applied across the kingdoms of life, transport reactions are scored from stoichiometry and energy alone, and neither an estimate of membrane potential or pH gradients are used. +The 2020 release combined two sources: the historical group contribution approach~\cite{jankowski2008} and eQuilibrator~\cite{flamholz2012}. This update carries four sources, each stored as its own record: group contribution, eQuilibrator~3.0~\cite{beber2022}, dGPredictor~\cite{wang2021}, and experimental values where they exist. Only the experimental values carry a published direction. The three predictors were released as energy estimators, so every direction reported here is an inference from their energies. All values are reported at pH~7.0, ionic strength 0.25\,M, pMg~3.0 and 298.15\,K; the basis for each condition and its effect on the released energies are given in Supplementary Methods~S2. \paragraph{Protonation.} eQuilibrator's Legendre transform requires a \emph{macroscopic} pKa ladder per compound: an ordered list in which consecutive species differ by one proton. The count of entries around the desired pH sets the proton count of the major microspecies, so the composition of the ladder changes the transformed energy directly. We generate the pKa ladders using ChemAxon Marvin~\cite{marvin} for use with eQuilibrator, but we're conscious that these data cannot be regenerated without a license, so we also include the pKas generated by an open-source alternative, MolGpKa~\cite{pan2021}, for every compound with a structure, and the measured ladders of Alberty~\cite{alberty2003}, versioned per structure source in an extended pKa dictionary. -\paragraph{Experimental anchors.} eQuilibrator and dGPredictor draw on the same set of experimental values derived from the NIST Thermodynamics of Enzyme-catalyzed Reactions (TECR) Database~\cite{goldberg2004}. The work was keyed to KEGG, and eQuilibrator uses MetaNetX mappings to integrate with other resources but there are inherent structural conflicts when we examine the mappings to ModelSEED identifiers. In order to draw direct comparisons between the results generated for any ModelSEED reactions we re-built the experimental anchors of TECR values, curating any conflicts to ensure a direct mapping between TECR values and ModelSEED reactions. In addition we evaluated openTECR, a communal re-curation of the same source, available at \url{https://github.com/opentecr}, to see if we can extend the set of anchors. +\paragraph{Experimental anchors.} eQuilibrator and dGPredictor both draw on the NIST Thermodynamics of Enzyme-catalyzed Reactions (TECR) Database~\cite{goldberg2004}, which is keyed to KEGG and reaches other namespaces through MetaNetX mappings that conflict structurally with ModelSEED identifiers. We therefore rebuilt the anchors, curating each conflict so that every TECR value maps directly onto a ModelSEED reaction, and evaluated openTECR (\url{https://github.com/opentecr}), a communal re-curation of the same source, to extend the set. -\paragraph{Uncertainty.} The three sources each adopt a different approach for estimating the uncertainty surrounding the predicted energy of reaction and as such the scale of uncertainty is not directly comparable (Figure~\ref{fig:thermo}C)) median values (kcal\,mol$^{-1}$) are 0.63 for eQuilibrator, 10.41 for group contribution, and 17.01 for dGPredictor. Comparing against the experimentally-anchored reactions, we find that group contribution overstates its error and eQuilibrator understates it. +\paragraph{Uncertainty.} The three sources each adopt a different approach for estimating the uncertainty surrounding the predicted energy of reaction, so the scale of uncertainty is not directly comparable between them (Figure~\ref{fig:thermo}C). We therefore calibrate each source's reported uncertainty against the experimental anchors rather than reading it at face value; the resulting medians and the direction of each source's bias are reported with the results. -\paragraph{Direction.} For historical reasons, we use the same heuristics for deriving reaction direction from the Group contribution approach~\cite{jankowski2008} unchanged from 2020, so that series stays comparable across releases. For eQuilibrator and dGPredictor we use one rule set built on the reversibility index from Noor \textit{et al.}~\cite{noor2012}: a reaction is called irreversible when $|\ln\Gamma|$ exceeds $\ln(1000)$ by more than one propagated standard deviation~\cite{beber2022}, $\Gamma$ being the fold-change of reactant concentration required to change reaction direction). However, we extend the work by Noor \textit{et al.} to require the uncertainty to clear the threshold as well. We report fewer directions than the published index would on the same energies, and we separate \emph{reversible} reactions from \emph{undetermined} cases. An exception we implement is for reactions whose uncertainty is small relative to the threshold and yet still straddles it: meaning the estimate lies within one standard deviation of the threshold, and these we report as reversible. \ No newline at end of file +\paragraph{Direction.} For historical reasons, we use the same heuristics for deriving reaction direction from the Group contribution approach~\cite{jankowski2008} unchanged from 2020, so that series stays comparable across releases. For eQuilibrator and dGPredictor we use one rule set built on the reversibility index from Noor \textit{et al.}~\cite{noor2012}: a reaction is called irreversible when $|\ln\Gamma|$ exceeds $\ln(1000)$ by more than one propagated standard deviation~\cite{beber2022}, $\Gamma$ being the fold-change of reactant concentration required to change reaction direction). However, we extend the work by Noor \textit{et al.} to require the uncertainty to clear the threshold as well. We report fewer directions than the published index would on the same energies, and we separate \emph{reversible} reactions from \emph{undetermined} cases. An exception we implement is for reactions whose uncertainty is small relative to the threshold and yet still straddles it: meaning the estimate lies within one standard deviation of the threshold, and these we report as reversible. + +\paragraph{Transport.} Compartments are \emph{relative} indices --- 0 inside, 1 outside --- not named compartments, so one reaction can be instantiated against a bacterial membrane, a mitochondrion or a plastid envelope when a template builds a model. Because the database is not committed to a cell type, transport reactions are scored from stoichiometry and energy alone: neither membrane potential nor pH gradient enters the estimate. The consequence is specific. A primary active transporter carries ATP hydrolysis in its stoichiometry, every predictor estimates that tightly, and the reaction earns a confident grade from chemistry that was never in question --- 96\% of gold-graded transport reactions are ATP-coupled and only 2.1\% translocate a proton. Such grades measure confidence in the coupled chemistry, not evidence that the transporter runs that way in a given organism; every reaction carries an \texttt{is\_transport} field so that this can be accounted for downstream. Reactions whose only change is of compartment carry no net chemistry and none is graded gold. dGPredictor commits on 1.4\% of transport reactions against 17.3\% elsewhere, so transport grades also rest on fewer independent sources. diff --git a/Papers/NAR_Update_2026/latex/sections/M10_results_structure_curation.tex b/Papers/NAR_Update_2026/latex/sections/M10_results_structure_curation.tex index 49f0a645..d79d2424 100644 --- a/Papers/NAR_Update_2026/latex/sections/M10_results_structure_curation.tex +++ b/Papers/NAR_Update_2026/latex/sections/M10_results_structure_curation.tex @@ -1,3 +1,3 @@ \subsection{Structure curation}\label{sec:results-structure} -We built and ran our curation pipeline on all structures that would impact the set of roughly 9,000 reactions used in the ModelSEED and PlantSEED reconstruction templates. The conflict pipeline resolved every previously conflicting compound in three curation rounds. i) Instances where the conflict existed at the level of mobile hydrogens were resolved automatically using historical formula ii) instances where the structures were structurally different were resolved manually iii) instances where resolution was not straight-forward were entered in the mass-balance exclusion registry where we fix an arbitrarily chosen formula because no source formula was defensible. \ No newline at end of file +We prioritised the curation pipeline on the compounds whose structures affect the ${\sim}9{,}000$ reactions carried by the ModelSEED and PlantSEED reconstruction templates, since an error in one of those propagates into every model built from them. The scope is a curation priority, not a statement about how much of the database is usable: the templates are discussed in Section~\ref{sec:results-mapping}. The conflict pipeline resolved every previously conflicting compound in three curation rounds. i) Instances where the conflict existed at the level of mobile hydrogens were resolved automatically using historical formula ii) instances where the structures were structurally different were resolved manually iii) instances where resolution was not straight-forward were entered in the mass-balance exclusion registry where we fix an arbitrarily chosen formula because no source formula was defensible. \ No newline at end of file diff --git a/Papers/NAR_Update_2026/latex/sections/M13_results_mapping.tex b/Papers/NAR_Update_2026/latex/sections/M13_results_mapping.tex index 8c8b46b7..9eab7941 100644 --- a/Papers/NAR_Update_2026/latex/sections/M13_results_mapping.tex +++ b/Papers/NAR_Update_2026/latex/sections/M13_results_mapping.tex @@ -2,4 +2,6 @@ %% connectivity_ontology} sections. \subsection{Atom mapping}\label{sec:results-mapping} -Atom mappings are published for 32,877 reactions, 59\% of the database: 25,058 clean, and 7,819 that required salvage (which are tagged). As the mappings can be used to span a metabolic reconstruction, reactions needing one repair are likely to carry others. \ No newline at end of file +Atom mappings are published for 32,877 reactions, 59\% of the database: 25,058 clean, and 7,819 that required salvage (which are tagged). As the mappings can be used to span a metabolic reconstruction, reactions needing one repair are likely to carry others. + +\paragraph{Reconstruction scope.} A template carries far fewer reactions than the database holds --- the v7.0 ModelSEED templates span 8,597 in union, about 9,000 with PlantSEED\_v3 roles --- but that gap does not mean the remainder is unusable. The database is a reference namespace across all kingdoms and includes biochemistry from secondary databases and published models that no single template targets. Promotion into a template is limited by mass and charge balance, which 58\% of reactions satisfy, and by association to an annotated functional role, which 41\% carry. Gene association is not the constraint: 2,258 of the 8,584 reactions in the Gram-negative template are typed as gapfilling and require no gene. 14,041 reactions now satisfy both balance and role annotation, 8,089 of them outside every current template. diff --git a/Papers/NAR_Update_2026/latex/sections/M14_discussion.tex b/Papers/NAR_Update_2026/latex/sections/M14_discussion.tex index e809f663..505d6eda 100644 --- a/Papers/NAR_Update_2026/latex/sections/M14_discussion.tex +++ b/Papers/NAR_Update_2026/latex/sections/M14_discussion.tex @@ -1,3 +1,3 @@ \section{Discussion} -Unlike the 2020 release, this update adopts a multi-source policy rather than promoting a single thermodynamic value, publishing four sources, each with their own uncertainty and estimated direction, because the sources may disagree in ways a single value conceals. Grading against experiment enables us to validate the approach: A researcher can take the highest-graded source and knows what the grade means. In being explicit about disagreement between sources, we expect this release to serve the widening range of communities that build on shared biochemistry: genome-scale reconstruction, flux analysis, pathway design and the training of predictive models. \ No newline at end of file +Unlike the 2020 release, this update adopts a multi-source policy rather than promoting a single thermodynamic value, publishing four sources, each with their own uncertainty and estimated direction, because the sources may disagree in ways a single value conceals. Grading against experiment enables us to validate the approach: A researcher can take the highest-graded source and knows what the grade means. Two limits are worth naming. Transport reactions are scored without a membrane potential or a pH gradient, because both are properties of an organism and a condition rather than of a reference reaction; a compartment-aware scoring layer is the natural next release, and it needs the organism context a model supplies and a database cannot. And roughly half the database still carries no thermodynamic direction at all, bounded by the structures needed to compute one. In being explicit about disagreement between sources, we expect this release to serve the widening range of communities that build on shared biochemistry: genome-scale reconstruction, flux analysis, pathway design and the training of predictive models. \ No newline at end of file