Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions Biochemistry/REACTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,25 @@ files. They are described here:

### Format of stoichiometry field

### Compartment indices

The compartment index `m` that appears in the formats below is **relative, not
absolute**. `0` is the inside of the cell relative to the reaction and `1` is
the outside; neither names a particular compartment. The reaction
`(1) cpd00002[0] + (1) cpd00009[1] => ...` therefore describes "hydrolyse ATP
on the inside and move phosphate in from the outside" without committing to
which membrane that is, and the same record can be instantiated against a
bacterial plasma membrane, a mitochondrial inner membrane or a plastid
envelope. A model template resolves the indices to named compartments (`c0`,
`e0`, and organelles) when it builds a model.

This is what Rhea expresses as `in`/`out` and maps onto the two indices
directly. It also means a reaction record carries no membrane potential and no
pH gradient, since both are properties of an organism and a condition rather
than of the reaction; thermodynamic values distributed with the database are
computed from stoichiometry and energy alone, and transport reactions are
flagged so that this can be accounted for downstream.

### Format of reaction definition using compound IDs
Each compound participating in the reaction is in this format:

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,12 +4,14 @@
%% Supplementary Methods S2.
\subsection{Multi-source thermodynamics}\label{sec:methods-thermo}

The 2020 release combined two sources: the historical group contribution approach~\cite{jankowski2008} and eQuilibrator~\cite{flamholz2012}. This update carries four sources, each stored as its own record: group contribution, eQuilibrator~3.0~\cite{beber2022}, dGPredictor~\cite{wang2021}, and experimental values where they exist. Only the experimental values carry a published direction. The three predictors were released as energy estimators, so every direction reported here is an inference from their energies. All values are reported at pH~7.0, ionic strength 0.25\,M, pMg~3.0 and 298.15\,K; the basis for each condition and its effect on the released energies are given in Supplementary Methods~S2. A key caveat: as we publish a generalized database that can be applied across the kingdoms of life, transport reactions are scored from stoichiometry and energy alone, and neither an estimate of membrane potential or pH gradients are used.
The 2020 release combined two sources: the historical group contribution approach~\cite{jankowski2008} and eQuilibrator~\cite{flamholz2012}. This update carries four sources, each stored as its own record: group contribution, eQuilibrator~3.0~\cite{beber2022}, dGPredictor~\cite{wang2021}, and experimental values where they exist. Only the experimental values carry a published direction. The three predictors were released as energy estimators, so every direction reported here is an inference from their energies. All values are reported at pH~7.0, ionic strength 0.25\,M, pMg~3.0 and 298.15\,K; the basis for each condition and its effect on the released energies are given in Supplementary Methods~S2.

\paragraph{Protonation.} eQuilibrator's Legendre transform requires a \emph{macroscopic} pKa ladder per compound: an ordered list in which consecutive species differ by one proton. The count of entries around the desired pH sets the proton count of the major microspecies, so the composition of the ladder changes the transformed energy directly. We generate the pKa ladders using ChemAxon Marvin~\cite{marvin} for use with eQuilibrator, but we're conscious that these data cannot be regenerated without a license, so we also include the pKas generated by an open-source alternative, MolGpKa~\cite{pan2021}, for every compound with a structure, and the measured ladders of Alberty~\cite{alberty2003}, versioned per structure source in an extended pKa dictionary.

\paragraph{Experimental anchors.} eQuilibrator and dGPredictor draw on the same set of experimental values derived from the NIST Thermodynamics of Enzyme-catalyzed Reactions (TECR) Database~\cite{goldberg2004}. The work was keyed to KEGG, and eQuilibrator uses MetaNetX mappings to integrate with other resources but there are inherent structural conflicts when we examine the mappings to ModelSEED identifiers. In order to draw direct comparisons between the results generated for any ModelSEED reactions we re-built the experimental anchors of TECR values, curating any conflicts to ensure a direct mapping between TECR values and ModelSEED reactions. In addition we evaluated openTECR, a communal re-curation of the same source, available at \url{https://github.com/opentecr}, to see if we can extend the set of anchors.
\paragraph{Experimental anchors.} eQuilibrator and dGPredictor both draw on the NIST Thermodynamics of Enzyme-catalyzed Reactions (TECR) Database~\cite{goldberg2004}, which is keyed to KEGG and reaches other namespaces through MetaNetX mappings that conflict structurally with ModelSEED identifiers. We therefore rebuilt the anchors, curating each conflict so that every TECR value maps directly onto a ModelSEED reaction, and evaluated openTECR (\url{https://github.com/opentecr}), a communal re-curation of the same source, to extend the set.

\paragraph{Uncertainty.} The three sources each adopt a different approach for estimating the uncertainty surrounding the predicted energy of reaction and as such the scale of uncertainty is not directly comparable (Figure~\ref{fig:thermo}C)) median values (kcal\,mol$^{-1}$) are 0.63 for eQuilibrator, 10.41 for group contribution, and 17.01 for dGPredictor. Comparing against the experimentally-anchored reactions, we find that group contribution overstates its error and eQuilibrator understates it.
\paragraph{Uncertainty.} The three sources each adopt a different approach for estimating the uncertainty surrounding the predicted energy of reaction, so the scale of uncertainty is not directly comparable between them (Figure~\ref{fig:thermo}C). We therefore calibrate each source's reported uncertainty against the experimental anchors rather than reading it at face value; the resulting medians and the direction of each source's bias are reported with the results.

\paragraph{Direction.} For historical reasons, we use the same heuristics for deriving reaction direction from the Group contribution approach~\cite{jankowski2008} unchanged from 2020, so that series stays comparable across releases. For eQuilibrator and dGPredictor we use one rule set built on the reversibility index from Noor \textit{et al.}~\cite{noor2012}: a reaction is called irreversible when $|\ln\Gamma|$ exceeds $\ln(1000)$ by more than one propagated standard deviation~\cite{beber2022}, $\Gamma$ being the fold-change of reactant concentration required to change reaction direction). However, we extend the work by Noor \textit{et al.} to require the uncertainty to clear the threshold as well. We report fewer directions than the published index would on the same energies, and we separate \emph{reversible} reactions from \emph{undetermined} cases. An exception we implement is for reactions whose uncertainty is small relative to the threshold and yet still straddles it: meaning the estimate lies within one standard deviation of the threshold, and these we report as reversible.
\paragraph{Direction.} For historical reasons, we use the same heuristics for deriving reaction direction from the Group contribution approach~\cite{jankowski2008} unchanged from 2020, so that series stays comparable across releases. For eQuilibrator and dGPredictor we use one rule set built on the reversibility index from Noor \textit{et al.}~\cite{noor2012}: a reaction is called irreversible when $|\ln\Gamma|$ exceeds $\ln(1000)$ by more than one propagated standard deviation~\cite{beber2022}, $\Gamma$ being the fold-change of reactant concentration required to change reaction direction). However, we extend the work by Noor \textit{et al.} to require the uncertainty to clear the threshold as well. We report fewer directions than the published index would on the same energies, and we separate \emph{reversible} reactions from \emph{undetermined} cases. An exception we implement is for reactions whose uncertainty is small relative to the threshold and yet still straddles it: meaning the estimate lies within one standard deviation of the threshold, and these we report as reversible.

\paragraph{Transport.} Compartments are \emph{relative} indices --- 0 inside, 1 outside --- not named compartments, so one reaction can be instantiated against a bacterial membrane, a mitochondrion or a plastid envelope when a template builds a model. Because the database is not committed to a cell type, transport reactions are scored from stoichiometry and energy alone: neither membrane potential nor pH gradient enters the estimate. The consequence is specific. A primary active transporter carries ATP hydrolysis in its stoichiometry, every predictor estimates that tightly, and the reaction earns a confident grade from chemistry that was never in question --- 96\% of gold-graded transport reactions are ATP-coupled and only 2.1\% translocate a proton. Such grades measure confidence in the coupled chemistry, not evidence that the transporter runs that way in a given organism; every reaction carries an \texttt{is\_transport} field so that this can be accounted for downstream. Reactions whose only change is of compartment carry no net chemistry and none is graded gold. dGPredictor commits on 1.4\% of transport reactions against 17.3\% elsewhere, so transport grades also rest on fewer independent sources.
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
\subsection{Structure curation}\label{sec:results-structure}

We built and ran our curation pipeline on all structures that would impact the set of roughly 9,000 reactions used in the ModelSEED and PlantSEED reconstruction templates. The conflict pipeline resolved every previously conflicting compound in three curation rounds. i) Instances where the conflict existed at the level of mobile hydrogens were resolved automatically using historical formula ii) instances where the structures were structurally different were resolved manually iii) instances where resolution was not straight-forward were entered in the mass-balance exclusion registry where we fix an arbitrarily chosen formula because no source formula was defensible.
We prioritised the curation pipeline on the compounds whose structures affect the ${\sim}9{,}000$ reactions carried by the ModelSEED and PlantSEED reconstruction templates, since an error in one of those propagates into every model built from them. The scope is a curation priority, not a statement about how much of the database is usable: the templates are discussed in Section~\ref{sec:results-mapping}. The conflict pipeline resolved every previously conflicting compound in three curation rounds. i) Instances where the conflict existed at the level of mobile hydrogens were resolved automatically using historical formula ii) instances where the structures were structurally different were resolved manually iii) instances where resolution was not straight-forward were entered in the mass-balance exclusion registry where we fix an arbitrarily chosen formula because no source formula was defensible.
Original file line number Diff line number Diff line change
Expand Up @@ -2,4 +2,6 @@
%% connectivity_ontology} sections.
\subsection{Atom mapping}\label{sec:results-mapping}

Atom mappings are published for 32,877 reactions, 59\% of the database: 25,058 clean, and 7,819 that required salvage (which are tagged). As the mappings can be used to span a metabolic reconstruction, reactions needing one repair are likely to carry others.
Atom mappings are published for 32,877 reactions, 59\% of the database: 25,058 clean, and 7,819 that required salvage (which are tagged). As the mappings can be used to span a metabolic reconstruction, reactions needing one repair are likely to carry others.

\paragraph{Reconstruction scope.} A template carries far fewer reactions than the database holds --- the v7.0 ModelSEED templates span 8,597 in union, about 9,000 with PlantSEED\_v3 roles --- but that gap does not mean the remainder is unusable. The database is a reference namespace across all kingdoms and includes biochemistry from secondary databases and published models that no single template targets. Promotion into a template is limited by mass and charge balance, which 58\% of reactions satisfy, and by association to an annotated functional role, which 41\% carry. Gene association is not the constraint: 2,258 of the 8,584 reactions in the Gram-negative template are typed as gapfilling and require no gene. 14,041 reactions now satisfy both balance and role annotation, 8,089 of them outside every current template.
2 changes: 1 addition & 1 deletion Papers/NAR_Update_2026/latex/sections/M14_discussion.tex
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
\section{Discussion}

Unlike the 2020 release, this update adopts a multi-source policy rather than promoting a single thermodynamic value, publishing four sources, each with their own uncertainty and estimated direction, because the sources may disagree in ways a single value conceals. Grading against experiment enables us to validate the approach: A researcher can take the highest-graded source and knows what the grade means. In being explicit about disagreement between sources, we expect this release to serve the widening range of communities that build on shared biochemistry: genome-scale reconstruction, flux analysis, pathway design and the training of predictive models.
Unlike the 2020 release, this update adopts a multi-source policy rather than promoting a single thermodynamic value, publishing four sources, each with their own uncertainty and estimated direction, because the sources may disagree in ways a single value conceals. Grading against experiment enables us to validate the approach: A researcher can take the highest-graded source and knows what the grade means. Two limits are worth naming. Transport reactions are scored without a membrane potential or a pH gradient, because both are properties of an organism and a condition rather than of a reference reaction; a compartment-aware scoring layer is the natural next release, and it needs the organism context a model supplies and a database cannot. And roughly half the database still carries no thermodynamic direction at all, bounded by the structures needed to compute one. In being explicit about disagreement between sources, we expect this release to serve the widening range of communities that build on shared biochemistry: genome-scale reconstruction, flux analysis, pathway design and the training of predictive models.