Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,4 +1,9 @@
\begin{figure*}[!t] \centerline{\includegraphics[width=\textwidth]{../figures/figure3_direction.pdf}} \caption{\textbf{Direction, agreement and evidence.} (\textbf{A}) Direction assigned by each source independently: \emph{eQ} eQuilibrator, \emph{dG} dGPredictor, \emph{GC} group contribution, and \emph{LLMs} the large language model ensemble, which supplies a direction and no energy; its calls are released with the database and take no part in the evidence grading. Bar length is the number of reactions the source covers, so the rows are directly comparable. (\textbf{B}) What each evidence grade rests on by the deciding source's own calibrated confidence, and (\textbf{C}) by what the other sources made of it. The two panels share a scale but not a palette, so each carries its own key; categories are ranked as in Supplementary Table~1. \emph{Neither way} is the case that table writes \texttt{---}: a cross-check ran and settled neither way, so it neither helped nor penalised.} \label{fig:direction} \end{figure*}
\begin{figure*}[!t]
\centerline{\includegraphics[width=\textwidth]{../figures/figure3_direction.pdf}}
\caption{
\textbf{Direction, agreement, and evidence.} (\textbf{A}) Direction assigned by each source independently: \emph{eQ} eQuilibrator, \emph{dG} dGPredictor, \emph{GC} group contribution, and \emph{LLMs} the large language model ensemble that predicts direction but not energy and are not considered in the evidence grading. Bar length is the number of reactions the source covers, facilitating comparison. (\textbf{B}) What each evidence grade rests on by the deciding source's own calibrated confidence, and (\textbf{C}) by what the other sources made of it. The two panels share a scale but not a palette, so each carries its own key; categories are ranked as in Supplementary Table~1. \emph{Neither way} is the case that table writes \texttt{---}: a cross-check ran and settled neither way, so it neither helped nor penalized.
} \label{fig:direction}
\end{figure*}

\subsection{Direction assignment and evidence grading}\label{sec:results-direction}

Expand All @@ -9,10 +14,20 @@ \subsection{Direction assignment and evidence grading}\label{sec:results-directi
%% Scripts/Thermodynamics/SourceGrading/grade_thermo_sources.py --cv
%% Anchors: 797 stereo-exact. The grade tiers are under active revision; the
%% rates below are for the scheme as it currently ships.
Each source now generates its own prediction for the direction each reaction proceeds, and the three sources differ more in what they will commit to than in what they say (Figure~\ref{fig:direction}A). eQuilibrator resolves 84.3\% of the reactions it covers, group contribution 42.9\% and dGPredictor 41.1\%. Across the database, 30,157 reactions receive a direction or a positive statement of reversibility from at least one source and 25,855 receive none.
Predictions of reaction directionality among the three sources differ more in their coverage than in their conclusions (Figure~\ref{fig:direction}A).
eQuilibrator, group contribution, and dGPredictor resolve 84.3\%, 42.9\%, and 41.1\% of the reactions they describe, respectively.
Across the database, 30,157 reactions receive a direction or a positive statement of reversibility from at least one source while 25,855 receive none.

\paragraph{eQuilibrator vs dGPredictor} dGPredictor is able to generate a predicted energy for more reactions ($\sim29,600$, against eQuilibrator's $\sim21,800$) and converts the least of that coverage into a usable reaction direction. For a typical dGPredictor reaction the error bar alone exceeds the decision threshold in the reversibility index, so 58.9\% of its reactions are returned undetermined against eQuilibrator's 15.7\%. Where both sources commit to a direction they agree 95.0\% of the time. dGPredictor is abstaining rather than contradicting, and it resolves 4,248 reactions eQuilibrator cannot, against 13,289 the other way.
\paragraph{eQuilibrator vs dGPredictor} dGPredictor predicts energy for more reactions ($\sim29,600$ versus $\sim21,800$ from eQuilibrator) and converts the least of that coverage into a usable reaction direction.
The error bar for most dGPredictor reactions exceed the decision threshold in the reversibility index, so 58.9\% of its reactions are returned undetermined against eQuilibrator's 15.7\%.
Consequently, dGPredictor resolves 4,248 reactions that eQuilibrator cannot while eQuilibrator resolves 13,289 reactions that dGPredictor cannot, although these sources agree in 95.0\% of reactions where both sources give a predictions.

\paragraph{Grading.} Of 33,099 graded reactions, 3,434 are gold, 18,388 silver and 11,277 bronze (Figure~\ref{fig:direction}D; see classification scheme in Supplementary Methods~S3). The tiers separate on what supports them: gold is 77\% self-certain and 68\% corroborated, while bronze is 99\% unconfident and split between reactions no second source could check (48\%) and reactions the other sources contradict (35\%). Because a measured reaction is always graded gold, the grades must be recomputed with openTECR withheld before they can be tested against it. Re-graded on the predictors alone, the 806 anchored reactions split into 357 gold, 437 silver and 12 bronze; the gold subset falls within 2\,kcal\,mol$^{-1}$ of measurement 96.9\% of the time, against 91.3\% for silver. Both tiers therefore track measurement closely enough that we recommend either for inclusion in a metabolic model, and reserve bronze for review rather than use.
\paragraph{Grading.} Of 33,099 graded reactions, 3,434 are gold, 18,388 silver and 11,277 bronze (Figure~\ref{fig:direction}D; see classification scheme in Supplementary Methods~S3).
The tiers separate on what supports them: gold is 77\% self-certain and 68\% corroborated, while bronze is 99\% unconfident, consisting of reactions that other sources either could not check (48\%) or contradicted (35\%).
Because a measured reaction is always graded gold, the measured reactions from openTECR must be withheld to preserve testing against it.
Re-graded on the predictors alone, the 806 anchored reactions become 357 gold, 437 silver and 12 bronze.
The gold subset falls within 2\,kcal\,mol$^{-1}$ of the measured value 96.9\% of the time against 91.3\% for silver, which both track measurement closely enough that we recommend including them in GEMs.
In contrast, we recommend bronze-rated directionalities for investigators to review but not use in modeling.

\paragraph{Recommendation.} Evidently where the sources agree on reaction direction, this becomes our recommendation, but because the sources can disagree, we prioritize the source from which we take the recommended reaction direction, so the priority order is eQuilibrator, dGPredictor, group contribution. eQuilibrator supplies 21,218 recommended reaction directions, dGPredictor 4,248 and group contribution 4,691; these numbers include the ones where there is agreement also. We record the source used to make the recommendation, so researchers may apply a different precedence.
\paragraph{Recommendation.} We recommend directionality when all sources agree, but preferentially recommend the predictions from eQuilibrator, dGPredictor, and group contribution in that order.
From this recommendation criteria, eQuilibrator supplies 21,218 recommended reaction directions, dGPredictor 4,248, and group contribution 4,691, with the source(s) prevenance for each recommendation being recorded.