Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -15,4 +15,6 @@ \subsection{Grading reaction direction}\label{sec:methods-grading}

We apply a classification approach with which we grade the predicted reaction direction based on the available evidence. Very often sources and heuristics disagree, as is the case here, and we qualify what the reaction direction would be, and how reliable the evidence is using several tiers for ease of interpretation: gold, silver, or bronze. Our evaluation splits two ways, on the confidence of a source's own claim, fitted as a probability against the experimental anchors, and what the other sources make of it, by a weighted comparison on the same scale. Grades, self-assessments, and cross-source verdicts are stored in the biochemistry database, and can be viewed in the UI. The process of grading reactions is described fully in Supplementary Methods~S3.

\subsection{Ensemble LLMs predictions}\label{sec:llm-grading} Thermodynamic assignment is bounded by estimate coverage: energies cannot be computed for an incomplete reaction in terms of structures. Furthermore, every reaction is a member of a pathway within a cell, and the overarching drive of the pathway, dynamically changing metabolite concentrations as downstream enzymes process reagents, means a reaction may be driven in a direction on a level that supersedes our evaluation~\cite{mavrovouniotis1993,xu2008,noor2014}. We run a complementary approach using an ensemble, a ``council'' of large language models (LLMs) to interpret the reaction based on its name and stoichiometry alone. Three LLMs independently make their predictions, a fourth audits, and a fifth adjudicates. We provide the prompts in the repository, and describe the process and its limits in Supplementary Methods~S4. We integrate these predictions transparently in our database, but do not include them in our grading of the thermodynamic evidence.
\subsection{Ensemble LLMs predictions}\label{sec:llm-grading} Thermodynamic assignment is bounded by structural coverage: an energy cannot be computed for a reaction whose participants lack complete structures, however well characterised the reaction is. That bound leaves 23,726 reactions with no direction from any predictor. We therefore run a complementary approach that does not depend on structures at all, using an ensemble --- a ``council'' --- of large language models (LLMs) to interpret a reaction from its name and stoichiometry. Three LLMs independently make their predictions, a fourth audits, and a fifth adjudicates; the prompts are in the repository and the process and its limits are described in Supplementary Methods~S4.

The ensemble directs 22,899 of the reactions no thermodynamic source will commit on, roughly doubling the fraction of the database that carries a direction. Its accuracy is not that of an energy calculation: scored against measurement it recovers the direction 77\% of the time against 98\% for eQuilibrator, and 85\% of its calls are simply the direction the equation is written in, so its errors fall almost entirely on reactions that run in reverse (Supplementary Methods~S4). We therefore release these calls as a separate, clearly-labelled layer, excluded from the evidence grading: they are for reactions a reconstruction would otherwise leave reversible by default, and should be reviewed before use rather than treated as evidence. Neither approach models the pathway context that can drive a reaction against its own standard-state energy as metabolite concentrations shift downstream~\cite{mavrovouniotis1993,xu2008,noor2014}; that remains a property of a model rather than of a reference database.
Original file line number Diff line number Diff line change
Expand Up @@ -6,4 +6,8 @@ \subsection{Ensemble} Eleven candidate LLMs were scored across the role of propo

\subsection{Reactions and status} We filtered out reactions for which the prediction would not have been appropriate, this includes reactions flagged as obsolete, symmetrical transport reactions, and pseudo-reactions that acted as lumped reactions, aggregating many species and for which there is no direction to infer, leaving $\sim46,000$ reactions for which direction is assigned. We store the output of the ensemble in our repository for review, including any objections raised by the auditor.

\subsection{Summary of results} As the set of reactions covers almost all those for which we generated thermodynamics data, we're able to draw a direct comparison. The ensemble commits to a direction far more readily than the thermodynamic rules do, abstaining on $\sim2\%$ of reactions. The confidence that the ensemble returns does not appear to distinguish correct calls when compared to the thermodynamics, and is not appropriate to use as a downstream filter when integrating reactions. Nevertheless, with the far wider coverage, researchers may find the results useful.
\subsection{Summary of results} As the set of reactions covers almost all those for which we generated thermodynamics data, we are able to draw a direct comparison. The ensemble commits to a direction far more readily than the thermodynamic rules do, abstaining on $\sim2\%$ of reactions, and it supplies a direction for 22,899 of the 23,726 reactions no thermodynamic source will commit on. That coverage is the reason to publish it.

\paragraph{Accuracy against measurement.} The confidence the ensemble returns does not distinguish its correct calls from its incorrect ones and is not appropriate as a downstream filter. The direction itself requires the same care. We score it against the measured energies in \texttt{opentecr\_comparison.csv}, taking a reaction to be measured-irreversible where $|\Delta_{\mathrm{r}}G'^{\circ}|$ exceeds $\tau = 2.0$\,kcal\,mol$^{-1}$, the same tolerance the evidence grading uses (Supplementary Methods~S3). On that basis the ensemble agrees with measurement on 165 of the 215 reactions where both commit (76.7\%, Cohen's $\kappa = 0.50$), against 54 of 55 for eQuilibrator (98.2\%, $\kappa = 0.92$). The result is not sensitive to the threshold: between a bare sign test and $RT\ln 1000$ the ensemble's agreement ranges from 65\% to 77\% and its $\kappa$ from 0.33 to 0.50, while eQuilibrator's stays above 96\%. Agreement with eQuilibrator across the whole database is 94.7\% of 8,084 reactions, but that figure should not be read as accuracy: the chance baseline is 88.3\% and $\kappa$ is 0.55, because both sources overwhelmingly report the forward direction.

The reason is visible in the call distribution. 85.3\% of the ensemble's calls are the direction the equation is written in ($>$), against 5.2\% reverse and 7.8\% reversible, and the errors are almost entirely one-sided: of the 215 reactions above, it is correct on 119 of the 120 that are measured forward and on 46 of the 95 measured reverse. The ensemble has learned the convention that reaction equations are written in the physiological direction, which is a real and useful prior, but it is not reading thermodynamics. This is why the calls are released as their own layer and take no part in the evidence grading, and why a reverse call --- the case where the model contradicts its own prior --- is the more informative of the two.