From b5bd459debb81e8fe091ff7bef3842a70baa7377 Mon Sep 17 00:00:00 2001 From: Andrew Freiburger Date: Tue, 22 Sep 2026 20:11:39 -0500 Subject: [PATCH 1/2] Reviewer 2.2: give the LLM ensemble a measured role, and disclose its bias The reviewer could not see what the ensemble is for: "Except the numbers in Figure 3, it is not apparent how LLMs play roles in reconstruction and other analyses." That is a fair reading of a section that argued against itself. The motivation opened on pathway-level driving force -- a reaction "may be driven in a direction on a level that supersedes our evaluation" -- cited three papers on exactly that, and then introduced an ensemble that sees only a reaction's name and stoichiometry and therefore has no pathway context at all. It then withdrew twice: the calls take no part in the grading (M06), and their confidence "is not appropriate to use as a downstream filter" (S04). A method motivated by something it does not do, and disclaimed for everything else. WHAT THE SECTION NOW CLAIMS. Coverage, which is measurable and real: structural completeness bounds thermodynamic assignment, leaving 25,855 reactions with no direction from any predictor, and the ensemble directs 22,902 of them -- roughly doubling the fraction of the database carrying a direction. That is the role. WHAT IT NOW DISCLOSES. Measuring the ensemble against measurement rather than against another prediction changes the picture, so the numbers are stated rather than left to be found: - against the measured energies in opentecr_comparison.csv, the ensemble recovers the direction on 115 of 158 reactions where both commit (72.8%), against 81 of 82 for eQuilibrator (98.8%). - agreement with eQuilibrator across the database is 94.7% of 8,085 reactions, but the chance baseline is 88.3% and Cohen's kappa is 0.55. The headline agreement is mostly a shared prior, not concordance. - 85.3% of its calls are the direction the equation is written in, and the errors are entirely one-sided: correct on all 72 reactions measured forward, and on 43 of the 86 measured reverse. The ensemble has learned that reaction equations are written in the physiological direction. That is a genuine prior and it is why the coverage is useful, but it is not thermodynamics, and it independently justifies the decision the paper had already made to keep these calls out of the grading. A reverse call -- the model contradicting its own prior -- is the more informative of the two, and the supplement now says so. CITATIONS. mavrovouniotis1993, xu2008 and noor2014 supported the pathway-driving-force motivation being removed, and each is cited nowhere else. Rather than drop them, they are kept on one sentence that states the point honestly as a limitation of BOTH approaches: pathway context can drive a reaction against its standard-state energy, and neither a thermodynamic estimate nor an ensemble reading a single equation models it. This also closes the loop with the transport limitation in reviewer 1's comment 2. Detail lives in Supplementary Methods S4, which has no page constraint; the main text carries the role and the caveat in four sentences. Co-Authored-By: Claude Opus 5 (1M context) --- .../latex/sections/M06_methods_council_direction.tex | 4 +++- .../latex/supplement/S04_direction_from_chemistry.tex | 6 +++++- 2 files changed, 8 insertions(+), 2 deletions(-) diff --git a/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex b/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex index e77300b2..9f57723f 100644 --- a/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex +++ b/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex @@ -15,4 +15,6 @@ \subsection{Grading reaction direction}\label{sec:methods-grading} We apply a classification approach with which we grade the predicted reaction direction based on the available evidence. Very often sources and heuristics disagree, as is the case here, and we qualify what the reaction direction would be, and how reliable the evidence is using several tiers for ease of interpretation: gold, silver, or bronze. Our evaluation splits two ways, on the confidence of a source's own claim, fitted as a probability against the experimental anchors, and what the other sources make of it, by a weighted comparison on the same scale. Grades, self-assessments, and cross-source verdicts are stored in the biochemistry database, and can be viewed in the UI. The process of grading reactions is described fully in Supplementary Methods~S3. -\subsection{Ensemble LLMs predictions}\label{sec:llm-grading} Thermodynamic assignment is bounded by estimate coverage: energies cannot be computed for an incomplete reaction in terms of structures. Furthermore, every reaction is a member of a pathway within a cell, and the overarching drive of the pathway, dynamically changing metabolite concentrations as downstream enzymes process reagents, means a reaction may be driven in a direction on a level that supersedes our evaluation~\cite{mavrovouniotis1993,xu2008,noor2014}. We run a complementary approach using an ensemble, a ``council'' of large language models (LLMs) to interpret the reaction based on its name and stoichiometry alone. Three LLMs independently make their predictions, a fourth audits, and a fifth adjudicates. We provide the prompts in the repository, and describe the process and its limits in Supplementary Methods~S4. We integrate these predictions transparently in our database, but do not include them in our grading of the thermodynamic evidence. \ No newline at end of file +\subsection{Ensemble LLMs predictions}\label{sec:llm-grading} Thermodynamic assignment is bounded by structural coverage: an energy cannot be computed for a reaction whose participants lack complete structures, however well characterised the reaction is. That bound leaves 25,855 reactions with no direction from any predictor. We therefore run a complementary approach that does not depend on structures at all, using an ensemble --- a ``council'' --- of large language models (LLMs) to interpret a reaction from its name and stoichiometry. Three LLMs independently make their predictions, a fourth audits, and a fifth adjudicates; the prompts are in the repository and the process and its limits are described in Supplementary Methods~S4. + +The ensemble directs 22,902 of the reactions no thermodynamic source will commit on, roughly doubling the fraction of the database that carries a direction. Its accuracy is not that of an energy calculation: scored against measurement it recovers the direction 76\% of the time against 99\% for eQuilibrator, and 85\% of its calls are simply the direction the equation is written in, so its errors fall almost entirely on reactions that run in reverse (Supplementary Methods~S4). We therefore release these calls as a separate, clearly-labelled layer, excluded from the evidence grading: they are for reactions a reconstruction would otherwise leave reversible by default, and should be reviewed before use rather than treated as evidence. Neither approach models the pathway context that can drive a reaction against its own standard-state energy as metabolite concentrations shift downstream~\cite{mavrovouniotis1993,xu2008,noor2014}; that remains a property of a model rather than of a reference database. \ No newline at end of file diff --git a/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex b/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex index c75ddfba..04f8853d 100644 --- a/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex +++ b/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex @@ -6,4 +6,8 @@ \subsection{Ensemble} Eleven candidate LLMs were scored across the role of propo \subsection{Reactions and status} We filtered out reactions for which the prediction would not have been appropriate, this includes reactions flagged as obsolete, symmetrical transport reactions, and pseudo-reactions that acted as lumped reactions, aggregating many species and for which there is no direction to infer, leaving $\sim46,000$ reactions for which direction is assigned. We store the output of the ensemble in our repository for review, including any objections raised by the auditor. -\subsection{Summary of results} As the set of reactions covers almost all those for which we generated thermodynamics data, we're able to draw a direct comparison. The ensemble commits to a direction far more readily than the thermodynamic rules do, abstaining on $\sim2\%$ of reactions. The confidence that the ensemble returns does not appear to distinguish correct calls when compared to the thermodynamics, and is not appropriate to use as a downstream filter when integrating reactions. Nevertheless, with the far wider coverage, researchers may find the results useful. \ No newline at end of file +\subsection{Summary of results} As the set of reactions covers almost all those for which we generated thermodynamics data, we are able to draw a direct comparison. The ensemble commits to a direction far more readily than the thermodynamic rules do, abstaining on $\sim2\%$ of reactions, and it supplies a direction for 22,902 of the 25,855 reactions no thermodynamic source will commit on. That coverage is the reason to publish it. + +\paragraph{Accuracy against measurement.} The confidence the ensemble returns does not distinguish its correct calls from its incorrect ones and is not appropriate as a downstream filter. The direction itself requires the same care. Scoring against the measured energies in \texttt{opentecr\_comparison.csv}, and taking a reaction to be measured-irreversible where $|\Delta_{\mathrm{r}}G'^{\circ}|$ exceeds 11.7\,kJ\,mol$^{-1}$, the ensemble agrees with measurement on 115 of the 158 reactions where both commit (72.8\%), against 81 of 82 for eQuilibrator (98.8\%). Agreement with eQuilibrator across the whole database is 94.7\% of 8,085 reactions, but that figure should not be read as accuracy: the chance baseline is 88.3\% and Cohen's $\kappa$ is 0.55, because both sources overwhelmingly report the forward direction. + +The reason is visible in the call distribution. 85.3\% of the ensemble's calls are the direction the equation is written in ($>$), against 5.1\% reverse and 7.8\% reversible, and the errors are entirely one-sided: of the 158 reactions above, it is correct on all 72 that are measured forward and on 43 of the 86 measured reverse. The ensemble has learned the convention that reaction equations are written in the physiological direction, which is a real and useful prior, but it is not reading thermodynamics. This is why the calls are released as their own layer and take no part in the evidence grading, and why a reverse call --- the case where the model contradicts its own prior --- is the more informative of the two. \ No newline at end of file From 9713d37f7ea9d54ecc6c07eef8919af0697eb8b1 Mon Sep 17 00:00:00 2001 From: Andrew Freiburger Date: Wed, 23 Sep 2026 09:41:59 -0500 Subject: [PATCH 2/2] LLM section: live population, one accuracy figure, and a threshold the paper already owns Three things found on review of this PR. - The main text and the supplement disagreed on the ensemble's accuracy: M06 said 76% and S04 said 72.8%, because they were computed at different margins (5 kJ and 11.7 kJ) and the 11.7 was never justified. Both now use tau = 2.0 kcal/mol (8.37 kJ), the tolerance the evidence grading already defines in S03, and S04 states the sensitivity: across a bare sign test up to RT ln 1000 the ensemble ranges 65-77% with kappa 0.33-0.50, eQuilibrator stays above 96%, and the errors are one-sided at every cut. At tau: 165 of 215 (76.7%, kappa 0.50) against 54 of 55 (98.2%, kappa 0.92); correct on 119 of 120 measured forward, 46 of 95 measured reverse. - 25,855 undirected reactions was the all-records count. Live: 23,729, of which the ensemble directs 22,902. The 22,902 is unchanged because the ensemble was only ever run on non-obsolete reactions (S04). - eQuilibrator's "81 of 82" against measurement counted obsolete duplicate anchors; on the live population and at tau it is 54 of 55. Co-Authored-By: Claude Fable 5.1 --- .../latex/sections/M06_methods_council_direction.tex | 4 ++-- .../latex/supplement/S04_direction_from_chemistry.tex | 6 +++--- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex b/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex index 9f57723f..2b262587 100644 --- a/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex +++ b/Papers/NAR_Update_2026/latex/sections/M06_methods_council_direction.tex @@ -15,6 +15,6 @@ \subsection{Grading reaction direction}\label{sec:methods-grading} We apply a classification approach with which we grade the predicted reaction direction based on the available evidence. Very often sources and heuristics disagree, as is the case here, and we qualify what the reaction direction would be, and how reliable the evidence is using several tiers for ease of interpretation: gold, silver, or bronze. Our evaluation splits two ways, on the confidence of a source's own claim, fitted as a probability against the experimental anchors, and what the other sources make of it, by a weighted comparison on the same scale. Grades, self-assessments, and cross-source verdicts are stored in the biochemistry database, and can be viewed in the UI. The process of grading reactions is described fully in Supplementary Methods~S3. -\subsection{Ensemble LLMs predictions}\label{sec:llm-grading} Thermodynamic assignment is bounded by structural coverage: an energy cannot be computed for a reaction whose participants lack complete structures, however well characterised the reaction is. That bound leaves 25,855 reactions with no direction from any predictor. We therefore run a complementary approach that does not depend on structures at all, using an ensemble --- a ``council'' --- of large language models (LLMs) to interpret a reaction from its name and stoichiometry. Three LLMs independently make their predictions, a fourth audits, and a fifth adjudicates; the prompts are in the repository and the process and its limits are described in Supplementary Methods~S4. +\subsection{Ensemble LLMs predictions}\label{sec:llm-grading} Thermodynamic assignment is bounded by structural coverage: an energy cannot be computed for a reaction whose participants lack complete structures, however well characterised the reaction is. That bound leaves 23,729 reactions with no direction from any predictor. We therefore run a complementary approach that does not depend on structures at all, using an ensemble --- a ``council'' --- of large language models (LLMs) to interpret a reaction from its name and stoichiometry. Three LLMs independently make their predictions, a fourth audits, and a fifth adjudicates; the prompts are in the repository and the process and its limits are described in Supplementary Methods~S4. -The ensemble directs 22,902 of the reactions no thermodynamic source will commit on, roughly doubling the fraction of the database that carries a direction. Its accuracy is not that of an energy calculation: scored against measurement it recovers the direction 76\% of the time against 99\% for eQuilibrator, and 85\% of its calls are simply the direction the equation is written in, so its errors fall almost entirely on reactions that run in reverse (Supplementary Methods~S4). We therefore release these calls as a separate, clearly-labelled layer, excluded from the evidence grading: they are for reactions a reconstruction would otherwise leave reversible by default, and should be reviewed before use rather than treated as evidence. Neither approach models the pathway context that can drive a reaction against its own standard-state energy as metabolite concentrations shift downstream~\cite{mavrovouniotis1993,xu2008,noor2014}; that remains a property of a model rather than of a reference database. \ No newline at end of file +The ensemble directs 22,902 of the reactions no thermodynamic source will commit on, roughly doubling the fraction of the database that carries a direction. Its accuracy is not that of an energy calculation: scored against measurement it recovers the direction 77\% of the time against 98\% for eQuilibrator, and 85\% of its calls are simply the direction the equation is written in, so its errors fall almost entirely on reactions that run in reverse (Supplementary Methods~S4). We therefore release these calls as a separate, clearly-labelled layer, excluded from the evidence grading: they are for reactions a reconstruction would otherwise leave reversible by default, and should be reviewed before use rather than treated as evidence. Neither approach models the pathway context that can drive a reaction against its own standard-state energy as metabolite concentrations shift downstream~\cite{mavrovouniotis1993,xu2008,noor2014}; that remains a property of a model rather than of a reference database. \ No newline at end of file diff --git a/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex b/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex index 04f8853d..93f0c844 100644 --- a/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex +++ b/Papers/NAR_Update_2026/latex/supplement/S04_direction_from_chemistry.tex @@ -6,8 +6,8 @@ \subsection{Ensemble} Eleven candidate LLMs were scored across the role of propo \subsection{Reactions and status} We filtered out reactions for which the prediction would not have been appropriate, this includes reactions flagged as obsolete, symmetrical transport reactions, and pseudo-reactions that acted as lumped reactions, aggregating many species and for which there is no direction to infer, leaving $\sim46,000$ reactions for which direction is assigned. We store the output of the ensemble in our repository for review, including any objections raised by the auditor. -\subsection{Summary of results} As the set of reactions covers almost all those for which we generated thermodynamics data, we are able to draw a direct comparison. The ensemble commits to a direction far more readily than the thermodynamic rules do, abstaining on $\sim2\%$ of reactions, and it supplies a direction for 22,902 of the 25,855 reactions no thermodynamic source will commit on. That coverage is the reason to publish it. +\subsection{Summary of results} As the set of reactions covers almost all those for which we generated thermodynamics data, we are able to draw a direct comparison. The ensemble commits to a direction far more readily than the thermodynamic rules do, abstaining on $\sim2\%$ of reactions, and it supplies a direction for 22,902 of the 23,729 reactions no thermodynamic source will commit on. That coverage is the reason to publish it. -\paragraph{Accuracy against measurement.} The confidence the ensemble returns does not distinguish its correct calls from its incorrect ones and is not appropriate as a downstream filter. The direction itself requires the same care. Scoring against the measured energies in \texttt{opentecr\_comparison.csv}, and taking a reaction to be measured-irreversible where $|\Delta_{\mathrm{r}}G'^{\circ}|$ exceeds 11.7\,kJ\,mol$^{-1}$, the ensemble agrees with measurement on 115 of the 158 reactions where both commit (72.8\%), against 81 of 82 for eQuilibrator (98.8\%). Agreement with eQuilibrator across the whole database is 94.7\% of 8,085 reactions, but that figure should not be read as accuracy: the chance baseline is 88.3\% and Cohen's $\kappa$ is 0.55, because both sources overwhelmingly report the forward direction. +\paragraph{Accuracy against measurement.} The confidence the ensemble returns does not distinguish its correct calls from its incorrect ones and is not appropriate as a downstream filter. The direction itself requires the same care. We score it against the measured energies in \texttt{opentecr\_comparison.csv}, taking a reaction to be measured-irreversible where $|\Delta_{\mathrm{r}}G'^{\circ}|$ exceeds $\tau = 2.0$\,kcal\,mol$^{-1}$, the same tolerance the evidence grading uses (Supplementary Methods~S3). On that basis the ensemble agrees with measurement on 165 of the 215 reactions where both commit (76.7\%, Cohen's $\kappa = 0.50$), against 54 of 55 for eQuilibrator (98.2\%, $\kappa = 0.92$). The result is not sensitive to the threshold: between a bare sign test and $RT\ln 1000$ the ensemble's agreement ranges from 65\% to 77\% and its $\kappa$ from 0.33 to 0.50, while eQuilibrator's stays above 96\%. Agreement with eQuilibrator across the whole database is 94.7\% of 8,085 reactions, but that figure should not be read as accuracy: the chance baseline is 88.3\% and $\kappa$ is 0.55, because both sources overwhelmingly report the forward direction. -The reason is visible in the call distribution. 85.3\% of the ensemble's calls are the direction the equation is written in ($>$), against 5.1\% reverse and 7.8\% reversible, and the errors are entirely one-sided: of the 158 reactions above, it is correct on all 72 that are measured forward and on 43 of the 86 measured reverse. The ensemble has learned the convention that reaction equations are written in the physiological direction, which is a real and useful prior, but it is not reading thermodynamics. This is why the calls are released as their own layer and take no part in the evidence grading, and why a reverse call --- the case where the model contradicts its own prior --- is the more informative of the two. \ No newline at end of file +The reason is visible in the call distribution. 85.3\% of the ensemble's calls are the direction the equation is written in ($>$), against 5.1\% reverse and 7.8\% reversible, and the errors are almost entirely one-sided: of the 215 reactions above, it is correct on 119 of the 120 that are measured forward and on 46 of the 95 measured reverse. The ensemble has learned the convention that reaction equations are written in the physiological direction, which is a real and useful prior, but it is not reading thermodynamics. This is why the calls are released as their own layer and take no part in the evidence grading, and why a reverse call --- the case where the model contradicts its own prior --- is the more informative of the two.