Reviewer 2.2: give the LLM ensemble a measured role, and disclose its forward bias - #298
Open
freiburgermsu wants to merge 2 commits into
Open
freiburgermsu wants to merge 2 commits into
freiburgermsu wants to merge 2 commits into
Conversation
… bias
The reviewer could not see what the ensemble is for:
"Except the numbers in Figure 3, it is not apparent how LLMs play roles in
reconstruction and other analyses."
That is a fair reading of a section that argued against itself. The motivation
opened on pathway-level driving force -- a reaction "may be driven in a
direction on a level that supersedes our evaluation" -- cited three papers on
exactly that, and then introduced an ensemble that sees only a reaction's name
and stoichiometry and therefore has no pathway context at all. It then withdrew
twice: the calls take no part in the grading (M06), and their confidence "is
not appropriate to use as a downstream filter" (S04). A method motivated by
something it does not do, and disclaimed for everything else.
WHAT THE SECTION NOW CLAIMS. Coverage, which is measurable and real: structural
completeness bounds thermodynamic assignment, leaving 25,855 reactions with no
direction from any predictor, and the ensemble directs 22,902 of them -- roughly
doubling the fraction of the database carrying a direction. That is the role.
WHAT IT NOW DISCLOSES. Measuring the ensemble against measurement rather than
against another prediction changes the picture, so the numbers are stated
rather than left to be found:
- against the measured energies in opentecr_comparison.csv, the ensemble
recovers the direction on 115 of 158 reactions where both commit (72.8%),
against 81 of 82 for eQuilibrator (98.8%).
- agreement with eQuilibrator across the database is 94.7% of 8,085
reactions, but the chance baseline is 88.3% and Cohen's kappa is 0.55.
The headline agreement is mostly a shared prior, not concordance.
- 85.3% of its calls are the direction the equation is written in, and the
errors are entirely one-sided: correct on all 72 reactions measured
forward, and on 43 of the 86 measured reverse.
The ensemble has learned that reaction equations are written in the
physiological direction. That is a genuine prior and it is why the coverage is
useful, but it is not thermodynamics, and it independently justifies the
decision the paper had already made to keep these calls out of the grading. A
reverse call -- the model contradicting its own prior -- is the more
informative of the two, and the supplement now says so.
CITATIONS. mavrovouniotis1993, xu2008 and noor2014 supported the
pathway-driving-force motivation being removed, and each is cited nowhere else.
Rather than drop them, they are kept on one sentence that states the point
honestly as a limitation of BOTH approaches: pathway context can drive a
reaction against its standard-state energy, and neither a thermodynamic
estimate nor an ensemble reading a single equation models it. This also closes
the loop with the transport limitation in reviewer 1's comment 2.
Detail lives in Supplementary Methods S4, which has no page constraint; the
main text carries the role and the caveat in four sentences.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e paper already owns
Three things found on review of this PR.
- The main text and the supplement disagreed on the ensemble's accuracy:
M06 said 76% and S04 said 72.8%, because they were computed at different
margins (5 kJ and 11.7 kJ) and the 11.7 was never justified. Both now use
tau = 2.0 kcal/mol (8.37 kJ), the tolerance the evidence grading already
defines in S03, and S04 states the sensitivity: across a bare sign test up
to RT ln 1000 the ensemble ranges 65-77% with kappa 0.33-0.50, eQuilibrator
stays above 96%, and the errors are one-sided at every cut. At tau: 165 of
215 (76.7%, kappa 0.50) against 54 of 55 (98.2%, kappa 0.92); correct on
119 of 120 measured forward, 46 of 95 measured reverse.
- 25,855 undirected reactions was the all-records count. Live: 23,729, of
which the ensemble directs 22,902. The 22,902 is unchanged because the
ensemble was only ever run on non-obsolete reactions (S04).
- eQuilibrator's "81 of 82" against measurement counted obsolete duplicate
anchors; on the live population and at tau it is 54 of 55.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Answers reviewer 2's comment 2. Text only —
M06(Methods) andS04(Supplement).The reviewer found a real gap
The section argued against itself. Its motivation opened on pathway-level driving force —
— cited three papers on exactly that, and then introduced an ensemble that sees "its name and stoichiometry alone" and therefore has no pathway context whatsoever. It then withdrew twice: the calls "do not [enter] our grading of the thermodynamic evidence" (M06), and their confidence "does not appear to distinguish correct calls … and is not appropriate to use as a downstream filter" (S04).
So the reader is handed a method motivated by something it does not do, and disclaimed for everything else. That is what the reviewer is reacting to, and the fix is to state a role the data supports.
The role: coverage
Thermodynamic assignment is bounded by structural completeness; the ensemble is not. It roughly doubles the fraction of the database carrying a direction. That is a real, quantified contribution and it is now what the section claims.
The disclosure: it is not reading thermodynamics
Checking the ensemble against measurement rather than against another prediction changes the picture materially.
Biochemistry/Thermodynamics/SourceGrading/opentecr_comparison.csvships a measuredopentecr_dG_kJfor 1,365 reactions, so this is directly testable:The 94.7% agreement with eQuilibrator across the whole database (8,085 co-commitments) reads far better than this — but its chance baseline is 88.3%, κ = 0.55. The headline agreement is mostly a shared prior.
The mechanism is unusually clean. 85.3% of the ensemble's calls are the direction the equation is written in, and the errors are entirely one-sided:
><><Perfect when the answer is "forward", a coin flip when it is not. The ensemble has learned that reaction equations are conventionally written in the physiological direction. That is a genuine and useful prior — it is why the coverage is worth having — but it is not thermodynamics.
This is the main reason to state it ourselves. It is recoverable from shipped data in an afternoon, so it is much better disclosed than discovered, and it independently justifies the decision the paper had already made to keep these calls out of the evidence grading. S04 also now notes the practical consequence: a reverse call is the model contradicting its own prior, and is the more informative of the two — 2,350 reactions a curator could review in a tractable pass.
The three orphaned citations
mavrovouniotis1993,xu2008andnoor2014supported the pathway-driving-force motivation being removed, and each is cited nowhere else in the manuscript (verified). Rather than drop them, they are kept on a single sentence that makes the point honestly — as a limitation of both approaches:This is true, keeps three relevant references, and closes the loop with the transport limitation in reviewer 1's comment 2 (PR #297).
Split of detail
Main text carries the role and the caveat in four sentences; the numbers, the confusion matrix and the κ discussion go to
S04, which has no page constraint.Verification
latexmk -pdfclean for bothmain.texandsupplementary.tex; zero undefined references or citations.Papers/NAR_Update_2026/analysis/review_transport_and_llm.py(added in the figures PR of this series), on the manuscript's own population — all records, obsolete included.Note on page count
Body text ends on page 6; one reference entry spills to page 7, where the baseline fit in six. PR #297 in this series is in the same position. The combined count needs re-checking once both merge, along with
\lastpage{6}— flagged rather than tuned per-branch, since neither branch can see the other's contribution.🤖 Generated with Claude Code
Review pass (2026-09-23)
Three defects found and fixed. (1) The main text and supplement disagreed on the ensemble's accuracy — M06 said 76%, S04 said 72.8% — because they used different margins, and S04's 11.7 kJ was never justified. Both now use τ = 2 kcal/mol, the tolerance S03 already defines, with a sensitivity statement: from a bare sign test up to RT·ln 1000 the ensemble ranges 65–77% (κ 0.33–0.50) and eQuilibrator stays above 96%; errors are one-sided at every cut. At τ: 165/215 (76.7%, κ 0.50) vs 54/55 (98.2%, κ 0.92); correct on 119 of 120 measured forward, 46 of 95 measured reverse. (2) 25,855 undirected was all-records → 23,729 live; the 22,902 is unchanged because the ensemble only ran on non-obsolete reactions. (3) eQuilibrator's 81/82 counted obsolete duplicate anchors → 54/55 at τ.
Re-verified unchanged: 94.7% agreement / 88.3% chance / κ 0.55, the 85.3 / 5.1 / 7.8 call split.