From b67c5745bb817a664a0eab81c611e630cfd63ced Mon Sep 17 00:00:00 2001 From: Andrew Freiburger Date: Fri, 2 Oct 2026 08:34:04 -0500 Subject: [PATCH] NAR 2026: sync M11_results_thermodynamics.tex with Overleaf edits Co-Authored-By: Claude Opus 5.5 (1M context) --- .../latex/sections/M11_results_thermodynamics.tex | 15 ++++++++++++--- 1 file changed, 12 insertions(+), 3 deletions(-) diff --git a/Papers/NAR_Update_2026/latex/sections/M11_results_thermodynamics.tex b/Papers/NAR_Update_2026/latex/sections/M11_results_thermodynamics.tex index 087dcc22..feaa0ee3 100644 --- a/Papers/NAR_Update_2026/latex/sections/M11_results_thermodynamics.tex +++ b/Papers/NAR_Update_2026/latex/sections/M11_results_thermodynamics.tex @@ -1,4 +1,8 @@ -\begin{figure*}[!t] \centerline{\includegraphics[width=\textwidth]{../figures/figure2_thermodynamics.pdf}} \caption{\textbf{Thermodynamic sources.} (\textbf{A}) Provenance of every pK$_\mathrm{a}$ ladder shipped in the database, across all 45,708 compounds and 56,002 reactions. ChemAxon Marvin supplies 79.4\% of compounds and 81.6\% of reactions. \emph{Hatching} marks the share reached through a second route: the calculator refuses polymers and organometallics as query molecules, so those are computed from SMILES instead. Marvin reported no dissociable proton between pH $-2$ and $16$ for 1,242 compounds. The remainder carry no structure to compute from, which is a curation gap rather than a protonation one. (\textbf{B}) Reactions that carry an energy computed from the source heuristic. (\textbf{C}) Reported uncertainty in kcal\,mol$^{-1}$, one axis per source because the three differ by more than an order of magnitude. The shaded band spans the 5th to 95th percentile, which is used in our classification scheme.} \label{fig:thermo} \end{figure*} +\begin{figure*}[!t] \centerline{\includegraphics[width=\textwidth]{../figures/figure2_thermodynamics.pdf}} + \caption{ + \textbf{Thermodynamic sources.} (\textbf{A}) Provenance of every pK$_\mathrm{a}$ ladder shipped in the database, across all 45,708 compounds and 56,002 reactions. ChemAxon Marvin supplies 79.4\% of compounds and 81.6\% of reactions. \emph{Hatching} marks the share reached through a second route: the calculator refuses polymers and organometallics as query molecules, so those are computed from SMILES instead. Marvin reported no dissociable proton between pH $-2$ and $16$ for 1,242 compounds. The remainder carry no structure to compute from, which is a curation gap rather than a protonation one. (\textbf{B}) Reactions that carry an energy computed from the source heuristic. (\textbf{C}) Reported uncertainty in kcal\,mol$^{-1}$, one axis per source because the three differ by more than an order of magnitude. The shaded band spans the 5th to 95th percentile, which is used in our classification scheme. + } \label{fig:thermo} +\end{figure*} \subsection{Thermodynamics}\label{sec:results-thermo} @@ -6,6 +10,11 @@ \subsection{Thermodynamics}\label{sec:results-thermo} %% Marvin 26.1 rebuild. Placeholders excluded per source: group contribution %% dg = 1e7 (26,555 records); eQuilibrator sigma >= 2500 kcal/mol (3,386), %% which is that source's refusal marker rather than an error bar. -The three predictors cover comparable fractions of the database and disagree substantially about how well they do it (Figure~\ref{fig:thermo}B). Uncertainty is reported on scales differing by more than an order of magnitude (Figure~\ref{fig:thermo}C): a median of 0.63\,kcal\,mol$^{-1}$ for eQuilibrator against 10.41\,kcal\,mol$^{-1}$ for group contribution and 17.01\,kcal\,mol$^{-1}$ for dGPredictor. +The three predictors cover comparable fractions of the database and disagree substantially about how well they do it (Figure~\ref{fig:thermo}B). Uncertainty is reported on scales differing by more than an order of magnitude (Figure~\ref{fig:thermo}C): median errors (kcal\,mol$^{-1}$) of 0.63 for eQuilibrator, 10.41 for group contribution, and 17.01 for dGPredictor. -\paragraph{Calibrating the reported errors against measurement.} We use the experimental anchors to evaluate the error estimates. For each source we collect one held-out prediction per anchored reaction, together with the uncertainty that source reports for it, and divide the residual against the measured \drGo{} by that uncertainty. Measured this way, eQuilibrator has a root mean square of 8.06 and a median of 2.27, with only 28.7\% of reactions within one reported standard deviation and 46.3\% within two, so it understates its error. Group contribution runs the other way, with a median of 0.32 and 73.6\% within one standard deviation, and dGPredictor is the best scaled of the three, with a root mean square of 1.15 and 87.9\% within one. A researcher choosing by smallest reported error would systematically choose the source that understates it, which is the argument against promoting any single source. The grading in the next section calibrates each source's confidence against measurement rather than taking its uncertainty at face value. +\paragraph{Calibrating the reported errors against measurement.} We use the experimental anchors to evaluate the error estimates. +For each source, we collect one held-out prediction per anchored reaction, together with the uncertainty that source reports for it, and divide the residual against the measured \drGo{} by that uncertainty. +Measured this way, eQuilibrator has a root mean square of 8.06 and a median of 2.27, with only 28.7\% of reactions within one reported standard deviation and 46.3\% within two, so it understates its error. +Group contribution, by contrast, has a median of 0.32 and 73.6\% within one standard deviation, and dGPredictor performs even better with a root mean square of 1.15 and 87.9\% within one standard deviation. +% A researcher choosing by smallest reported error would systematically choose the source that understates it, which is the argument against promoting any single source. +The grading in the next section calibrates each source's confidence against measurement rather than taking its uncertainty at face value.