Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
bee6401
Generalize ensemble summaries for prediction confidence and Rama8000
CedricHermansBIT Oct 6, 2026
58ef909
Test prediction confidence and Rama8000 ensemble summaries
CedricHermansBIT Oct 6, 2026
16add9a
Summarize prediction ensemble models and shared residue coverage
CedricHermansBIT Oct 6, 2026
837e2d6
Add prediction ensemble controls to Summary workflow
CedricHermansBIT Oct 6, 2026
6652f48
Do not silently include AlphaFold 3 in prediction ensembles
CedricHermansBIT Oct 6, 2026
b0a04d1
Analyse uploaded prediction ensembles and expose residue variability
CedricHermansBIT Oct 6, 2026
dfc3979
Link prediction variability map to residue inspection
CedricHermansBIT Oct 6, 2026
058bd0e
Style prediction ensemble workflow and variability map
CedricHermansBIT Oct 6, 2026
6b8dcc5
Test multi-file prediction ensemble aggregation
CedricHermansBIT Oct 6, 2026
e260c31
Exercise multi-file prediction ensemble in live browser
CedricHermansBIT Oct 6, 2026
e181a94
Fix comparison header browser assertion
CedricHermansBIT Oct 6, 2026
4aa2ed8
Preserve prediction ensemble hashes and reject duplicate models
CedricHermansBIT Oct 6, 2026
7846621
Document prediction ensemble capability in README
CedricHermansBIT Oct 6, 2026
ddcde20
Document prediction ensemble analysis and interpretation
CedricHermansBIT Oct 6, 2026
1002660
Make prediction ensemble summary robust to masked statistics
CedricHermansBIT Oct 6, 2026
f543b67
Allow practical multi-file prediction ensemble uploads
CedricHermansBIT Oct 6, 2026
8bbbe23
Document prediction ensemble inspection workflow
CedricHermansBIT Oct 6, 2026
12d5d25
Use multi-element Puppeteer evaluation for comparison headers
CedricHermansBIT Oct 6, 2026
bcf4531
Evaluate comparison headers as a node collection
CedricHermansBIT Oct 6, 2026
3f67cad
Assign ensemble test confidence and validation by residue identity
CedricHermansBIT Oct 6, 2026
3140f07
Merge main after Rama8000 integration
CedricHermansBIT Oct 6, 2026
54d1729
Merge README.md with conformational-change explorer
CedricHermansBIT Oct 6, 2026
43ec1f5
Merge docs/inspection-user-guide.md with conformational-change explorer
CedricHermansBIT Oct 6, 2026
e7cd315
Merge shinyRam/app.R with conformational-change explorer
CedricHermansBIT Oct 6, 2026
cc83438
Merge shinyRam/www/custom.js with conformational-change explorer
CedricHermansBIT Oct 6, 2026
4f479a9
Merge shinyRam/www/styles.css with conformational-change explorer
CedricHermansBIT Oct 6, 2026
09c2dee
Merge conformational change explorer into prediction ensemble
CedricHermansBIT Oct 6, 2026
4a1716b
Sync conformational comparison calculations from main
CedricHermansBIT Oct 6, 2026
2d4ba2c
Sync conformational comparison regression tests from main
CedricHermansBIT Oct 6, 2026
1147760
Restore main comparison browser coverage on prediction ensemble branch
CedricHermansBIT Oct 6, 2026
d85953a
Enforce 2 to 30 prediction ensemble bounds
CedricHermansBIT Oct 6, 2026
64adf48
Fix prediction ensemble cache and provenance handling
CedricHermansBIT Oct 6, 2026
3dc19d9
Test prediction ensemble model bounds
CedricHermansBIT Oct 6, 2026
b220a16
Assert nonzero prediction confidence disagreement
CedricHermansBIT Oct 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ RamplotR is an open-source R Shiny application that brings backbone geometry, a
| --- | --- |
| **Explore a structure** | Interactive φ/ψ plots with several reference-density datasets, native RamplotR density regions and a parallel six-class Rama8000 standard validation that matches the current cctbx/Phenix categories. |
| **Inspect residues in context** | Synchronized Ramachandran plot, searchable residue table, all-chain sequence navigator and NGL 3D viewer. The issue queue distinguishes Rama8000 outliers from native RamplotR `Not allowed` regions and missing angles. |
| **Work with predicted models** | Import AlphaFold DB models by UniProt accession or upload AlphaFold 2/3, ColabFold and ESMFold structures. Examine pLDDT, and view a linked PAE heatmap when compatible confidence data are available. |
| **Work with predicted models** | Import AlphaFold DB models or upload AlphaFold 2/3, ColabFold and ESMFold structures. Examine pLDDT/PAE and analyse multiple AF2/ColabFold or ESMFold seeds as a prediction ensemble with residue-level φ/ψ, Rama8000 and confidence agreement. |
| **Examine structural geometry** | Explore peptide ω, side-chain χ1 and descriptive Cβ measurements. Optionally attach the matching deposited structure's official wwPDB validation report for independent rotamer, clash and geometry annotations. |
| **Inspect experimental evidence** | Overlay a local CCP4/MRC cryo-EM map in the 3D viewer as a qualitative aid, without uploading the map to a separate service. |
| **Compare models** | Sequence-align chains from two structures; use the Conformational Change Explorer to navigate residue-level wrapped φ/ψ displacement, then inspect paired residues in linked 2D/3D views alongside native and Rama8000 category changes. |
Expand Down Expand Up @@ -105,6 +105,7 @@ The default RamplotR teal contour palette provides consistent, recognisable publ
- [AlphaFold, ColabFold and ESMFold confidence analysis](docs/prediction-confidence.md)
- [Geometry, official wwPDB evidence, cryo-EM overlays, ensembles and batch mode](docs/structural-verification.md)
- [Rama8000 standard validation and direct wwPDB comparison](docs/rama8000-validation.md)
- [Prediction ensemble analysis](docs/prediction-ensembles.md)
- [Independent wwPDB angle-validation protocol](docs/wwpdb-validation.md) · [Results](docs/validation-results.md)
- [Performance and large-structure benchmarks](docs/benchmark-results.md) · [Scaling results](docs/scaling-results.md)
- [Shinylive browser deployment and one.com caching](docs/shinylive-deployment.md)
Expand Down
23 changes: 23 additions & 0 deletions docs/inspection-user-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,29 @@ comparison can exceed the alignment size limit; select shorter chains instead.

Identical Ramachandran coordinates do not imply identical Cartesian structure, and an angular difference alone is not evidence of a clinically meaningful change. Distinct models from the same structure are related observations, not independent experiments.

## Prediction ensembles

When the loaded structure is a prediction, the **Summary** tab exposes a
prediction-ensemble workflow even if the current coordinate file contains only
one model. Upload additional AF2/ColabFold, ESMFold or other models whose
B-factor field contains pLDDT. The currently loaded compatible model can be
included as one ensemble member.

RamplotR reports circular phi/psi spread, residue coverage, Rama8000 agreement
and pLDDT spread across the uploaded models. A compact **Prediction variability
map** ranks residues by the larger circular SD of phi or psi. Clicking a map
cell or ensemble-table row selects the corresponding residue in the existing
inspector and linked 2D/3D views.

The variability map is a navigation tool. Model-to-model disagreement is
described as prediction uncertainty or heterogeneity, not molecular dynamics.
Duplicate coordinate files are rejected, and the model export records labels,
declared source and coordinate MD5 hashes.

AlphaFold 3 ensemble confidence is intentionally not inferred from B-factors in
this first version because each AF3 model needs its matching atom/token
confidence sidecar.

## Publication exports and reproducibility

The Summary tab includes an export disclosure. **Vector SVG** produces an editable figure; **High-resolution PNG** is suitable for manuscripts; **HTML report** includes the plot, summary counts, selected residue data and analysis settings. The residue CSV exports the current table filters; the comparison CSV exports the aligned comparison.
Expand Down
118 changes: 118 additions & 0 deletions docs/prediction-ensembles.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,118 @@
# Prediction ensemble analysis

RamplotR can analyse multiple independently generated prediction models as a
single **prediction ensemble**. The purpose is to expose where prediction
seeds/models agree or disagree in local backbone geometry and confidence.

This is deliberately not described as molecular dynamics. Variation between
predicted models can reflect model uncertainty, alternative plausible
solutions, stochastic sampling, different seeds or pipeline settings. It is
not experimental evidence that a protein moves between those conformations.

## Supported inputs

The first implementation accepts one coordinate file per model for:

- AlphaFold 2 / ColabFold;
- ESMFold;
- other prediction models that store pLDDT in the B-factor field.

The currently loaded compatible prediction can optionally be included as one
ensemble member.

AlphaFold 3 is intentionally excluded from automatic ensemble confidence
analysis for now. AF3 confidence is atom/token based and needs its matching
confidence JSON sidecar; RamplotR does not silently reinterpret an AF3 model as
AF2.

Each uploaded file must contain exactly one structural model. Up to 30 models
are analysed per run.

## Residue matching

Models are matched by:

- chain;
- residue number;
- insertion code;
- residue identity.

Missing residues remain missing and never receive fabricated coordinates,
angles or confidence scores. The Summary view reports how many residues are
present in every analysed model.

Duplicate coordinate files are rejected using MD5 hashes so repeated copies of
one model cannot artificially inflate ensemble agreement.

## Per-residue statistics

For every matched residue RamplotR reports:

- circular mean phi and psi;
- circular standard deviation of phi and psi;
- number of models contributing each angle;
- native RamplotR density-region agreement;
- Rama8000 Favored/Allowed/Outlier mode and agreement;
- mean, standard deviation, minimum and maximum pLDDT.

Circular statistics are essential because +179 degrees and -179 degrees are
neighbours rather than opposite conformations.

The prediction variability map uses the larger of phi-SD and psi-SD only as a
navigation aid:

| Maximum circular SD | Display |
| --- | --- |
| <5 degrees | stable |
| 5-15 degrees | moderate |
| 15-30 degrees | variable |
| >=30 degrees | high |

These bands are interface thresholds, not statistical significance cutoffs.

A separate outline marks residues whose Rama8000 category differs between
models.

## Model-level provenance

The model summary records:

- model/file label;
- residue count;
- number of residues with finite phi/psi;
- Rama8000 outlier count;
- mean and minimum pLDDT;
- declared prediction source;
- coordinate-file MD5 hash when available.

The two downloadable CSV files therefore preserve both residue-level ensemble
results and model-level provenance.

## Interpretation

A useful pattern is a residue with:

- large phi/psi spread;
- changing Rama8000 category;
- and/or large pLDDT variability across seeds.

That combination identifies a position worth inspecting in the 2D plot and 3D
structure. It does **not** by itself identify a biologically relevant
conformational switch.

Conversely, high pLDDT in every model does not imply that all models agree on
the same backbone geometry. RamplotR keeps confidence and geometric agreement
as separate measurements.

## Current limitations

- models are matched by residue identity rather than sequence realignment;
- AF3 multi-model confidence sidecars are not yet supported;
- ensemble members are not superposed or clustered in 3D yet;
- the current map summarizes local backbone variability, not global domain
motion;
- prediction ensembles are not a substitute for experimental ensemble or
dynamics data.

Pairwise structural differences can be explored separately in the Compare tab,
including the Conformational Change Explorer.
167 changes: 145 additions & 22 deletions shinyRam/R/ensemble.R
Original file line number Diff line number Diff line change
Expand Up @@ -39,13 +39,34 @@ ram_ensemble_summary <- function(models) {
membership <- matrix(FALSE,nrow=n,ncol=nmodels)
for(i in seq_along(models))
membership[,i] <- !is.na(match(all_keys,keys[[i]]))
if(!n) return(data.frame(chain=character(),resi=integer(),

optional_all <- function(field)
all(vapply(models,function(tbl) field %in% names(tbl),logical(1)))

if(!n) {
empty <- data.frame(chain=character(),resi=integer(),
insertion_code=character(),resn=character(),models_present=integer(),
phi_models=integer(),psi_models=integer(),
phi_mean=numeric(),phi_sd=numeric(),psi_mean=numeric(),psi_sd=numeric(),
classified_models=integer(),class_consistency=numeric(),
region_mode=character(),changes_class=logical(),
stringsAsFactors=FALSE))
stringsAsFactors=FALSE)
if(optional_all("rama8000_region")) {
empty$rama8000_models <- integer()
empty$rama8000_consistency <- numeric()
empty$rama8000_mode <- character()
empty$rama8000_changes <- logical()
}
if(optional_all("plddt")) {
empty$plddt_models <- integer()
empty$plddt_mean <- numeric()
empty$plddt_sd <- numeric()
empty$plddt_min <- numeric()
empty$plddt_max <- numeric()
}
return(empty)
}

get <- function(col, type="numeric") {
result <- matrix(if(type=="character") NA_character_ else NA_real_,
nrow=n,ncol=nmodels)
Expand All @@ -56,45 +77,147 @@ ram_ensemble_summary <- function(models) {
}
result
}
phi <- get("phi");psi <- get("psi");region <- get("region","character")
# Use the first occurrence of each ID across all models, not row number.
phi <- get("phi"); psi <- get("psi"); region <- get("region","character")

ids <- do.call(rbind,lapply(models,function(x)
x[,required[1:4],drop=FALSE]))
first <- !duplicated(unlist(keys,use.names=FALSE))
flat_keys <- unlist(keys,use.names=FALSE)
first <- !duplicated(flat_keys)
info <- ids[first,,drop=FALSE]
info <- info[match(all_keys,unlist(keys,use.names=FALSE)[first]),,drop=FALSE]
info <- info[match(all_keys,flat_keys[first]),,drop=FALSE]

calc <- function(matrix_values) {
result <- t(vapply(seq_len(n),function(i)
ram_ensemble_circular(matrix_values[i,]),numeric(2)))
colnames(result) <- c("mean","sd")
result
}
ph <- calc(phi);ps <- calc(psi)
region_mode <- vapply(seq_len(n),function(i) {
values <- region[i,]; values <- values[!is.na(values)]
if(!length(values)) NA_character_ else
names(sort(table(values),decreasing=TRUE))[[1L]]
},character(1))
agreement <- vapply(seq_len(n),function(i) {
values <- region[i,];values <- values[!is.na(values)]
if(!length(values)) NA_real_ else max(table(values))/length(values)
},numeric(1))
mode_and_consistency <- function(matrix_values) {
mode <- vapply(seq_len(n),function(i) {
values <- matrix_values[i,]; values <- values[!is.na(values)]
if(!length(values)) NA_character_ else
names(sort(table(values),decreasing=TRUE))[[1L]]
},character(1))
consistency <- vapply(seq_len(n),function(i) {
values <- matrix_values[i,]; values <- values[!is.na(values)]
if(!length(values)) NA_real_ else max(table(values))/length(values)
},numeric(1))
list(mode=mode,consistency=consistency,
count=as.integer(rowSums(!is.na(matrix_values))))
}

ph <- calc(phi); ps <- calc(psi)
native <- mode_and_consistency(region)
out <- data.frame(info,
models_present=as.integer(rowSums(membership)),
phi_models=as.integer(rowSums(is.finite(phi))),
psi_models=as.integer(rowSums(is.finite(psi))),
phi_mean=ph[,"mean"],phi_sd=ph[,"sd"],
psi_mean=ps[,"mean"],psi_sd=ps[,"sd"],
classified_models=as.integer(rowSums(!is.na(region))),
class_consistency=as.numeric(agreement),region_mode=region_mode,
changes_class=!is.na(agreement) & agreement<1,
classified_models=native$count,
class_consistency=as.numeric(native$consistency),
region_mode=native$mode,
changes_class=!is.na(native$consistency) & native$consistency<1,
stringsAsFactors=FALSE,check.names=FALSE)
out[order(-as.integer(out$changes_class),
-pmax(replace(out$phi_sd,is.na(out$phi_sd),0),
replace(out$psi_sd,is.na(out$psi_sd),0)),

if(optional_all("rama8000_region")) {
standard <- get("rama8000_region","character")
stat <- mode_and_consistency(standard)
out$rama8000_models <- stat$count
out$rama8000_consistency <- stat$consistency
out$rama8000_mode <- stat$mode
out$rama8000_changes <- !is.na(stat$consistency) & stat$consistency<1
}

if(optional_all("plddt")) {
confidence <- get("plddt")
finite_count <- rowSums(is.finite(confidence))
safe_stat <- function(fun) vapply(seq_len(n),function(i) {
values <- confidence[i,is.finite(confidence[i,])]
if(!length(values)) NA_real_ else fun(values)
},numeric(1))
out$plddt_models <- as.integer(finite_count)
out$plddt_mean <- safe_stat(base::mean)
out$plddt_sd <- vapply(seq_len(n),function(i) {
values <- confidence[i,is.finite(confidence[i,])]
if(length(values)<2L) NA_real_ else stats::sd(values)
},numeric(1))
out$plddt_min <- safe_stat(base::min)
out$plddt_max <- safe_stat(base::max)
}

spread <- pmax(replace(out$phi_sd,is.na(out$phi_sd),0),
replace(out$psi_sd,is.na(out$psi_sd),0))
standard_change <- if("rama8000_changes" %in% names(out))
as.integer(out$rama8000_changes) else rep(0L,nrow(out))
out[order(-standard_change,-as.integer(out$changes_class),-spread,
out$chain,out$resi,out$insertion_code),,drop=FALSE]
}

ram_prediction_ensemble_analyze <- function(pdbs, classifier, source,
labels = NULL,
max_models = 30L) {
if(!is.list(pdbs) || length(pdbs) < 2L)
stop("A prediction ensemble requires at least two predicted structures.")
permitted <- c("alphafold2","esmfold","other_prediction")
if(length(source)!=1L || !source %in% permitted)
stop("Prediction ensembles currently support AF2/ColabFold, ESMFold or other pLDDT-in-B-factor models.")
max_models <- suppressWarnings(as.integer(max_models))
if(length(max_models)!=1L || is.na(max_models) ||
max_models < 2L || max_models > 30L)
stop("Prediction ensemble max_models must be between 2 and 30.")
count <- min(length(pdbs),max_models)
if(is.null(labels)) labels <- paste("Model",seq_along(pdbs))
labels <- as.character(labels)
if(length(labels)!=length(pdbs) || any(!nzchar(labels)))
stop("Every prediction model needs a label.")

models <- lapply(seq_len(count),function(i) {
pdb <- ram_model_at(pdbs[[i]],1L)
torsions <- ram_extract_torsions(pdb)
classified <- classifier(torsions)
confidence <- ram_prediction_from_atoms(pdb,torsions,source)
keys <- ram_prediction_key(classified$chain,classified$resi,
classified$insertion_code)
confidence_keys <- ram_prediction_key(confidence$chain,confidence$resi,
confidence$insertion_code)
if(anyDuplicated(confidence_keys))
stop("Prediction residue identifiers must be unique.")
classified$plddt <- confidence$plddt[match(keys,confidence_keys)]
classified$confidence_category <- ram_plddt_category(classified$plddt)
classified$model_label <- labels[[i]]
classified
})
summary <- ram_ensemble_summary(models)
model_summary <- do.call(rbind,lapply(seq_len(count),function(i) {
table <- models[[i]]
finite_angles <- is.finite(table$phi) & is.finite(table$psi)
data.frame(
model=labels[[i]],
residues=nrow(table),
finite_phi_psi=sum(finite_angles),
rama8000_outliers=if ("rama8000_region" %in% names(table))
sum(table$rama8000_region=="Outlier",na.rm=TRUE) else NA_integer_,
plddt_mean=if ("plddt" %in% names(table) && any(is.finite(table$plddt)))
base::mean(table$plddt[is.finite(table$plddt)]) else NA_real_,
plddt_min=if ("plddt" %in% names(table) && any(is.finite(table$plddt)))
base::min(table$plddt[is.finite(table$plddt)]) else NA_real_,
stringsAsFactors=FALSE
)
}))
list(
summary=summary,
model_summary=model_summary,
models=models,
labels=labels[seq_len(count)],
analyzed_models=count,
available_models=length(pdbs),
common_residues=sum(summary$models_present==count),
source=source,
limited=count<length(pdbs)
)
}

ram_ensemble_analyze <- function(pdb,classifier,max_models=30L) {
if(!exists("ram_model_count",mode="function") ||
!exists("ram_model_at",mode="function") ||
Expand Down
Loading
Loading