Peer-reviewed

MetaMP Finds Membrane-Protein Agreement Depends on Label Rules

A peer-reviewed study combines four protein resources and keeps disagreements visible for benchmarking and expert review.

MetaMP, a platform designed to combine membrane-protein annotations, shows that agreement can change with the rule used to compare labels. On a deliberately difficult benchmark made up of disputed entries, OPM had the highest strict agreement, while MetaMP had the highest agreement after a biologically motivated rule collapsed related labels. The contrast is central to the platform's approach: it keeps source-specific values and visible conflicts available for review.

Using the strict comparison, OPM matched expert benchmark labels in 96 of 121 cases, or 79.34%. MetaMP matched 25 cases, or 20.66%, and MPstruc matched 17, or 14.05%. Under benchmark-aware matching, MetaMP matched 107 cases, or 88.43%, ahead of OPM at 98 cases, or 80.99%, and MPstruc at 95, or 78.51%.

The benchmark was not a general slice of membrane-protein records. It was assembled only from the 121 MPstruc-OPM conflicts, which represented 2.96% of 4,089 harmonized records. The authors therefore use it as a difficult-case test of reconciliation, not as an estimate of the overall prevalence of disagreement.

A shared record that preserves the trail

MetaMP was designed to harmonize annotations from four complementary sources: MPstruc, RCSB PDB, OPM and UniProt. Its live snapshot contained 4,095 rows covering 4,089 unique PDB codes. The system retained 282 annotation attributes, including 107 nominal fields and 175 quantitative ones. The counts describe that live snapshot and are not presented as exhaustive coverage of PDB.

The data layer uses an ETL workflow, which extracts, harmonizes and loads records while retaining source-specific values instead of overwriting them with a single answer. It also records visible conflicts and legacy or replacement mappings. The result is a data trail that keeps source divergence visible.

Different methods, different matches

The same caution applies to topology, the pattern of transmembrane segments. The exploratory comparison was made at the PDB-entry level and used 106 expert-curated records, with different predictors covering different subsets. TMAlphaFold-linked TMDET matched 54 of 58 records, or 93.10%, on its smaller overlap. MetaMP TMDET matched 76 of 106, or 71.70%. OPM-derived counts matched 19 of 95, or 20.00%, but showed the strongest linear association with expert counts, with Pearson r = 0.634. These results describe different methods and overlaps, not a representation-independent ranking of topology predictors.

A classifier for prioritization, not a final answer

MetaMP's classifier workflow evaluated 36 model bundles across six classifier families, two training modes and three feature representations built from seven OPM-derived descriptors. After exclusions, the training matrix contained 3,966 rows spanning 3,963 unique PDB codes. The selected semi-supervised No-DR Decision Tree reached an internal accuracy of 0.937 and a weighted F1 score of 0.928. On the expert benchmark with benchmark-aware matching, it scored 0.884 accuracy and 0.887 weighted F1; on exact labels, those figures fell to 0.207 and 0.142.

That gap matters. The strongest supervised UMAP Logistic Regression model had internal accuracy of 0.883 and weighted F1 of 0.833, benchmark-aware expert accuracy of 0.884 and weighted F1 of 0.885, and exact-label accuracy of 0.124 with F1 of 0.073. The models are presented as assistive baselines for standardized MPstruc labels, not definitive annotation engines.

Faster testing, limited causal evidence

The platform was also tested by 24 unpaid online volunteers: 13 male, 10 female and one who did not disclose gender. The evaluation used a guided training phase followed by independent testing for each task, with training and testing compared within participants.

Across three tasks, mean performance was 4.21 points on a 0-to-6 scale, with a standard deviation of 0.98 and a 95% confidence interval from 3.80 to 4.62. Average combined training and testing time was 9.35 minutes. Testing took 1.22 minutes per task, compared with 1.90 minutes for training, a reported reduction of 35.97%.

Statistical tests detected a difference between paired training and testing times. The Wilcoxon signed-rank test returned W = 46.0, P = 0.0020 and r = 0.61. Completion time also differed across tasks in Friedman and repeated-measures ANOVA analyses, with P values of 0.0137 and 0.0180. But the analysis of whether faster completion was linked to correctness was exploratory: the logistic coefficient was minus 0.1555 (P = 0.0504), while a clustered GEE sensitivity estimate was minus 0.1426 (P = 0.0126).

Because participants received guidance and the comparison was within participants, the time difference cannot by itself establish a causal training effect. The adapted SUS-style usability score averaged 72.81, with a median of 76.25 and a 95% confidence interval from 66.16 to 79.47. The point estimate was above the reference threshold of 68, but the interval overlapped it.

What the benchmark cannot settle

Several limits narrow the interpretation. The benchmark labels were assigned by a single domain expert without an independent second-annotator check. They are therefore expert-informed reference assignments, not formally adjudicated gold standards. The benchmark's selection from MPstruc-OPM conflicts also means that its scores should not be read as a general measure of annotation quality.

The classifier targets standardized MPstruc labels and is intended to support prioritization and expert review. The topology analysis compared entry-level outputs with heterogeneous overlap and representations. The authors do not present it as a universal ranking of topology predictors.

The software is open-source under the MIT License, and its repository includes the source code, database schema, Docker configuration and expert benchmark annotations. The authors report no dedicated external or internal funding, no external sponsor involvement and no competing interests.

Paper data and sources

Original title: MetaMP Ecosystem for Unified, Auditable, and Benchmark-Ready Data for Reliable Membrane Protein Annotation.
Authors: Ebenezer Awotoro, Chisom Anyabolu, Florian Schwarz et al.
Journal/Repository: Computational and structural biotechnology journal
Status: Peer-reviewed
First online: 2026-08-20
DOI: 10.34133/csbj.0165
Original paper

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.