Peer-reviewed

Forensic protein tool narrows microbial sources, but misses some

MARLOWE performed strongly on pure cultures and primary organisms in simulated mixtures, but secondary contributors and a comparison with MiCId showed clear limits.

MARLOWE, a computational method that uses de novo peptide tags, or short protein-sequence clues from mass spectrometry, to characterize unknown forensic biological samples, performed best on pure cultures and primary organisms in simulated mixtures. It correctly characterized 94% of pure bacterial cultures at species level. In mixtures, it put the primary organism first at genus and broader taxonomic-group levels and within the top two at species level. But across 225 MS/MS files, the correct source fell inside the top-ranked group in 63% of cases, and MiCId had higher specificity in the same comparison.

How the ranking works

MARLOWE's workflow is built to assign an unknown sample to potential organisms in a broad sequence database. It performs de novo peptide identification, extracts high-confidence tags, filters peptides by strength, matches them to the database, corrects for peptides shared by multiple organisms, and scores taxonomic groups. The workflow produces ranked organism and taxonomic-group results.

To control the evidence entering that ranking, MARLOWE retained a tag only if it contained at least five consecutive residues with local confidence of 80 or more on a 100-point scale, and if the overall peptide-spectrum match confidence averaged at least 50. It also assigned peptide strength according to how many taxa were in the database and how many contained a given peptide, so shared clues and more distinctive clues were treated differently. The test used KEGG Genomes Release 91.0, accessed in July 2019, with 5,851 microorganisms. The implementation used R version 4.2.2 and three R packages.

The biodiversity evaluation included 66 pure bacterial cultures and 75 computationally constructed mixed samples. Those mixtures were divided into three 25-sample settings with 20, 30 or 40 proteins from a secondary source. Each simulated dataset used a different random sampling, and 40 proteins was the cap for that source.

Strongest on the main source

In the pure-culture results, species-level characterization was correct in 94%. Genus and broader taxonomic-group assignments were each 100% correct with the correct entry ranked first. Every species that was correctly characterized appeared in the top-two list.

The second source was harder to recover. In the simulated mixtures, MARLOWE identified the primary organism in 100% of cases when judged by the top-ranked genus and taxonomic group, and placed the primary species within the top two. The secondary contributor appeared in the top five in 72% to 76% of cases at genus and taxonomic-group levels, and in 60% to 68% of cases at species level.

The harder comparison

MARLOWE's limits were clearer in its comparison with MiCId. The two tools analyzed the same files against the same KEGG Genomes database, after files without clear ground truth or with sparse or low-quality data had been discarded. Across 225 MS/MS files, MARLOWE put the correct source inside the top-ranked taxonomic group in 63% of cases and placed it in some taxonomic group in 74%. At a 5% false discovery rate, MARLOWE's species-level specificity, a measure of avoiding false-positive identifications, was 91.4%, compared with 96.7% for MiCId. MiCId reached 97.8% at genus level.

A separate 42-sample test of the B. cereus group found that the true contributor ranked first at species level in 88% of samples. But a six-replicate B. cereus subset was much less steady: only two replicates ranked the true contributor first, although four placed it within the top three. The paper discusses those low scores alongside a database containing 14 B. cereus strains divided among five taxonomic groups.

A lead, not a verdict

The findings leave MARLOWE in a useful but limited role. The study presents it as a computational way to narrow potential source organisms, while the mixture evaluation was restricted to 75 computationally constructed samples. The authors call for larger real or synthetic mixtures and evaluation of more complex samples, and note that database size, contents and quality influence identifications. It is therefore a candidate-generation tool for investigative leads, not a definitive forensic attribution.

The paper reports public raw mass-spectrometry data under four ProteomeXchange accessions, PXD001860, PXD003669, PXD014522 and PXD023033. MARLOWE is offered as open-source software in the R packages MakeSearchSim, CandidateSearchDatabase and OrgIDPipeline.

Paper data and sources

Original title: MARLOWE: taxonomic characterization of unknown samples for forensics using de novo peptide identification.
Authors: Sarah C Jenson, Fanny Chu, Gelio Alves et al.
Journal/Repository: Scientific reports
Status: Peer-reviewed
First online: 2026-08-20
DOI: 10.1038/s41598-026-50102-3
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.