An mRNA mapping brainteaser:
solving the RNase crossword puzzle
using pre- and post-column digestion
solving the RNase crossword puzzle
using pre- and post-column digestion
As mRNA therapeutics and vaccines move from promise to mainstream medicine, the analytical toolkit used to characterize them must keep pace. Sequence is now an established critical quality attribute (CQA) and mass spectrometry (MS) is the method of choice when modifications are unknown or cannot yet be annotated by next-generation sequencing (NGS). Yet for a technique still in its infancy when applied to molecules as large as messenger RNA, every percentage point of sequence coverage counts.
One of our recently published studies tackles this challenge head on. In short, we introduced a complementary LC-MS platform that raises oligonucleotide identification from the MS/MS level to the full-MS level by coupling a custom-built post-column RNase A cartridge to an ion pairing-reversed phase (IP-RP) LC-MS system. The result is an elegant two-enzyme “crossword” strategy that, when combined with conventional CID-based RNA mapping, achieves 89 % sequence coverage of the firefly luciferase (FLuc) mRNA open reading frame, among the highest reported for this benchmark molecule.
Why mRNA mapping is hard
mRNA is structurally complex. Beyond its open reading frame (ORF), it carries a 5′-cap, 5′- and 3′-untranslated regions (UTRs) as well as a 3′-poly A tail. In vitro transcribed (IVT) mRNA used in therapeutic contexts often incorporates modified nucleosides (such as N1-methylpseudouridine, m¹Ψ) to improve stability and reduce immunogenicity. Unintended modifications (e.g. oxidations, deaminations, mRNA-lipid adducts) can arise during synthesis, formulation or storage, and these are invisible to NGS.

mRNA primary structure. A ribonucleotide consists of a nucleobase (blue) linked to a ribose (green) and a 5’-phosphate. In an RNA polymer, the ribonucleoside units (nucleobase and ribose) are connected via phosphodiester bridges. A = adenosine, C = cytidine, G = guanosine, ORF = open reading frame, U = uridine, UTR = untranslated region.
RNA mapping by LC-MS relies on ribonuclease (RNase) digestion to convert the intact mRNA into a mixture of oligonucleotides amenable to mass analysis. Matching recorded masses and MS/MS fragmentation patterns to an in silico digest then reconstitutes sequence coverage. However, two problems limit this approach:
- Non-unique oligonucleotides: short sequences that occur multiple times along the mRNA and therefore cannot pinpoint a specific site;
- Isomeric oligonucleotides: sequences with identical mass that can only be distinguished by MS/MS, a process prone to incorrect annotation when the constituent monomers are drawn from a limited four-nucleotide alphabet.
Our strategy: two enzymes, one crossword
The core innovation in this study is a post-column immobilized enzyme reactor (IMER) packed with biotinylated RNase A coupled to streptavidin-sepharose beads. The cartridge is custom-built to a miniaturized format to avoid peak broadening or retention-time distortion when placed downstream of the analytical LC column.
Our workflow begins with a pre-column digestion of FLuc mRNA using RNase 4 (an enzyme that cleaves between UA and UG dinucleotide sites) generating comparatively long oligonucleotides (median length ~15 nucleotides) that are largely unique. These oligonucleotides are then separated using ion pair reversed-phase (IP-RP) chromatography and, as they elute, pass through either an empty reference cartridge or our RNase A-IMER. RNase A cleaves 3′ of every pyrimidine (C, U), so each RNase 4-derived oligonucleotide is further fragmented into a characteristic set of short sub-sequences (its so-called “RNase A fingerprint”).
Analyzing the two LC-MS runs in parallel, with chromatogram alignment by parametric time warping, effectively solves a nucleotide-level crossword: the RNase 4 fragments define the long-range context and the RNase A sub-fragments pin down the internal sequence.

Cartridge design and workflow in cartridge-based RNA mapping. The RNase A cartridge was constructed in three steps: (1) production of biotin-labelled RNase A, (2) reaction of biotin-labelled RNase A with streptavidin-Sepharose beads yielding RNase A-coupled beads and (3) cartridge packing (inserted box). The workflow starts with mRNA digestion using, e.g., RNase T1 or RNase 4. The digest is then separated twice via LC-MS using either a non-functionalized, or “empty”, post-column cartridge [packed with streptavidin-Sepharose beads, also referred to as “cartridge” (CART)] or a post-column RNase A cartridge [packed with RNase A-coupled beads, also referred to as immobilized enzyme reactor (IMER)], yielding the mMWs of the RNase T1- or RNase 4-derived oligonucleotides or the RNase A-derived fragments, respectively. Data processing occurs via MSsenger® and, following chromatogram alignment, grouping of the mMWs of the fragments with the mMWs of their respective oligonucleotides allows RNA mapping. c 2′,3′-cyclic phosphodiester at 3′-terminus of the nucleic acid sequence, HILIC hydrophilic interaction chromatography, IP-RP ion pairing-reversed phase, p 3′-linear phosphate at 3′-terminus of the nucleic acid sequence.
A key practical detail is that RNase A produces 2′,3′-cyclic phosphodiester intermediates rather than 3′-linear phosphates under the rapid flow-through conditions of the IMER, which we have confirmed experimentally. This is consistent with the known enzymatic mechanism of bovine pancreatic RNase A and confirms that post-column digestion and IP-RP separation conditions are well balanced, efficient enough to generate relevant fragments, yet controlled enough to preserve the chromatographic profile.
Three annotation procedures, one answer
To extract the most information from our dual LC-MS dataset, we implemented three complementary in silico annotation procedures within our in-house R package, MSsenger®:
- APR1 (forward procedure, IMER data only): searches for co-eluting features representing an intact RNase 4 oligonucleotide and its RNase A fragments in the IP-RP-IMER-MS chromatogram.
- APR2 (forward procedure, cross-run): extends APR1 by combining oligonucleotide annotation from the reference (CART) run with RNase A fragment annotation from the IMER run, capitalizing on the high-quality chromatogram alignment.
- APR3 (backward procedure): designed for the many large oligonucleotides whose intact MS signal is too weak or too broad to be reliably extracted. Instead of starting from the oligonucleotide, APR3 searches for clusters of co-eluting RNase A fragments and in-source fragmentation ions that together tile a contiguous region of the mRNA sequence consistent with a single RNase 4 cleavage product. Statistical validation uses a resampling procedure against random isomeric sequences.
APR3 proved remarkably powerful: it assigned 194 oligonucleotides versus 92 from the combined forward approaches, and the SC reached with APR3 alone (71 %) nearly equaled that from all three cartridge procedures combined (72 %). When cartridge-based RNA mapping was combined with CID-based RNA mapping (using both RNase 4 and RNase T1 digests) our total SC rose to 89%.
Validation and complementarity
Our two approaches showed to be genuinely complementary, not redundant. Of the 95 oligonucleotides assigned by CID-based RNA mapping, 12 were not recovered by the cartridge procedure and vice versa. Long RNase 4-derived oligonucleotides spread their MS signal across many charge states and are difficult to fragment cleanly by CID. So this is precisely where the IMER excels. In contrast, shorter oligonucleotides with strong MS signals but ambiguous CID spectra benefit most from the independent orthogonal evidence that cartridge-based annotation provides.
The platform was first validated on a mixture of three synthetic isomeric heptamers (CAACCUG, CUACCAG, ACCAUCG). The IMER generated a distinct RNase A spectral fingerprint for each, clearly distinguishing sequences that are indistinguishable by mass alone.
The 5′-capping efficiency (98 %) and poly A tail length (mode: 124 A units, essentially monodisperse) of the FLuc mRNA preparation were also determined from the full MS data, illustrating the breadth of structural information accessible within a single LC-MS experiment.

Peak annotation via cartridge-based search in the FLuc mRNA digest LC-MS profile. Chromatographic peak regions are indicated by numbers (bold italic black font). In the first half of the IP-RP-IMER-MS chromatogram, many peaks were equally well annotated by the forward (APR1, APR2) and backward (APR3) search procedures (the displayed resampling-based significance for a particular annotation is the lowest obtained among these three procedures) whereas peaks in the second half of the chromatogram were mainly annotated via the APR3 approach (two peaks could only be annotated via the forward procedures and are indicated by ‘(f)’). In addition, certain peaks that were not retrieved by any of the three procedures, were manually annotated (grey-colored sequences) explaining why some of the annotated peaks lack a significance value. The displayed RNase 4-derived oligonucleotide sequences could not always be assigned to chromatographic peaks and only their feature-assigned fragments were taken for SC calculation.
Implications for mRNA drug development
Regulatory agencies expect biotech companies to confirm the sequence of their therapeutic mRNA and detect modifications that NGS cannot see. Our multi-RNase approach directly addresses this need with several practical advantages:
- Full-MS-level identification: by reading an RNA fragment pattern rather than relying on CID alone, the approach sidesteps the well-documented pitfalls of MS/MS-based oligonucleotide annotation (isomeric ion series, i-ion ambiguity, co-elution artifacts).
- Scalable to any mRNA length: our strategy can be applied to longer transcripts such as Cas9 mRNA and self-amplifying RNA (saRNA). Algorithmic improvements in molecular feature extraction and the use of alternative RNase combinations (e.g. RNase U2 in the post-column cartridge) are identified as clear paths to further gains.
- Modification localization at full MS: a future application is the use of paired oligonucleotide/fragment mass differences to localize unknown modifications without needing to record a CID spectrum. This could be a significant advantage for tracking labile adducts.
- Automated and reproducible: IMER showed no carry-over and consistent performance across multiple days of injection. Chromatogram alignment by parametric time warping is automated in our MSsenger® tool, making the workflow amenable to routine analytical pipelines.
- Low material requirement: the entire experiment runs on less than 10 µg of digested mRNA per injection, solving a practical constraint in early-stage drug development where material is limited.
Conclusion
For analytical scientists working in mRNA drug development, our work offers a ready-to-implement blueprint. For the broader biotech community, it represents a meaningful step towards the database-driven, automated LC-MS characterization pipelines that comprehensive mRNA quality control will eventually require.
Want to learn more about our advanced mRNA analytics? Let’s talk!
