Ribonucleic acid (RNA) is a polymer made of nucleotide (nt) building blocks, which are composed of a ribose sugar, a monophosphate group and one of the four nucleobases: adenine (A), guanine (G), cytosine (C) or uracil (U) which replaces thymine (T) found in DNA. The ribose (in RNA) features a 2’-hydroxyl group while the deoxyribose (in DNA) lacks this, which is the main cause for their structural and functional differences. A basic RNA nucleoside consists of a nucleobase and the ribose (thus simply a nucleotide excluding the phosphate group). To form a chain, the nucleotides are linked by phosphodiester bonds between the 5’-phosphate group and the 3’-hydroxyl group of the ribose. In this tutorial we will mainly use the 1-letter code designations (A, G, C and U) of these building blocks (see Table 1).
| Nucleobase | Nucleoside | 1-letter code | 3-letter code |
| Adenine | Adenosine | A | ADE |
| Guanine | Guanosine | G | GUA |
| Cytosine | Cytidine | C | CYT |
| Uracil | Uridine | U | URA |
| Thymine | Thymidine | T | THY |
RNA fulfills a plethora of different functions in all kingdoms of life, and these functions are linked to its specific nucleotide sequence as well as to functional RNA structures. The simplest RNA structures are single-strands and base-paired regions forming helical double strands. A-form helices are stabilized by the Watson-Crick (WC) base-pairs that are characterized by the typical hydrogen bond donor/acceptor complementarity between the G and C and between the A and U nucleobases (Figure 1).
Figure 1: Schematic representation and atom nomenclature for nucleic acids. The figure shows the common RNA nucleobases in a canonical Watson-Crick base-pair (top left for A with U and bottom left for G with C), the DNA nucleobase T (top right) and the nucleic acid backbone, composed of alternating phosphate and ribose groups (bottom right). The atom nomenclature as used in this tutorial is given. It should be mentioned that the nomenclature might slightly differ in other sources. In case of deoxyribose (DNA) the H2' and the O2'-HO2' are replaced by H2'1 and H2'2.
In addition to the canonical Watson-Crick base-pairs observed in the A-from helices, a wide range of alternative pairings such as reverse Watson-Crick, Hoogsteen and the so called “Wobble” base-pair can be observed (Figure 2 and Figure 3). These base-pairs are either associated with alterations in the local backbone and base-orientation or allow more complex structural shapes. The Wobble base-pair formed by G and U (Figure 2, right) is also frequently found in RNA double-stranded regions. In folded RNAs, many more base-pair combinations are possible, each one having a unique hydrogen bonding pattern. Hydrogen bonding can also involve the ribose and the phosphate backbone. Moreover, beside the standard RNA nucleosides, numerous modifications exist (such as pseuouridine and N6-methyladenosine), which can affect the RNA structure, stability and function (Database of RNA modifications).
The NMR-observable nuclei in an RNA molecule are 1H and 31P with nearly 100% natural abundance of the spin 1/2 isotope, 13C (~1% natural abundance spin 1/2) and 15N (0.4% natural abundance spin 1/2). The 1H nuclei in RNA can be divided into two groups: the chemically “labile”, nitrogen-bound (imino and amino) and oxygen-bound (hydroxyl) protons that can exchange with the protons of the water solvent and non-exchangeable carbon-bound protons of the nucleobases and the ribose. Apart from oxygen, virtually all atoms of an RNA molecule are NMR observable. Their respective characteristic chemical shift ranges in an NMR spectrum are depicted in Figure 4.
Whether an exchangeable proton is detectable by NMR depends on the lifetime of the proton before it exchanges with a water proton. This is mainly a concern for the imino protons that can exchange rapidly (up to ms timescale) unless they are protected from exchange by hydrogen bond formation or buried in the core of a folded RNA. For single-stranded RNA or RNA without a defined structure, the imino protons are therefore invisible to NMR. Imino proton signals appear in a spectral region of 10 to 15 ppm that is devoid of other 1H signals (Figure 4) and the peak position is very sensitive to the surrounding structure. This makes them excellent reporters for RNA structure and folding. Thus, NMR is the most straightforward tool to characterize RNA structures present in solution as it relies on the unique spectral pattern of proton signals as structure reporters that are inherently invisible for X-ray crystallography and cryo-electron microscopy.
NMR spectroscopy requires RNA samples of sufficient purity, homogeneity, and concentration to yield interpretable spectra with sufficient signal intensity. Limitations in molecular weight of the RNA under investigation are set by the number of signals in each spectral region, at some point leading to signal overlap, and by the relaxation properties of a given RNA, leading to signal intensity loss with increasing size by T2 relaxation and limited experiment repetition rate due to T1 relaxation. Each of these limitations can be pushed (in some cases dramatically so) towards larger RNA sizes by several advanced methods, however, an RNA molecule of up to 80 nucleotides (nts) with average nucleotide composition is regarded as suitable for NMR structural characterization.
| Routine | Challenging | Example | |
| High resolution structure determination (unlabeled RNA) | <10 nt | 10 nt ~3 kDa | 8mer UGUU tetraloop: bmrb 15157 |
| High resolution structure determination (13C15N labeled RNA) | 25 nt ~8 kDa | 40 nt ~12 kDa | 155mer HIV packaging signal: 2N1Q |
| Secondary structure verification (unlabeled) | 25 nt ~8 kDa | 40 nt ~12 kDa | Recognition of structured RNAs by proteins |
| Secondary structure verification (15N labeled) | 40 nt ~12 kDa | 80 nt ~23 kDa | |
| Integrative modelling RNAs, RNA-RNA and RNA-protein complexes/ | / | MDa | Telomerase complex, Ox40 3’UT |
| Ligand binding studies (unlabeled RNA) | 30 nt ~10 kDa | 80 nt ~23 kDa | SL1 targeting |
| Ligand binding studies(labeled RNA, segmental labeling) | 40 nt ~13 kDa | 80 nt ~23 kDa | Pseudoknot targeting, riboswitches |
| Detection of RNA modifications / editing (2’O-Me, Thioate, protonation) | 30 nt ~10 kDa | 70 nt ~20 kDa | Ribosomal substrate RNA for ErmB methyltransferase |
Computational biology offers several powerful tools to predict the RNA structure. However, the acceptable accuracy is currently limited to 2D secondary structure prediction – that is, the number and arrangement of Watson-Crick base-pairs formed by a given nucleotide sequence. Typically, the first question to NMR is whether the predicted RNA fold can be confirmed experimentally.
In brief, starting with an RNA sample containing all isotopes at natural abundance, of which the purity, concentration and structural homogeneity is verified (as described in chapter 2). A first assessment is to simply count the number of signals in a 1D 1H spectrum of the imino proton region. Each resolved signal corresponds to one imino proton involved in a hydrogen bond: one U:H3 for each A-U (or U-A) base-pair, one G:H1 for each G-C (or C-G) base-pair, and two imino protons, U:H3 and G:H1, for each G-U (U-G) wobble base-pair. The 2D 1H1H NOESY is then used to determine the sequence of base-pairs, which already yields a conclusive picture of the expected-vs-detected secondary RNA structure elements. A more detailed description for the prediction and experimental determination of the RNA secondary structure is presented in chapter 2 “NMR of unlabeled RNA”.
The first step of investigating an RNA by NMR is to define the minimal construct that still represents the native RNA element of interest. The minimal construct should be stable and uniformly folded, of high purity and of sufficient concentration. RNA construct design should be guided by available structural and phylogenetic information, computer folding predictions available from e.g. RNAfold, and sample preparation requirements. In many cases, the RNA construct needs to be adapted for certain NMR experiments, and in some cases, it is recommended to screen different RNA constructs by 1D 1H NMR to find the ideal one.
To achieve stable and uniform folding of an RNA, some simple design rules may be helpful, but should be applied with care, as they are not universally applicable:
1. Single-stranded overhangs should be avoided
2. Apical loops should be replaced by stable tetraloops such as GAGA or UUCG
3. Exposed palindromic sequences should be avoided
4. Termini of helices should be stabilized by 1 to 2 G-C base-pairs
5. The folding prediction should yield one well-defined minimal energy structure
To achieve sufficient sensitivity in NMR experiments, it is recommendable to prepare RNA samples at concentrations between 0.3 and 1.0 mM. Table 3 summarizes the commonly used NMR tubes with their recommended sample volumes, and the absolute quantities of RNA to achieve a concentration of 0.5 mM RNA, which corresponds to 5 mg/ml for a 10 kDa RNA, and 10 mg/ml for a 20 kDa RNA.
| NMR tube: | 5 mm tube | 3 mm tube | 5 mm Shigemi |
| Sample volume (in µl): | 550 | 180 | 300 |
| Quantity required (in nmol): | 275 | 90 | 150 |
Higher concentration increases the signal-to-noise ratio but can also promote the formation of dimers or aggregates, therefore concentrations higher than 1 mM are usually not advisable. Additionally, most measurements involving carbon and nitrogen nuclei will require isotope labeling to achieve sufficient signal-to-noise ratio and reduced measurement times.
High purity and sufficient concentration of an RNA sample can be achieved either by chemical or by enzymatic synthesis. The respective advantages, requirements and limitations of each method are shortly listed below, for a more extensive overview please refer e.g. to Schnieders et al.
The RNA is synthesized in the 3’ to 5’ direction, nucleotide by nucleotide from phosphor-amidite building blocks while attached to a solid phase. After each coupling step, a purification step is performed.
Advantages:
Disadvantages:
A recombinant phage-derived RNA polymerase (most common: T7) transcribes the RNA from a DNA-template containing an additional promotor sequence. Nucleotide triphosphate building blocks are added in 5’ to 3’ direction, forming a phosphodiester linkage and releasing pyrophosphate. The reaction was first described in 1987 by Milligan et al. and is commonly referred to as “in-vitro transcription” (IVT). Meanwhile, several IVT protocols have been optimized for state-of-the-art NMR applications (e.g. Karlsson et al., or Schnieders et al.)
Advantages:
Disadvantages:
After RNA synthesis, procedure-specific by-products, reactants and impurities need to be removed prior to the NMR experiment. Depending on the nature, molecular weight, charge, and solubility of these impurities, different purification procedures are recommended, many of which are modular and easily combined:
Dialysis
Isopropanol/Ethanol precipitation
Anion-exchange chromatography
Phenol/Chloroform/Isoamylalcohol extraction
Preparative PAGE
Reversed-phase HPLC
Size Exclusion Chromatography (SEC)
Homogeneous folding of purified RNAs is commonly achieved by a short heat denaturation step followed by rapid cooling on ice. For most stable RNA constructs, this procedure will prove successful: Misfolded species from the synthesis or the purification procedure are unfolded by heating the sample briefly to ~95 °C. Subsequent, rapid cooling promotes folding into low-energy monomeric species. In general, it is recommended to perform RNA folding in sterile deionized water to minimize the effect of salt and buffer components on RNA folding and integrity during denaturation. After RNA folding, exchange into the desired NMR buffer is performed by (centrifugal) dialysis. RNA folding should be verified by native gel electrophoresis directly prior to the NMR experiment. In addition, the expected RNA folding should be validated by 1D 1H NMR experiments.
Common problems in RNA folding and their solutions:
In general, RNA is unstable in basic environments due to self-cleavage of its phosphodiester bond by the cis-acting 2’-hydroxyl group, which is more nucleophilic at higher pH. Thus, RNAs are often stored and measured in slightly acidic conditions. NMR buffers consisting of proton-free or perdeuterated components are preferable. A commonly used NMR buffer for RNA is 10 to 25 mM NaPi or KPi at pH around 6.5. Usually, 50 mM NaCl or KCl is added to electrostatically neutralize the polyanionic RNA phosphate backbone and allow for stable folding. As stated before, extensive tertiary interactions and very compact RNA structures are only stable in the presence of divalent cations, usually Mg2+. Care must be taken if the RNA requires Ca2+ for folding, since calcium phosphate is insoluble. There are also RNA motifs that are stabilized by protonation; thus, it may be also beneficial to optimize the pH of the buffer.
The RNA sample can be prepared either in H2O (with 5-10 % D2O for the NMR lock) or in 100 % D2O buffers. The initial experiments should be carried out in H2O, as exchangeable protons are only detectable in H2O. The exchangeable protons are imino protons, which are crucial for secondary structure elucidation and amino protons, as bridging NOE contacts for imino-aromatic contacts. In D2O these protons are exchanged with deuterium atoms and become unobservable, but on the other hand D2O allows for better detection of ribose signals that resonate close to the water signal (e.g. in 1H1H NOESY or 1H13C HSQC spectra) and this can be important for high resolution structure determination.
The first goal of the NMR investigation should be to confirm or - if necessary - correct the predicted secondary structure. To this end signals of the base-paired imino-protons of the bases G:H1 and U:H3, which would be undetectable in D2O, are crucial. Even in H2O the imino protons are only detectable, when being base-paired, as their exchange with water is otherwise broadens the signals beyond observability. In the base-paired state however the hydrogen bonding decreases the exchange rate. A less typical exchange that might nonetheless occur under certain conditions is exchange of the purine (A and G) H8 protons with D2O. However, this is a process that can takes weeks to month to complete. Further protons that are only observable in H2O are the amino protons of A:H61/H62, C:H41/H42 and G:H21/H22. These protons form hydrogen bonds with carbonyl groups when in a base-pair, and therefore give further information about the secondary structure. Moreover, non WC base-pairs give different patterns than their canonical counterparts.
The advantage of D2O measurements is the absence of the otherwise dominating H2O signal at 4.7 ppm, that is close to the proton signals of the ribose ring and the pyrimidine (C and U) H5 protons. A significant number of these signals might overlap with the broad H2O peak and their intensities affected by the water suppression. Therefore, measurement in D2O is advisable for those interested to resolve these signals. Given the fact that these peaks are already quite close, D2O should especially be considerations for further studies, when 2D spectra are to be measured.
Figure 6: Impact of the solvent on the ribose signals in 2D 1H13C HSQC spectra measured either in H2O (left) and D2O (right) at 303 K. In the H2O version the water band at 4.7 ppm in the F2 dimension is well suppressed. Other spectra can feature a significant broader solvent band. The RNA peak intensity difference is due to different sample concentrations