Reverse Complement Calculator
The reverse complement of a DNA or RNA sequence is its partner strand read 5′ to 3′. Get it, the complement or reverse, with IUPAC codes and GC content.
Calculator
Spaces, line breaks and numbers are ignored, so a FASTA or GenBank paste works as it is. Ends marked as in 5′-ATGC-3′ are read, and an unmarked sequence is read 5′ to 3′. It runs in your browser, but the sequence is also written into the page address so that a link reproduces it: do not paste a confidential one.
20 bases read as DNA, 5′ to 3′.
- Length Bases on this strand, ambiguity codes included and gaps left out. As a duplex it is the same number of base pairs.
- 20 nt
- GC content G, C and S over the 20 bases whose pairing strength is known. S is always G or C and W never is, so both count.
- 40.0 %
- Ambiguity codes Positions written with a code such as R, Y or N rather than one base.
- 0
- Palindromic Yes when the sequence is its own reverse complement, as the EcoRI site GAATTC is.
- No
Both strands
Composition
| Code | Stands for | Count | Share |
|---|---|---|---|
| A | Adenine | 7 | 35.0% |
| C | Cytosine | 4 | 20.0% |
| G | Guanine | 4 | 20.0% |
| T | Thymine | 5 | 25.0% |
Citing this tool
Last updated . Add the date you accessed it as well, which a citation of a page that can change asks for. If a specific result matters, cite the permalink from the tool’s share row instead of this page: it reproduces the exact parameters.
The equation
Watson and Crick (1953), with the NC-IUB codes of Cornish-Bowden (1985)
The other strand, read the right way round
The reverse complement of a DNA sequence is its partner strand written 5′ to 3′: replace every
base with the one it pairs with, A with T and G with C, then read the result backwards. The reverse
complement of 5′-ATGC-3′ is 5′-GCAT-3′. As a formula, r[i] = c(s[n + 1 − i]): the base
at position i of the answer is the complement of the base at position i counted from the other
end of the input.
The reversal is there because the two strands of a double helix run in opposite directions, as Watson and Crick described in 1953. Each strand has a 5′ end and a 3′ end, and by convention a sequence is written from its 5′ end to its 3′ end, so writing the partner in that same convention means starting from what was the right-hand end. The DNA double helix explorer shows the two antiparallel backbones this comes from.
Complement, reverse and reverse complement are three different sequences
Only one of the three is the partner strand written in the standard direction, which is the form a primer order or a sequence database expects. Starting from 5′-ATGCGT-3′:
| Operation | Result | What it is |
|---|---|---|
| Complement | 3′-TACGCA-5′ | The partner strand as it lies under the original, reading 3′ to 5′ |
| Reverse | 3′-TGCGTA-5′ | The original strand written backwards: the same molecule, and not its partner |
| Reverse complement | 5′-ACGCAT-3′ | The partner strand read 5′ to 3′, as sequences are written and primers ordered |
The complement and the reverse complement are the same molecule written in opposite directions. That is why the calculator labels every result with the way it reads: a complement copied out and then treated as if it read 5′ to 3′ is a common error, and nothing about the letters themselves gives it away.
Worked example: the T7 promoter primer
The T7 promoter primer is 5′-TAATACGACTCACTATAGGG-3′, 20 bases. Its complement, written under it base by base, is 3′-ATTATGCTGAGTGATATCCC-5′. Reading that from its own 5′ end, on the right, gives the reverse complement of the T7 promoter primer, 5′-CCCTATAGTGAGTCGTATTA-3′, which is what the calculator shows when it opens.
Notice that the three G bases at the 3′ end of the primer have become three C bases at the 5′ end of the answer. The 3′ end of one strand always sits opposite the 5′ end of the other, so the last bases of the input are the first bases of the output. The primer holds 4 G and 4 C, so 8 of its 20 bases are G or C, a GC content of 40.0 percent, and the reverse complement is 40.0 percent as well.
IUPAC nucleotide codes and their complements
A degenerate primer uses the IUPAC codes for a position that may hold more than one base, and each code complements to the code for the partners of its bases. R, A or G, pairs with Y, T or C. B, not A, pairs with V, not T. S, W and N are their own complements, because the partners of G or C are C or G, and so on. The codes are the NC-IUB recommendations of 1984 (Cornish-Bowden, 1985), and each complement follows from the bases its code stands for.
| Code | Stands for | Named from | Complement |
|---|---|---|---|
A | A | Adenine | T |
C | C | Cytosine | G |
G | G | Guanine | C |
T | T | Thymine | A |
U | U | Uracil | A |
R | A or G | puRine | Y |
Y | C or T | pYrimidine | R |
S | C or G | Strong, three hydrogen bonds | S |
W | A or T | Weak, two hydrogen bonds | W |
K | G or T | Keto | M |
M | A or C | aMino | K |
B | C, G or T | not A, B follows A | V |
D | A, G or T | not C, D follows C | H |
H | A, C or T | not G, H follows G | D |
V | A, C or G | not T or U, V follows U | B |
N | any base | aNy | N |
The table gives T as the partner of A, and on an RNA strand it is U. Take the 16S rRNA primer 27F, 5′-AGAGTTTGATCMTGGCTCAG-3′, where M means A or C. Its reverse complement is 5′-CTGAGCCAKGATCAAACTCT-3′, and the M has become K, G or T, the partners of A and C. A tool that knows only A, C, G and T has to stop at the M, drop it or leave it as it is, and an M left as it is in a reverse complement is wrong, because the partner of M is K.
X is not an IUPAC nucleotide code: the recommendations advise against it, because X is the symbol for xanthine. RepeatMasker still writes it over masked repeats when asked to, so the calculator reads it as an unknown base, pairs it with X, and says so.
GC content with ambiguity codes
GC content is the share of bases that are G or C:
GC = (G + C + S) / (A + C + G + T + U + S + W) × 100%. S counts because it is always G
or C, and W counts on the other side because it never is. A code such as M or N could be either, so
it is left out of the headline figure, which is the rule Biopython’s gc_fraction
applies by default, and reported as a range instead.
For 27F, 9 of the 19 bases whose strength is known are G or C, so the headline is 47.4 percent. The range is 45.0 to 50.0 percent: taking the M as A gives 9 of 20, and taking it as C gives 10 of 20. That range is the more honest answer for a degenerate primer, because the tube really does hold both molecules. Because every G sits opposite a C, the GC content of the reverse complement is always the same as that of the sequence it came from.
RNA, and DNA that pairs with RNA
An RNA sequence is complemented the same way with U in place of T. The calculator reads a sequence containing U and no T as RNA, so the start of the human HBB coding sequence as mRNA, 5′-AUGGUGCAUCUGACUCCUGAGGAG-3′, has the reverse complement 5′-CUCCUCAGGAGUCAGAUGCACCAU-3′.
Setting “Read as” to DNA writes T wherever the partner of A is needed, which gives the DNA oligo that pairs with an RNA, the sequence an antisense probe or a reverse transcription primer is ordered as. For the same stretch that is 5′-CTCCTCAGGAGTCAGATGCACCAT-3′. A sequence containing both T and U is read as DNA, and the calculator says so.
A strand written 3′ to 5′
Questions sometimes give a strand the other way round, as the template strand 3′-TACCACGTAGAC-5′. Marked ends are read, so the calculator turns this one round to 5′-CAGATGCACCAT-3′ before converting, and its reverse complement, 5′-ATGGTGCATCTG-3′, is the partner read 5′ to 3′: the coding strand, here the first twelve bases of the HBB coding sequence, starting with the ATG start codon. A sequence with no marks is read 5′ to 3′, the convention that sequence databases such as GenBank follow.
Palindromes: sequences that are their own reverse complement
Some sequences read the same on both strands, and the calculator flags them. The EcoRI site 5′-GAATTC-3′ and the BamHI site 5′-GGATCC-3′ are their own reverse complements, which is why enzymes that cut as a pair of identical subunits recognise sites of this shape. Codes compare as codes, so the HinfI site 5′-GANTC-3′ counts too: as a pattern it is its own reverse complement, although one particular site such as GAATC is not, because its partner reads GATTC.
Lists, FASTA files and GenBank records
A FASTA file with several records is converted record by record, in order, and each header is
kept with the operation added, as in >515F (Parada) reverse complement. A GenBank
record is read from its ORIGIN line to its closing //, so the annotation above it and
the numbers down the side do no harm, and a second record after the // is refused
rather than left out. For a list of primers or index sequences, one per line, tick
“One sequence per line”: each line is converted on its own and stays on its line, so the result can
be pasted back beside the list it came from.
Spaces, line breaks and numbers are layout and are ignored. Apart from marks on the ends, any other character that is not a nucleotide code, an X or a gap stops the conversion, with its line and character position. Dropping it quietly would return a sequence one base short whenever that character stood for a base, and a primer ordered one base short is a different primer.
What this does not cover
It converts sequences and counts bases. Pairing follows Watson and Crick’s rules only, so the G·U wobble pairs an RNA duplex can also form are not considered. It does not estimate a melting temperature, which depends on the salt and strand concentrations as well as the sequence. The primer Tm calculator works it out. It does not look for hairpins or primer dimers, and it has no letters for modified bases such as inosine or locked nucleic acids. Reading a sequence as codons and translating it into protein is the job of the DNA to protein translator. Once an oligo arrives, the volume to dissolve it in follows from its mass and molecular weight with the molarity calculator, a working stock comes from the dilution calculator, and measuring a DNA or RNA prep is a job for the nucleic acid quantification calculator.
The conversion runs in your browser, but the sequence is also written into the page address so that a link reproduces it, up to about 8,000 characters. The privacy notice explains where a page address can travel, so do not paste a sequence that has to stay confidential.
Common mistakes
- Complementing without reversing. 3′-TACG-5′ is the partner of 5′-ATGC-3′, but written out as TACG and read 5′ to 3′ it describes a different molecule. The reverse complement is GCAT.
- Reversing without complementing. CGTA is 5′-ATGC-3′ read backwards. Taken as 5′ to 3′ it is a different molecule, and it pairs with neither strand.
- Leaving an ambiguity code as it was. R becomes Y and M becomes K. Only S, W and N stay the same.
- Writing T in an RNA answer, or U in a DNA primer. A pairs with U in RNA and with T in DNA, so choose what the result is for rather than what the input was.
- Reading a strand given 3′ to 5′ as if it were 5′ to 3′. Mark the ends, or turn it round first.
- Converting a list as one sequence. Several primers pasted one per line are read as one long sequence unless “One sequence per line” is ticked, and the result is then every primer’s reverse complement run together, the last primer first.
Common questions
What is the reverse complement of a DNA sequence?
It is the other strand of the double helix, written 5′ to 3′. Swap every base for its partner, A for T and G for C, then read the result backwards, so the reverse complement of 5′-ATGC-3′ is 5′-GCAT-3′. The reversal is needed because the two strands run in opposite directions, which also means each strand of a duplex is the reverse complement of the other.
What is the difference between the complement and the reverse complement?
The direction they are read in. The complement pairs each base where it stands, so it reads 3′ to 5′ under the original, and the reverse complement is that same strand turned round to read 5′ to 3′, the way sequences are written and primers are ordered. For 5′-ATGCGT-3′ the complement is 3′-TACGCA-5′ and the reverse complement is 5′-ACGCAT-3′.
How are IUPAC ambiguity codes complemented?
Each code becomes the code for the partners of its bases, so R (A or G) becomes Y (T or C), M becomes K, B becomes V and D becomes H, while S, W and N are their own complements. The reverse complement of the 16S rRNA primer 27F, AGAGTTTGATCMTGGCTCAG, is therefore CTGAGCCAKGATCAAACTCT, with the M turned into a K. The codes are the NC-IUB recommendations of 1984 (Cornish-Bowden, 1985).
How do I find the reverse complement of an RNA sequence?
Pair A with U instead of T, then reverse as usual, so the reverse complement of 5′-AUGGUGCAUC-3′ is 5′-GAUGCACCAU-3′. The calculator does this by itself when a sequence contains U and no T. Setting “Read as” to DNA gives the DNA that pairs with the same RNA, 5′-GATGCACCAT-3′, which is the sequence to order for a DNA probe or a reverse transcription primer.
How is GC content calculated?
Count the G and C bases and divide by the number of bases: 8 of the 20 bases of the T7 promoter primer are G or C, so its GC content is 40.0 percent. With ambiguity codes, S counts as G or C and W as A or T, while a code such as N that could be either is left out and reported as a range. A sequence and its reverse complement always have the same GC content, because every G sits opposite a C.