DNA to Protein Translator
Translate DNA or mRNA into amino acids with the standard genetic code: the protein from the first AUG to the stop codon, and all 3 or 6 reading frames.
Calculator
Reads A, C, G and T or U in either case, and the IUPAC ambiguity codes such as N. Spaces, line numbers and one FASTA header line are ignored.
628 bases read as RNA.
- Protein length Amino acids from the first AUG to the stop codon. The stop codon adds none.
- 147 aa
- Starts at base The A of the AUG, counted from the 5′ end of the mRNA.
- 51
- Reading frame Frames +1, +2 and +3 start at bases 1, 2 and 3 of the mRNA.
- +3
- Stop codon Bases 492 to 494.
- UAA
- Mass Average mass of the chain as translated, from standard atomic weights: nothing modified, and the first methionine kept.
- 15,998 Da
- Longest ORF The same ORF as the protein: frame +3, from base 51.
- 147 aa
- Sequence Bases read, after spaces, numbers and any FASTA header were dropped.
- 628 nt
Protein from the first AUG
Met-Val-His-Leu-Thr-Pro-Glu-Glu-Lys-Ser-Ala-Val-Thr-Ala-Leu-Trp-Gly-Lys-Val-Asn-Val-Asp-Glu-Val-Gly-Gly-Glu-Ala-Leu-Gly-Arg-Leu-Leu-Val-Val-Tyr-Pro-Trp-Thr-Gln-Arg-Phe-Phe-Glu-Ser-Phe-Gly-Asp-Leu-Ser-Thr-Pro-Asp-Ala-Val-Met-Gly-Asn-Pro-Lys-Val-Lys-Ala-His-Gly-Lys-Lys-Val-Leu-Gly-Ala-Phe-Ser-Asp-Gly-Leu-Ala-His-Leu-Asp-Asn-Leu-Lys-Gly-Thr-Phe-Ala-Thr-Leu-Ser-Glu-Leu-His-Cys-Asp-Lys-Leu-His-Val-Asp-Pro-Glu-Asn-Phe-Arg-Leu-Leu-Gly-Asn-Val-Leu-Val-Cys-Val-Leu-Ala-His-His-Phe-Gly-Lys-Glu-Phe-Thr-Pro-Pro-Val-Gln-Ala-Ala-Tyr-Gln-Lys-Val-Val-Ala-Gly-Val-Ala-Asn-Ala-Leu-Ala-His-Lys-Tyr-His
How the protein was read
- First AUG at base 51, so the reading frame is +3
- Stop codon UAA at bases 492 to 494
- ORF length = 494 − 51 + 1 = 444 bases
- amino acids = 444/3 − 1 = 147
Reading frames
The three frames of this strand, in one-letter code. A stop is *, and each run from an AUG to its stop is highlighted. Choose a frame to see its codons.
TFASDTTVFTSNLKQTPWCI*LLRRSLPLLPCGAR*TWMKLVVRPWAGCWWSTLGPRGSLSPLGICPLLMLLWATLR*RLMARKCSVPLVMAWLTWTTSRAPLPH*VSCTVTSCTWILRTSGSWATCWSVCWPITLAKNSPHQCRLPIRKWWLVWLMPWPTSITKLAFLLSNFY*RFLCSLSPTTKLGDIMKGLEHLDSA**KTFIFIA
HLLLTQLCSLATSNRHHGASDS*GEVCRYCPVGQGERG*SWW*GPGQAAGGLPLDPEVL*VLWGSVHS*CCYGQP*GEGSWQESARCL**WPGSPGQPQGHLCHTE*AAL*QAARGS*ELQAPGQRAGLCAGPSLWQRIHPTSAGCLSESGGWCG*CPGPQVSLSSLSCCPISIKGSFVP*VQLLNWGIL*RALSIWILPNKKHLFSLQ
ICF*HNCVH*QPQTDTMVHLTPEEKSAVTALWGKVNVDEVGGEALGRLLVVYPWTQRFFESFGDLSTPDAVMGNPKVKAHGKKVLGAFSDGLAHLDNLKGTFATLSELHCDKLHVDPENFRLLGNVLVCVLAHHFGKEFTPPVQAAYQKVVAGVANALAHKYH*ARFLAVQFLLKVPLFPKSNY*TGGYYEGP*ASGFCLIKNIYFHC
Codons in frame +3
This frame starts at base 3, so bases 1 and 2 are not read.
The last 2 bases do not make a full codon.
- 3 AUUIleUGCCysUUCPheUGATerCACHisAACAsnUGUCysGUUValCACHisUAGTer
- 33 CAAGlnCCUProCAAGlnACAThrGACAspACCThrAUGMetGUGValCAUHisCUGLeu
- 63 ACUThrCCUProGAGGluGAGGluAAGLysUCUSerGCCAlaGUUValACUThrGCCAla
- 93 CUGLeuUGGTrpGGCGlyAAGLysGUGValAACAsnGUGValGAUAspGAAGluGUUVal
- 123 GGUGlyGGUGlyGAGGluGCCAlaCUGLeuGGCGlyAGGArgCUGLeuCUGLeuGUGVal
- 153 GUCValUACTyrCCUProUGGTrpACCThrCAGGlnAGGArgUUCPheUUUPheGAGGlu
- 183 UCCSerUUUPheGGGGlyGAUAspCUGLeuUCCSerACUThrCCUProGAUAspGCUAla
- 213 GUUValAUGMetGGCGlyAACAsnCCUProAAGLysGUGValAAGLysGCUAlaCAUHis
- 243 GGCGlyAAGLysAAALysGUGValCUCLeuGGUGlyGCCAlaUUUPheAGUSerGAUAsp
- 273 GGCGlyCUGLeuGCUAlaCACHisCUGLeuGACAspAACAsnCUCLeuAAGLysGGCGly
- 303 ACCThrUUUPheGCCAlaACAThrCUGLeuAGUSerGAGGluCUGLeuCACHisUGUCys
- 333 GACAspAAGLysCUGLeuCACHisGUGValGAUAspCCUProGAGGluAACAsnUUCPhe
- 363 AGGArgCUCLeuCUGLeuGGCGlyAACAsnGUGValCUGLeuGUCValUGUCysGUGVal
- 393 CUGLeuGCCAlaCAUHisCACHisUUUPheGGCGlyAAALysGAAGluUUCPheACCThr
- 423 CCAProCCAProGUGValCAGGlnGCUAlaGCCAlaUAUTyrCAGGlnAAALysGUGVal
- 453 GUGValGCUAlaGGUGlyGUGValGCUAlaAAUAsnGCCAlaCUGLeuGCCAlaCACHis
- 483 AAGLysUAUTyrCACHisUAATerGCUAlaCGCArgUUUPheCUULeuGCUAlaGUCVal
- 513 CAAGlnUUUPheCUALeuUUALeuAAGLysGUUValCCUProUUGLeuUUCPheCCUPro
- 543 AAGLysUCCSerAACAsnUACTyrUAATerACUThrGGGGlyGGAGlyUAUTyrUAUTyr
- 573 GAAGluGGGGlyCCUProUGATerGCAAlaUCUSerGGAGlyUUCPheUGCCysCUALeu
- 603 AUAIleAAALysAACAsnAUUIleUAUTyrUUUPheCAUHisUGCCys
Numbers are the position of each row’s first base. Shaded codons belong to an open reading frame, the AUG that opens one is outlined, and each stop reads Ter.
Citing this tool
Last updated . Add the date you accessed it as well, which a citation of a page that can change asks for. If a specific result matters, cite the permalink from the tool’s share row instead of this page: it reproduces the exact parameters.
The equation
Crick, Barnett, Brenner and Watts-Tobin (1961), with NCBI translation table 1
How DNA is translated into protein
To translate DNA into protein, write the coding strand as mRNA by putting U in place of T,
find the first AUG, and read the bases three at a time from there, looking each codon up in
the genetic code until you reach one of the three stop codons, UAA, UAG or UGA. Each codon
adds one amino acid and the stop codon adds none, which gives the count in the equation
above: amino acids = ORF length in bases ÷ 3 − 1.
The code is read three bases at a time, without overlap, from a fixed starting point, as Crick, Barnett, Brenner and Watts-Tobin inferred from frameshift mutations in 1961. The first codon was decoded the same year, when Nirenberg and Matthaei found that an RNA made only of U produced a chain made only of phenylalanine. The table this tool reads is NCBI’s translation table 1, the standard code, with the amino acid symbols of the IUPAC-IUB Joint Commission on Biochemical Nomenclature.
Reading the results
Paste a sequence or choose an example, then say what it is: a coding strand or mRNA, or a template strand written in either direction. With three reading frames the tool reads the strand you gave and reports the protein from the first AUG. With six it adds the other strand and reports the longest open reading frame instead, because it can no longer know which strand is transcribed.
- Protein length counts amino acids from the start codon to the stop codon, which adds none. Starts at base is the A of that AUG, and Reading frame says which of the three ways of cutting the strand into codons it lies in.
- Stop codon names the codon that ends the chain and where it is, or says there is none before the end of what you pasted. Mass is the average mass of the chain as translated, in daltons.
- Longest ORF is the longest run from an AUG to a stop in any of the three frames. When it differs from the protein, look at both. Calvo, Pagliarini and Mootha (2009) found upstream open reading frames, an AUG before the main coding sequence and out of frame with it, in about half of human and mouse transcripts, so on a real transcript the first AUG is not always where the main protein starts.
- The working shows the arithmetic, the frame overview shows every frame in one-letter code with each open reading frame highlighted, and the codon view puts every amino acid under its codon, with the position of each row’s first base at its left.
Worked example: human beta-globin
The sequence the calculator opens on is the whole mRNA of human beta-globin, RefSeq NM_000518.5, 628 bases long. Reading it the way a ribosome does:
- The first AUG is at base 51. Frame +3 reads codons that start at bases 3, 6, 9 and so on, and 51 is one of them, so that is the frame the protein is in.
-
From there the codons read
AUG GUG CAU CUG ACU CCU GAG, which is Met-Val-His-Leu-Thr-Pro-Glu, the start of the beta chain. - The first stop codon in that frame is
UAA, at bases 492 to 494. -
The reading frame is
494 − 51 + 1 = 444bases, which is 148 codons, so the chain is444/3 − 1 = 147amino acids long.
That matches NCBI’s own translation of the record, NP_000509.1, residue for residue. The mass readout gives 15,998 Da for the chain as translated, the figure UniProt lists for beta-globin, P68871. The residue masses come from the same standard atomic weights the Molar Mass Calculator uses.
The mRNA is longer than the protein at both ends. The 50 bases before the AUG are the 5′ untranslated region, and in frame +3 they hold a UAG at bases 30 to 32, so the protein cannot begin any earlier in that frame. The 134 bases after the stop are the 3′ untranslated region. Choose frames +1 and +2 to see what the same mRNA says when it is read out of frame: 7 and 14 stop codons, and no open reading frame longer than 39 amino acids.
One base, one amino acid: the sickle cell mutation
The sickle cell allele of beta-globin differs from the usual one at a single base. Choose the “Sickle cell mutation” example and base 70 changes from A to U, which turns the seventh codon from GAG into GUG and its amino acid from glutamic acid into valine. Every other codon and every other amino acid stays the same, and the mass readout falls from 15,998 to 15,968 Da.
HGVS nomenclature writes the change as c.20A>T and p.Glu7Val, counting from the A of the start codon. It is also called Glu6Val, because the methionine that starts the chain is removed from the finished protein and the older numbering begins at the valine after it.
Coding strand, template strand and mRNA
The mRNA has the same sequence as the coding strand of the DNA, with U in place of T, because it is made by pairing with the other strand, the template. So a coding strand only needs its Ts changed, while a template has to be complemented. Direction matters too: the two strands run in opposite directions, as the DNA double helix explorer shows, so a template written 5′ to 3′ also has to be reversed before it reads the way the mRNA does.
The “Template strand” example is the template 3′-TAC GGT CTT AAG ATT-5′.
Complementing it base by base gives 5′-AUG CCA GAA UUC UAA-3′, which reads
Met-Pro-Glu-Phe and then stops: four amino acids from 15 bases, since
15/3 − 1 = 4. The same template written 5′ to 3′ is
TTAGAATTCTGGCAT, and “Template strand, written 5′ to 3′” reverses and
complements it into the same mRNA.
Reading frames and open reading frames
A sequence can be cut into codons in three ways, starting at its first, second or third base, and the other strand gives three more, which is why there are six reading frames. The tool labels them +1 to +3 on the strand you gave and −1 to −3 on the other strand, each counted from its own 5′ end.
An open reading frame here means an AUG and every codon after it in the same frame, up to and including the first stop codon. An AUG inside one does not start another, so internal methionines stay part of the chain. In random sequence about one codon in 21 is a stop, because 3 of the 64 are, so a frame that runs for a hundred codons without one stands out, and the longest open reading frame is the usual first guess at a gene in unannotated DNA.
It is a guess, and the “Gene on the other strand” example shows why. It is the beta-globin coding sequence reverse complemented, so the gene now runs backwards. In three frames nothing real appears, yet frame +1 happens to contain no stop codon at all, and its first AUG, at base 94, runs 117 codons to the end of the sequence. The example opens in six frames, and there the real chain of 147 amino acids appears in frame −1, longer than the accident and closed by its stop codon.
The standard genetic code
Four bases read three at a time give 4³ = 64 codons. Sixty-one of them encode
the twenty amino acids and three are stops, so most amino acids have more than one codon.
Leucine, serine and arginine have six each, while methionine and tryptophan have one each.
The table is generated from the one the tool translates with, so the two cannot disagree.
| First two bases | Third base U | Third base C | Third base A | Third base G |
|---|---|---|---|---|
| UU | UUU Phe (F) | UUC Phe (F) | UUA Leu (L) | UUG Leu (L) |
| UC | UCU Ser (S) | UCC Ser (S) | UCA Ser (S) | UCG Ser (S) |
| UA | UAU Tyr (Y) | UAC Tyr (Y) | UAA Ter (*), stop | UAG Ter (*), stop |
| UG | UGU Cys (C) | UGC Cys (C) | UGA Ter (*), stop | UGG Trp (W) |
| CU | CUU Leu (L) | CUC Leu (L) | CUA Leu (L) | CUG Leu (L) |
| CC | CCU Pro (P) | CCC Pro (P) | CCA Pro (P) | CCG Pro (P) |
| CA | CAU His (H) | CAC His (H) | CAA Gln (Q) | CAG Gln (Q) |
| CG | CGU Arg (R) | CGC Arg (R) | CGA Arg (R) | CGG Arg (R) |
| AU | AUU Ile (I) | AUC Ile (I) | AUA Ile (I) | AUG Met (M), start |
| AC | ACU Thr (T) | ACC Thr (T) | ACA Thr (T) | ACG Thr (T) |
| AA | AAU Asn (N) | AAC Asn (N) | AAA Lys (K) | AAG Lys (K) |
| AG | AGU Ser (S) | AGC Ser (S) | AGA Arg (R) | AGG Arg (R) |
| GU | GUU Val (V) | GUC Val (V) | GUA Val (V) | GUG Val (V) |
| GC | GCU Ala (A) | GCC Ala (A) | GCA Ala (A) | GCG Ala (A) |
| GA | GAU Asp (D) | GAC Asp (D) | GAA Glu (E) | GAG Glu (E) |
| GG | GGU Gly (G) | GGC Gly (G) | GGA Gly (G) | GGG Gly (G) |
In one-letter code a stop is written *, and in three-letter code Ter, the forms NCBI and the HGVS nomenclature use. X, or Xaa, is an amino acid the sequence does not decide: it appears when a codon holds an ambiguity code such as N and its possible readings disagree. GCN still reads as alanine, because GCA, GCC, GCG and GCU all encode it.
What this does not cover
- Other genetic codes. Only the standard code is used. Mitochondria and some organisms read a few codons differently: in NCBI’s vertebrate mitochondrial code, table 2, UGA encodes tryptophan, AUA encodes methionine, and AGA and AGG are stops.
- Other start codons. Only AUG opens a reading frame here. NCBI’s standard code also lists UUG and CUG as possible starts and its bacterial code adds GUG among others, and a chain that starts at one of them still begins with methionine, or formylmethionine in bacteria, because the initiator tRNA carries it whatever the codon.
- How a ribosome chooses its AUG. In eukaryotes the ribosome usually starts at the first AUG it meets while scanning from the 5′ end, but a weak sequence around that AUG lets some ribosomes pass it. Bacteria place the ribosome with a Shine-Dalgarno sequence a few bases upstream, so their start need not be the first AUG, and one bacterial mRNA can carry several genes. The first AUG shown here is the scanning rule and nothing more.
- Introns. The genomic DNA of a human gene is usually interrupted by introns, which are spliced out before translation, so translating it directly gives the wrong protein. Paste the mRNA or the coding sequence, which a RefSeq record such as NM_000518.5 gives.
- Recoding. A few proteins carry selenocysteine at a UGA codon, and some archaea and bacteria put pyrrolysine at UAG. Programmed frameshifts and RNA editing change proteins too. The tool reads every UGA and UAG as a stop.
- What happens to the chain afterwards. The mass is of the chain as translated, with its first methionine and nothing added or cut. Signal peptides, disulfide bonds, glycosylation and other modifications all change it; for beta-globin, losing the methionine alone takes the chain from 15,998 to 15,867 Da. To measure how much of a protein you have, absorbance at 280 nm and the Beer-Lambert law calculator are the usual route, with an extinction coefficient that comes mostly from the protein’s tryptophan and tyrosine.
Common mistakes
- Translating the template strand as it stands. The template is the complement of the mRNA, so its bases read as codons give a different protein. The example template begins TAC, which would be tyrosine; its mRNA begins AUG, methionine.
- Complementing a 5′ to 3′ template without reversing it. The result looks like a plausible mRNA and runs the wrong way.
- Starting at base 1. Translation starts at the first AUG. Read from base 1, the beta-globin mRNA begins Thr-Phe-Ala-Ser and meets a stop after 20 codons, while the real protein starts at base 51 in another frame.
- Counting the stop codon as an amino acid. 444 bases is 148 codons but 147 amino acids, because the last codon is the stop.
- Reading past the stop. Translation ends at the first stop codon in the frame. Anything after it, a later AUG in the same frame included, is not part of that protein.
- Mixing T and U. DNA has T and mRNA has U. The tool reads either, and writes codons with T for a coding strand typed as DNA and with U otherwise, but an answer should use the letter the question asks for.
- Expecting the finished protein to keep its methionine. Many lose it straight after translation, beta-globin included, which is why its residue numbers are often one lower than its codon numbers.
Common questions
How do you translate mRNA into amino acids?
Find the first AUG, then read the bases three at a time and look each codon up in the genetic code until you reach UAA, UAG or UGA, which end the chain without adding an amino acid. AUG GCU UGG UAA reads Met-Ala-Trp and then stops. The codon view here puts each amino acid directly under its codon, so the reading can be checked one codon at a time.
How do I get the mRNA from the template strand?
Pair every base with its complement, with U opposite A, and keep track of direction, because the mRNA runs antiparallel to the template. A template written 3′-TAC GGA TTC-5′ gives the mRNA 5′-AUG CCU AAG-3′, which reads Met-Pro-Lys. A template written 5′ to 3′ has to be reversed as well as complemented. Say which you have under “The sequence is” and the tool does both steps.
How many amino acids does a 300-base mRNA code for?
At most 100, and 99 if the 300 bases run from the A of the start codon to the last base of the stop codon, because the stop codon adds no amino acid. In general the count is the length of the open reading frame divided by three, minus one. Human beta-globin’s reading frame runs from base 51 to base 494 of its mRNA, 444 bases or 148 codons, and its chain is 147 amino acids long.
Why are there six reading frames?
Because each strand can be cut into codons starting at its first, second or third base, and DNA has two strands. A gene usually sits in just one frame of one strand, and the other frames are broken up by stop codons: in random sequence about one codon in 21 is a stop, since 3 of the 64 are. That is why a long run without a stop in one of the six frames is the usual first clue to a gene in unannotated DNA.
Does every protein start with methionine?
Nearly every chain does as it is made, because the initiator tRNA that reads the start codon carries methionine, or formylmethionine in bacteria, but many proteins lose it straight afterwards. Human beta-globin is one: translation begins Met-Val-His and the finished chain starts at the valine. That is why its sickle cell mutation, in the seventh codon, is also called Glu6Val, while HGVS numbering from the start codon gives p.Glu7Val.