B41[6% A SYMMETRICAL PATTERN IN THE GENETIC CODE The triplets of links in the genetic material which svecify the detalled structure of proteins heave now been identified in sufficient numbers to fall into a regular pattern. Richard V. ick The author is a biologist in the laboratory of Biology, Netional Cancer Institute, Bethesda, Maryland. Page < R.V. Sek The chemical structure of the genetic mechanism must have several very special characteristics. It must contain large, complex molecules capable of representing @ large amount of information intisir possible alternative configurations. It must be autocatalytic - capable of producing exact replicas of itsel?. It mist do this with great reliability, yet be able to mutate ocecasionally, amd still be autecatalytic in its mutated form. The Watson-Crick model demonstrates how these properties folio from the structure of DHA. But this accounts only for the self-reproduction and mutation of the genetic material. It must also have quite a different property. It must catalyse specifically some material other than itself, which enters into the living chemistry of the organism and results in the phenotypic expression of that "gene". If it is a gene for black rather than brown hair, for example, it must snecifically produce some chemical which affects the chain of events which finally results in the production of pigment and its deposition in the hair. The "one gene-one enzyme" formulation has led to the expectation that this non-genic product is protein. This second problem is more intricate than the first. The reproduction of the genes was pictured as some sort of mold-snd-cast system, and so it has proved to be. But for a “semplate” to determine the structure of some product fundamentally dissimilar from itself requires a more elaborate apparatus. Proteins are long, umbranched chains, made of twenty different kinds of links - amino acids. The exmct order of hundreds of these links appears to be genetically determined for each of thousands of different proteins. The DNA of the chromosomes also occurs as long, unbranched chains, but with only four different kinds of links - nucleotides. There are two kinds of nucleotides, two purines and two pyrimidines, which differ markedly in size. The specific matching of one puxine with one pyrimidine - edenine with thymine, cytosine with guanine - is the basis for the Watson-Crick model. But how could there possibly be such a matching between the nucleotides and the amino acids, especially in view of their different. numbers? This problem was posed in 1954 by Gamow (1) as a mathematical challenze te Page 3 R.V.Eck biochemistry. For each link in a protein chain being constructed, there are twenty possibilities. How can this be specified by a DNA chain containing only four different kinds of links? Evidently more than one, seemingly at least three mucleotides would heve to combine in the determination of a single aminc acid link, perhaps in some overlapping fashion. This mathematical puzzle remains valid, and several facets of it have been clarified. The DHA produces "messenger REA" apparently by some means similar to the Watson-Crick mechanism. RNA contains nucleotides corresponding to ah, with thymine replaced by uracil. The messenger RHA goes go the protein-producing organelles in the cytoplasm and determines the production of specific protein molecules, according to its detailed sequence of nucleotides. {ft does not do this directly, but vie a number of adapter molecules, called “transfer RNA's". Presumably each amino acid has one or more trensfer RNA which specifically attaches to it. Another part of the transfer RNA attaches to its specific nuclectide “codon" (triplet?) wherever it may occur in the polyribonucleotide chain (2). Thus each amino acid is held in its proper place while the polymerizing mechanism links it to the two adjacent amino acids ’ determined by the sequence of nucleotides in the information-carrying RNA. The elucidation of a codon pattern should provide evidence on the mechaniam of transfer RNA specificity. How, structurally, dees the specific /RHA secognize its proper codon combination? An ansver to this question is suggested in this paper. The Giscovery two years ago that a non-living system could be made to synthesize an artificial protein, polyphenylalanine, using as messenger RNA synthetic polyuridylic acid (3), has made possible a rapidly developing experimental attack on this problem. The in vitro synthesis of proteins using various synthetic copolymers of two or more ribonucleotides as artificial messenger RNA bas been reported from tvo laboratories. Tables have been published giving groups of three nucleotides ("triplets") which have been identified as "coding for" the various amino acids (4). There has been much interest and speculation in the mmber of codons and the possible relationships among then. Fasious peiteras have teen opopowwed winleb usue dis aocmmasiave the nieeholeres AM Mee te Re CUT Le Kage 4 R.V-Eek that date, and predict certain others. The discovery of additional triplets has then required that many of these hypotheses be abandoned or much modified. Judging by the rate at which new ones have been discovered it appears that nearly all, if not all 4 x4 xh « 64 mathematically possible triplets will ultimately be identified es codons. Will these appear to be a chactic jumble, or will they fall into some regular pattern? | Some regularity can be seen in the list of published triplets. No more amino acids are usually assigned to a nucleotide combination than there can be alternative rearrengements of it. This supports the expectation that the maximm mmber will not exceed 64. There are only three exceptions to this, two of which may be due to laboratory errors. The third exception will be considered later. It frequently appears that the two or more triplets found to code for the same amino acid have two of their three nucleotides in common. Roberts has used this property to derive a"ioublet" code » Which implies that the third nucleotide is irrelevant (5). This, of course, gives only 16 combinations and requires some supplementary explanation to account for the 20 amino acids. It has also been suggested that the code may be partly doublet and partly triplet. The Purine-Pyrimidine Pattern In the most recent lists from the two laboratories there are a total of 49 different triplets, 26 of them reported by both (4). This seems like a large enough proportion of the 64 to outline en over-all pattern, if one exists. The data can be arranged in various ways, and if one tabulates these triplets as in table 1, such a pattern emerges clearly. The pattern is this: All 64 triplets occur. Each amino acid is represented by one or more pairs of triplets which are identical except for one nuclectide. The non-identical nucleotides in each pair are the two purines or the two pyrimidines. For exmmple, ACC and AUC have been found to cede for histidine. GGU codes for tryptophan, and this pattzern predicts that GAU will also be found to ecde for tryptophan. This pattern seems acceptable sterecchemically. Furthermore, it is complete ond earns 7 + at 4 ye qs : sigs 7 a4 % ? oa: self-consistent. “here are ezactly 32 vairs. each using ef leest one seporied triples. Fase 5 R. Vu. Beck All 64 possible permutations are used. Of the 49 reported triplets, four were discarded: two because theve were four amine acids assigned where there could be only three permutations, one because it wes assigned to GUG In conflict with phenylalanine, and only one because it was inconsistent with the pattern presented here. The 19 remaining combinations have been assigned to the various amino acids in such a wey as to completc the pattern symmetrically. For example, arginine has a pair, CCG and CUG, and one unpaired triplet, AAC. The missing triplet could be AAA or AGG. But AAA is already fully “occupied”, and AGG is not. Arginine is therefore assigned the missing AGG. As this process continues, completing all the pairs, the number of remaining alternatives is greatly reduced, but ali the missing triplets can ke accounted for with only one discrepancy as notel. This lest polnat is Illustrated in table 2 in which the sane Jascignuents ere re- tabulated. Here one can readily confirm that each triplet has emctly as many amino acids assigned to it as there can be permutations. Any pattern of this sort is open to the suspicion that it may merely resemble the true pattern. 2% would be pointless to compute a “probability” that it could have occurred “by chance", because the data obviously fall into some sort of pattern - they are not rendom. Eut there is a suiteble test. One can attemnt to construct similar- appearing patterns of 32 pairs in which AC end GU are paired, or AU and CG. ‘This attempt was made; these patterns cannot be constructed without discarding an umreasonebic number of well-established date. For example, the matching triplet for phenylalanine ~- UUU would have to be UGU or UAU respectively. Both laboratories have identified each of these with three other amino acids. Oniy UCU, as in the proposed pattern, is free to represent phenylalanine. About seven such conflicts developed in each attempt. in the purins-pyrimidine pattern only one triplet assignment had to be discarded in this way, and it was reported from one leboretory only. This appears to be a moderately strong indication that this pattern (which incidentally “mikes sense" chemically) is not just ea contvived modification of some "partially doublet” code. eV Bele This test indicates that if there is 2 pattern of this general sort in the 64 possible triplets, the existing 46 data (after excluding the three which arc inconsistent in any case) are more than sufficient to determine that pattern. It appears that this many data could fit only into a true pattern. Previously, when there were too few data it was possible to devise en almost endless number of patterns in which they could be accomodated. Determination of Orter The pattern at this stage depends only on the published tables of triplets, not on date from amino acid "“metants", ete. If these additionsl clues are used, a beginning can be made in determining the order of the nucleotides within each triples. In the experiments which have yielded the triplet codons, no order can be determined. Thus, in table 1 "aac" means, "AAG, AGA, or GAA". In table 2, "ACG means "ACG, AGC, CAG, CGA, GAC, and GCA". In table 3, however, these orders have been assigned, and "CAA-threonine" means that exact order, with the provision the; the evilence is not rigorously conclusive, and it might be AAC. If the purine- pyrimidine link is always in the same noasitjon, and if this position were known, the sequence would be detexmined in all these cases where toe remaining tyo are the same (ACA = ACA). The experimentally deternincd sequence AU for *vrosine and GUU for cysteine (6) requires one of the isoleucine codons to be WA ana valine to be UUG if the specie] link is in the middie, or UAU and UGU if it is at the Fight end. it evidently cannot be at the left end. For illustration, it is assumed to be in the middie. Aside from these, the orderings in table 3 are not rigorously determined. Any nucleotide pair might be exchanged with the diagonally corresponding one and still be consistent with table 2. For example, lysine-AAg and isoleucine-UAA might be inter- chenged. However, some possibilities seem much more plausible than others. It seems reasonable to expect that a mutetion will often involve the change of a single link. The amino acids which could xeplace one another in this way might ke expected to be found most frequently as "allele" pairs in homologous protein sequences. For examle. we abe Ya alamine can change to serine if one ef ites miciesatides esenges from G te th Baie ean Page 7 R. V.Eck occur in two different pairs. However, if the alanines at C...G were exchanged with serine and arginine at G...C, there would then he no codons of alanine and serine having two letters in common. Since alenine-serine is the most Prequently-cccurring "allele" pair (7), the arrangement as shown is much preferred. Similarly, if valine and cysteine were exchanged, the numerous allele pairs val-ala, val-ilu, val-leu, and others would not be producible by single-link "interchanges". (This, with the previously-mentioned determination of the sequence valine. = WG, constitutes support for this procedure.) By comparing each possible alternative in this way with a list of about three hundred “alleles” from hemoglobin and other proteins, the tentative arrangement shown in table 3 was derived. There are some other uncertain details in table 3. Different amino acids might have been discarded. For example, glutamine at AGG might have been retained, and arginine at AAG discarded (table 2). Furthermore, glycine at GAG might exchange places with glutamine, if it were retained at AGG. There are, however, only a few such alternatives, and at each choice there seemed some good clue to the selection. Another source of uncertainty is the possibility of experinental error. Of the 32 pairs, five consist of an unreported (predicted) triplet and a triplet reported fram only one laboratory. Any of these might prove to be in error. Even if there were no basis for choice in the alternative positions indicated in table 3, it would contain much information about sequence. For each triplet in table 2 there are one, three, or six possible sequences. Table 3 reduces these to two alternatives at most. This pattern predicts all the amino acids coded by the remaining 19 ordered triplets, suggesting that there may be no "nonsense" combinations. This is not a strong inference, hovever. If one of the reported triplets is erroneous its assigned pair might represent "nonsense". Or, in a few cases the pairs might be sub-divided, one of the two triplets being "noasense". we ge R.V.eok Predictions Concerning Transfer RHA's It is consistent with this pattern that there could te 6! transfer RHA's,one for each triplet. The pattern of 32 pairs suggests, hovever, tha) there may be only 32 transfer RHA's and that each one responds indiscriminately to both of its specific triplets. fhe specificity of the attachment site would reside in: one of the four uvucleotides in one position, a purine or a pyrinidine in another position, and one of the four nucleotides in a third positien. These three determ nants vould cceur in three specific (not necessarily adjacent) pesitions on the mereenger RNA chein. The two triplets of each pair vould presumably te indistinguisbnole to the transfer RRA, which might xecognize the third (middle?) nucleotide oniy by its size. in this sense each codon would consist of a pair of tripleis. One might face selously eall this a “two-and-a-haif-Letter™ code. Gn this iuterpretaiion there woulda be only one wransfer RHA fer aspartic acid, cysteine, { BLuts wine, histidine, methionine, phenylalanine. trvotophan, tyrosine, and valine. There would be two for cach of the ovher amino acids except serine, which would have three. (Barring errors, as mentioned above.) Tye xpeximental determinetion of the number of transfer REA‘s for each amino acid vould ke @ poverful check on the correctness of rhis pattern, and therefore on the validity of the individual repovted triplets. Phe two specific eransfer RAA'ts reported for icucine are consistent with this prediction (2). A strong check vill alee be provided by each subsequens Gisecovery of a triplet Vor the evidence on vhich Roberts based his "@oublet” cade (5}. “NLS PEtitern as a cede whie’ bale p ~ fete “a ty $9 3 thy doublet ard partly triplet. This coulé te teste: = yay we Baye Spas Pe 4 fam = aS om wf ee sy af pepe Dana for For NEA con be found for alanine 2 BAVCILNS, A50L0 0A: or vee BO ae, veoh bet + + 1 Yn wt 2 sel ye: ee es 4% couliie, ooh tareontoc, eml only tve for sexine, these will Dy 2: Zz st at, te” 2 a benvener: “doi bleta” + Sang ’ meert fe one oe % 5 a so CO! semoe., On the obher Ran ids ok Poke Ay om Te : ele gga ga oe Ly Cet ee Poder PA Neat aay SE SE ie ee Fage 9 R.V.Eck There are four crucial tests which could te made with isolated transfer RNA's. Is the transfer RIA of proline which responds to CAC the same as the one which responds to CGC? Similarly for cee ee and AUG; glycine GCG an@ GUG; end leucize UAU and UGU. If these shovld prove to be identical, this pattem vould be validated. Unsettled Questions The finding that leucine as well as phenylalenine (4) is ceded by UUU is of considerable interest. Is it an accident caused by some abnormal condition in the in vitro situation, or is it of fundemental significance? It reises the possibility that pexhaps in the presence of some other source of information not normally present in the artificial situation there might be e second pattern of 32 codons. Perhaps some of the triplets discarded in making this pattsera are other ambiguities of this kind rather than errors. Evidence which seems contrary to this is the finding that the same transfer RNA of leucine responds to UGU and to WU (2). This pattern seems to be good evidence for a triplet code, since in it triplets are necessary and sufficient. Eut if the above-mentioned ambiguities prove to be fundamental, they would require extra information, which might reside in still other links (as an overlapping quadruples code?). If this pattern vere the complete code, any simple type of overlepping should be immediately evident by substituting the triplets for the amizo acids in a few of the known protein sequences. This does not appear to be the case. If it were possible to make regularly crdered synthetic polyuers such as poly- dinucleotides, etc., the question ef vhether the three dlements of each triplet are adjacent in the chain could be settled. Also, such polymers could he used to study the possibility of overlapping codes. I% seams possible that some regular BIA's covld be synthesized from synthetic DHA's using the mechanism reported by Chamberlin and Perg and by Otska et al. (8). Page 10 R.oV.Eck In retrospect, it seems that this simple pattern could have been discovered with fewer clues, and we may wonder vhy it was not found earlier. In the lasttwo years there have been a remarkable number of ad hoc proposals to account for the data currently at hand. When there vere about tvelve triplets identified it was expected by some that the total number would be emctly 20 - one for each amino acid. Later a number of special combinations were considered, such as combining the three nucleotides without regard to order, and others of this sort which were mathematically possible tut structurally unimaginable. Then there was the "high-U" code, chemically imlausible but mathematically capable of accounting for the results to that date. Recently it was proposed that there may be some simple pattern in vivo but that some circumstance ds obscuring it in the in vitro experiments. Ry permitting some normally hidden potentialities to be expressed, this would produce too many triplets. All of these were attempts to account for the number 20 and to guess the pattern at a stage when an indefinitely large number of patterns were yet mathematically possible. Por this reason the probabilities were strongly ageinst success in this approach. fhe point of view which led to this pattern was from the opposite direction: Whatever the pattern may te, there are 64 triplets. In the end some of them may be “nonsense”, or even non-existent in nature. Some may prove to be equivalent to others, etc. But whatever those details may te they will consist of some sub-pattern of the 6h mathematically possible triplets. Mou, with a total of 49 different triplets identified, surely the pattern must be visible! I tabulated these triplets in various weys and shortly this detail appeared: Five amino acids had two triplets, with two letters in common. Another had three, the third being unrelated to the first two. Of these six examples four contained the alternatives C vs. U (e.g. aspartic acid-ACG end AUG). It was not surprising that none contained © in such alternatives, since the G polymers had given the most experimental difficulty and most of the combingtions having two G's vere as yet unassigned. From this observation table 1 followed directly and the puzzle practically solved itself. Faege 12 R. VeEck A similar history has occurred in the related problem of overlapping or non- overlapping codes. Gamow suggested that the necessary amount of information might be reduced by seme systematic constraint on the sequences of amino ecids. This could be caused by the same nucleotides serving in more than one triplet simultaneously (1). At that time there was a moderate amount of protein sequence dataaetiable. Several curious overlapping codes were proposed, each of which ws followed enthusiastically end then disproved mathematically, using the then currently available data as it con- tinued to increase in amount. Then Brenner (9) concluded thet all overlapping triplet codes were inconsistent with the data, and the attention of most protein eryptographers was diverted to non-overlapping codes, where it has remained ever since. However, the problem was not approached in its most general form. Brenner's computation, the results of the single step mutations, and the results of Crick et al (10) have been taken to disprove overlapping codes, but this is true only for a certain sub-class of such codes. Furthermore it includes the unstated assumption that there is no other source of genetic information. (The results of Crick et al (10) have also been taken as evidence for a triplet code but this inference is valid only for non-overlapping codes. ) A large number of overlapping codes are still mathematically possible (11), and several are even structurally plausible which involve a regular folding or coiling of the RNA strand so that the non-adjacent nucleotides of the codons assume their specific positions. The substance of the argument against overlapping codes, whether from amino acid sequences or from the single-step mutation data, is that such codes could not contain enough information to account for the observed mumber of variations in protein sequences. However, all proposed ccdes have required, explicitly or implicitly, some additional unknown source of information such as "commas", spacers", "stepping by threes", “forbidden combinations", etc. As long as the nature of that additional mechanism is undiscovered the possibility remains that it could contain enough information to supplement: an overlapping code. Clear evidence of patterns in the constraints on protein scquences would be a basis for an attack on this problem. The amount of protein sesuence vta nov available my be barely sufficient for this (7). In numerical proportion this Page 12 R.V.EBck problem is at a mich earlier stage than that of the nucleotide triplets. Here we required about 45 of the 64 possibilities before the pattern revealed itself. In the protein cryptogream less than two thousand of the potential eight thousand tripeptide sequences have been reported. If, say, three thousand of the eight thousand were "forbidden" according to some pattern, could we see that pattern? Perhaps this seemingly obscure problem will seem simple when it is solved. Summary The accumulation of experimental results from the system of Nirenberg and Matthaei has now reached about 75% of the total possible, if the triplet” concept is correct. Considered as a mathematical puzzle this has proved to be sufficient to determine an apparently unique soluticn: The sixty-four combinations of four nucleotides taken three at a time, are resolved into thirty-two pairs. ‘The second member of eachmir is identical with the first, except that in one position a purine is replaced by the other purine or a pyrimidine by the other pyrimidine. Almost all of the reported triplets £it into this pattern, and it predicts which amino acids will be found te correspond to the remaining nineteen unidentified triplets. This pattern accounts for several of the observations concerning regularities inthe data. It partially determines the order of the nucleotides in each triplet and suggests a structurel basis for transfer RNA specificity. Whether the three nucleotides of each triplet are adjacent in the nucleic acid chain, and whether they somehow impose constraints on the possible sequences of amino acids which they determine, is yet to be worked out. 1. 6. Te Page 13 R.V-Bek Bibliography G. Gamow, Rature 173, 318 (195k). B. Weisblum, S. Benzer, and R. W. Holley, Proc. Mat. Acad. Sci. 18, 1h49 (19623. M. W. Nivenberg and J. H. Matthaei, Proc. Nat. Acad. Sci. 47, 1588 (1961). O. W. Jones, Jr. and N. W. Nirenberg, Proc. Nat. Acad. Sei. 48, 2115 (1962); J. Wahba, S. Gardner, C. Basilio, R. S. Miller, F. Speyer, and P. tengyel, ibid. 49, 116 (1963); Two additional triplets were reported at a meeting: J. Abelson, Science 139, 774 (1963). R. B. Roberts, Proc. Rat. Acad. Sci. 48, 897 (1962). A. J. Wahba, C. Basilio, J. FP. Speyer, P. Lengyel, R. 8S. Miller, and S. Ochoa, Proc. Hat. Acad. Sci. 48, 1683 (1962). R. V. Eck, J. Theoret. Biol. 2, 139 (1962). M. Chamberlin and P. Berg, Proc. Nat. Acad. Sci. 48, 81 (1962); B. Otaka, H. Mitsul, and S. Osawa, ibid. 48, 425 (1962). S. Brenner, Proc. Nat. Acad. Sci. 43, 687 (1957). F. H. C. Crick, L. Barnett, S. Bremer, ani R. J. Watts-Tobin, Nature, 192, 1227 (1961). R. Wall, Feture 193, 1268 (1962). Re Vek. A symmetrical pattern in the genetic code Alanine cCat# CAGH Leucine UAUF CUURy CUO CuUaG¢ = CRE uaueF ccue discard} Arginine cosy; aAAGH Lysine ASAT AAU! (ACAS cucé AaGt AGA AGU? discard} Asparagine ACASS CAUE Hethionine Aut ANIA = CGU? ACG? Aspartic acid sca¢ Phenylalanine vupitt neg UcuF Cysteine SUG Proline coos = caciyt CU? cues# cece Glutamic acid AAG! Gat Serine coug cAGe cae ARG? GGD? cow? CEG? cau? Glutamine Aace? § (aced Threonine ACC AACE# (cced GC? discard) Avc#~ acc? discs rd) Glycine GUG?# GAGE Tryptophan Goux} GcGy GGacr GAU? Histidine ACCH# Tyrosine Auang AUCH ACU? Isoleucine Auuet aaud Yaline quu# ACU? = AGIN? ecu? Jable 1. in this pattem there are 32 psirs having the two purines or the two pyrimidines in a certain position (illustrated as if jn the center). %% includes all 64 possible configueabions of three nuclectides. It eccomiodates 45 of the 49 published wWiplets. Taree are discarded because they are intern- ally inconsistent with the published list (too many amino acids for one trip- let). One is discarded because it is inconsistent with this pattern, 19 remaining triplets ave predicted, as shown. The actual order of the nucleotides is still undetermined, except AUD for tyrosine and GUU for cysteine. *t = Reported by Nirenkerg's group. # = Reported by Ochoa's group. 2 = Predicted by this pattern. “ 2 =. s3. teas, ch wee = cs we yen Bo € a 4& Symmetrical Pacvern in the Genetie Code Reported Predicted Discard AAA Syst}! Cec Prot} GGG Gly WUT Phe%} Leu* AAC Asn Glatt Torey Lys AAG Aref Giut# Lys® AAU Asn# Zlué teretsf CCA Bist?! Prro%! thr? CCG Alat# Arett Prot fhr# CCU Prot? gay Isu GGA Gly? Arg Glu Ging Gec Gly# Ala Ser GGU Gly Trytf Glu WA iutf Lent¢ tyref WUC Lew# Phe# sert# WUE Cys Leutt vel? ACG Alaf Aso ser Gln Met Tar ACU Asn# His Tart Tiu Ser tyr AGO Asp# Glut? Mett# Tiu Lys Try CGU Ala# Aref Ser* Cys Val Asn Zable 2. Entries of table 1 rearranced » to euphasize the number of emino acids reported and predicted for each triplet. Theze are as rany entries for cach triplet es there are possible permutations of the three nucleo- tides ~- one, three, or six. * = Reported by Hirenberg's group. # = Reported by Ochoa's group. & symmetrical necters im the genetie code at nagee naged - ef ten he me a“ APOE ee MAU oop Ly Ae & AGG AGU a acer 2064! ACU nis af ELS yf ue i HAD Ale pas fa wee ys CAF | cacne CGA cage ° B PULET ER. ser ccAY hehe : i ats fy cuaAd the Chics? aS 4 SESS § z i me et COATT GAAS SX GALI bey 2S o gpepsp GAU a CE fe be i a are en : 8S oa oS 4G * Gae eau GCA accede Geof ... GU 2... FR ee eye weigy a fF Gi . eos : ¥ Si quasy A AA# wa = BST as * glu a ra] rere UCA ap = eC UCS val. ENR ° GUC sulk me LEU WA! nS 3 Re Table 3. The purine-pyrimidine paivs in the genetic code, with orier tentatively determined. In one position of each pair, the ivo purines (A and G), and the two pyrimidines (C end U) ave equivalent, reducing the 64 triplets to 32 "codons". Amino acids in capitels have theiz sequences unambiguously determined, assuming that the special position is in the middle. Amino acids in lower esse could possibly belong to the pairs aiagonally opposite, e.g., CAA-gin; ALC-thr, ete. The frequencies of amino acid "alleles" suggest the assigrments indicated. + 2 Reported by Nireaberg's group. yo Reported by Ochea's group.