Analysis of the Escherichia coli Genome: DNA Sequence of the Region from 84.5 to 86.5 Minutes Donna L. Daniels, Guy Plunkett Ill, Valerie Burland, Frederick R. Blattner The DNA sequence of 91.4 kilobases of the Escherichia coli K-12 genome, spanning the region between rmnC at 84.5 minutes and rrnA at 86.5 minutes on the genetic map (85 to 87 percent on the physical map), is described. Analysis of this sequence identified 82 potential coding regions (open reading frames) covering 84 percent of the sequenced interval. The arrangement of these open reading frames, together with the consensus promoter sequences and terminator-like sequences found by computer searches, made it possible to assign them to proposed transcriptional units. More than half the open reading frames correlated with known genes or functions suggested by similarity to other sequences. Those remaining encode still unidentified proteins. The sequenced region also contains several RNA genes and two types of repeated sequence elements were found. Intergenic regions include three “gray holes,” 0.6 to 0.8 kilobases, with no recognizable functions. Complete genomic sequences, including those of viruses, plasmids, organelles, and . FOL of lambda prophage and F factor without treatment by mutagens. Other common lahorarorv strains of E. coli have all been { & -E ua S a4 =o ol Eg o bd ad =) Sj Ela are we Hee wnace phramn. Sy ~ We have chosén the E. coli K-12 strain MG1655 to repres sequencing original The aut! re in the Labor: etics, Univer- Sity of Wisconsin, 44 lenry Malt, Madison, WI 53706. RESEARCH ARTICLES telatively small team of technicians aided by student workers. At this point examina- tion of the sequence data was limited to quality control checks. Ambiguities, where several determinations of an individual nu- cleotide (nt) differed (12), averaged about 1 per 100 nt. A second team, working with computer assistance, conducted the finishing stage (13). Human editing of the computer-gen- erated alignments reduced ambiguities to about 1 in 200 nt and the autoradiograph lanes where data required proofreading were identified. Deferral of proofreading until after initial assembly saved time and re- duced costs. In regions where data te- mained ambiguous, the finishing team re- quested additional data, which could in- volve special treatments, from the data production team. Next, a computer-aided examination for ORF’s, codon usage fre- quencies, and similarities to database en- tries was used to further refine the se- quence. A translated frame could often be distinguished by its codon distribution or by similarity of its predicted amino acid se- quence to a known protein. The sequence was scrutinized for potential insertion or ° . t t peo. - a“ CleClrOpniuicer wuimvugir a was used to resolve sequence, and autora- diograph films were scanned photoelectri- cally into computers where individual se- quences were merged into overlapping con- tiguous segments (the assembly process). The production stage was effected by a Baw mew ee SCIENCE * VOL. 257 * 7 AUGUST 1992 NEW YORK 10021-6399 THE ROCKEFELLER UNIVERSITY Transcription units were Suggested by the arrangement of genes. To locate pro- moters of transcriptional units, a matrix search derived from in vitro measurements of Moyle et al. (15) was used to obtain consensus sequence matches which were 771