10: 10: 10: ii: ii: 76 5 U41 RRO1685-03 Dartmouth College Bionet Training Schedule - Novice :00-10: 00-10: 30-10: 45-11: 15-11 15 30 45 15 735 August 7, 1986 Introduction - Overview of Bionet Systen: Login, System Commands, Mail, File structure, Databases Break Overview of Programs: GENED - Sequence Data Entry and Editing GEL - Sequencing Gel Management Program SEQ - DNA Sequence Analysis Program PEP - Protein Sequence Analysis Program SIZER/MAP - Restriction Enzyme Fragment Sizing and Mapping CLONER - Recombinant DNA Simulation System QUEST - Database Simiiarity Searching IFIND - Database Similarity Searching 12:00-1:00 Hands on session: 1:00-2:15 2:15-3:15 3:15-3:30 3:30-4:15 4:15-5:00 77 5 U41 RRO1685-03 Novice Training Continued Lunch GENED and GEL Programs SEQ and PEP Programs Break SIZER and MAP Programs QUEST and IFIND Programs 78 5 U41 RRO1685-03 Dartmouth College BIONET Training Schedule - Advanced August 8, 1986 9:00-10:30 GEL - Searching and elimating vector sequences; Semi-automatic vs. automatic merging 10:30-10:45 Break 10:45-12:00 CLONER - Simulation of the construction of pUC9 12:00-1:00 Lunch Hands on session: 1:00-2:30 PEP - Comparison of the Search and Align algorithms for protein sequence homology searching; setting chemical similarity matching for homology searches 2:30-3:15 QUEST - Searching using complex keys 3:15-3:30 Break 3:30-5:00 IFIND - Similarity searching between a trans- dated portion of DNA and a QUEST retrieved portion of the NBRF database 79 5 U41 RRO1685-03 Stanford University Bionet Training Schedule August 27, 1986 The morning session will be geared to the novice user: 9:00-10:15 10:15-10:30 10: 30-12:00 12:00-1:00 1:00-2:15 2:15-3:30 3:30-3:45 3:45-5:00 Introduction Overview of Bionet System: Logging on, System Commands, Mail, BBoards Break Directories, File Structure and Location; Xsearch and Find Lunch GENED - Sequence entry and use of ESEQ editor SEQ - Restriction enzymes site searching Break PEP - Designing probes with PEP; Hybrid protein construction and hydropathicity analysis 10: 10: 12: 80 5 U41 RR01685-03 Stanford University BIONET Training Schedule cont'd August 28, 1986 :00-10:30 Sequence Alignment Algorithms 30-10:45 Break 45-12:00 Database Searches (QUEST/IFIND) 00-1:00 Lunch :00-2:30 GEL - Sequencing Gel Management Program :30-3:45 SIZER/MAP - Restriction Enzyme Fragment Sizing and Mapping :45-4:00 Break :00-5:00 CLONER ~- DNA Cloning Simulation 81 5 U4] RRO1685-03 Stanford University Bionet Training Schedule - cont’d August 29, 1986 9:00-12:00 Advanced Topics Including: QUEST Searching using complex keys IFIND similarity searching using a QUEST retrieved portion of a database. 12:00-1:00 Lunch 1:00-3:00 Editors, Batch Jobs 3:00-5:00 File transfer; up and downloading of files and programs between PC’s and Bionet 82 5 U41 RRO1685-03 VIll. BIONET APPLICATION BIONET™ Dear Researcher: You are invited to apply for access to the BIONET '™ National Computer Resource for molecular biology. Enclosed are a description of BIONET, an application form, and order form for BIONET documentation. The BIONET Resource is a central computer facility serving the computational needs, for both research and communication, of the molecular biology community. The Resource is funded by a five year, cooperative agreement with the Biomedical Research Technology Program, Division of Research Resources, National Institutes of Health. IntelliGenetics'™, Inc. of Mountain View, California will provide the computer facilities, core software, and support. Responsibility for overseeing the Resource rests with a National Advisory Committee (NAC), comprised of Drs. Joshua Lederberg (Chair, Rockefeller), Saul Amarel (Rutgers), Fotis Kafatos (Harvard), Allan Maxam (Harvard Medical School), Thomas Rindfleisch (Stanford), Richard Roberts (Cold Spring Harbor), and Charles Yanofsky (Stanford). The BIONET Resource has three goals: e To provide computational assistance in data analysis and problem solving for molecular biologists and researchers in related fields. e To serve as a focus for development and sharing of new software tools. e To promote collaboration and rapid sharing of information among a national community of scientists. Please read the enclosed User Agreement closely. By signing it, you will be agreeing to adhere to both the letter and the spirit of the guidelines described. Each principal investigator must complete an application to be eligible to use the BIONET Resource. Access cannot be passed on from one principal investigator to another. Each scientist who qualifies for and currently has his or her own source of funding is considered a principal investigator. Please type the information on your application form for legibility and accurate processing. Processing time will take approximately four weeks after receipt of your application. If you are applying from a commercial or foreign organization, be sure that your application contains sufficient supporting material to allow the National Advisory Committee to make its judgements. If your application is approved, we will send you a welcome notification, the "Introduction to BIONET" documentation, and instructions for logging on the BIONET 83 5 U41 RRO1685-03 computer via the UNINET telecommunications network. We will also provide initial on- line training at your convenience. Communications is a critical component of the BIONET Resource. On approval of your application, we will send you information on using Electronic Mail, Electronic Bulletin Boards, and File Transfer programs. These features will allow you to exchange information and ideas instantly with the BIONET staff and other users. An annual fee of $400 is currently being charged to all U.S. users. This fee, which covers a portion of our telecommunication charges for Uninet access, is your total cost for the BIONET Resource. Aside from the manuals, there are no other charges for this service. Foreign investigators, including Canadians, on the BIONET system will be responsible for supporting their own york communications. Sincerely, Mary Lou Warn Administrator BIONET 84 5 U41 RR01685-03 APPLICATION CHECKLIST _ Provided user information (page 1) — —Completed INTENDED USE OF BIONET and current grant support statement (page 3) — —Marked DRR Scientific Classifications (page 4) _ —BIONET User agreement read and signed by Principal Investigator and responsible grant administrative officer (page 6) _ —Filled out documentation order form (if desired) __Copy made for your records. Mail the completed application to: BIONET Application IntelliGenetics, Inc. 1975 El Camino Real West Mountain View, CA 94040 Incomplete applications cannot be processed and will be returned. Please send all inquiries about this application to the above address. Include your name, phone number and application date in all correspondence. Applications are processed once a month. The cut-off date is the 20th. Applications received on or after that date will be processed the following month. 85 5 U41 RRO1685-03 Application Form for the BIONET Resource See reverse for description and eligibility of user classi fications Date of Application: Principal Investigator (full name and title): Affiliation: Department, School and Institution: Mailing Address (Include a Street Address for parcels shipped UPS): Area code and phone number: Applying for Class I, II, II or IV Use?: (See Reverse for more information) Would you like more information on the BIONET Satellite program?: (See Reverse for more information) Type of terminal or terminal emulator to be used: (example: Tektronix 4023 or IBM-PC with VT100 emulator) Type of communications software to be used: (example: KERMIT or Smarterm) (note: KERMIT, a public domain communications software, is available from our lending library. Please indicate if you would like to borrow a disk for copying.) Additional users: List up to 5 BIONET users under your direction. (Each Principal Investigator is allocated a fixed amount of space on the computer and only one user in a group can be logged in at one time.) Highlight the primary contact person for your group if not yourself. NAC acceptance rules require individual qualifying Pl’s to apply separately. Additional Pl’s listed on this application will not be given access. Name Title Phone Position 86 5 U41 RRO1685-03 Criteria for Eligibility The four classes of user status are described below. Most users will be Class I users or IV users. Please call or write if you would like to be considered for Class II or III status. CLASS I: Researchers from academic and non-profit institutions who can demonstrate that they are supported by governmental, philanthropic, or unrestricted institution funds and that their research can be assisted by the resource facilities. Exceptions will be considered on a case-by-case basis. These users will have access to the programs in the Core, Database, and Contributed Libraries, and to the electronic mail and bulletin board facilities. An annual fee of $400 is charged for this access. Pi’s in foreign countries, including Canada, will not pay the subscription fee but must pay their own telecommunication costs. CLASS II: Scientists who wish to participate in developing the BIONET Resource by providing new programs to the community. Acceptance as a Class II user is determined in part by the relevance of their programs. These programs should help achieve the goals described in the cover letter. Class II user must meet the same eligibility requirements as the Class I users. However, they will also receive support from the BIONET staff in developing and making their contributed software accessible to the Bionet community. The Class II user will not be required to pay the subscription fee. Please include a description, in detail, of what you intend to contribute, what support you will need from the resource and how the work will benefit the BIONET community. Also include a list of current publications in the area of intended use (for the past two years only). CLASS III: People responsible for Department, School or Campus-wide computer facilities who wish to provide information about or access to BIONET to the community they serve. These users must provide evidence of their position and responsibilities for providing computer facilities for a local community of scientists with access to BIONET. CLASS IV: Scientists who wish to take advantage of only the electronic communications facilities - electronic mail, bulletin boards, and file transfer programs - will be given restricted access for an annual fee of $100. These users must meet the eligibility requirements of the Class I user. BIONET SATELLITE PROGRAM In addition to the above classes; BIONET, in cooperation with IntelliGenetics, is now able to offer an on-site BIONET package. Utilizing existing Digital Equipment VAX or 2060 computers, or SUN Microsystems, all of the programs, bulletin boards and electronic mail functions would be accessible at your location. Your local scientific community would benefit by having a direct and more powerful access to the resource. Accessing this service requires the purchase of a software license from IntelliGenetics. A special purchase program has been arranged to make it easy for academic institutions to join the BIONET Satellite program. If you are interested, please contact us directly or mark the appropriate response on the reverse. 87 5 U41 RRO1685-03 Intended use of BIONET. Include a Research Title of 80 characters or less, and a Research Abstract with a minimum of 3 lines and a maximum of 350 characters. Class I and I applicants, in addition, should include additional information described in the Criteria for Eligibility on page 2. You may attach a separate sheet if you pre fer. Current grant support in area of intended use. Include each federal grant by Principal Investigator, title, funding institution, grant number and duration of support and a brief (three to ten line) abstract of the research. If funding is from institutional or other unrestricted funds, provide information on sources of funding sufficient for the NAC to determine if conditions for access have been met. If this funding ts scheduled to end within 12 months, state whether a renewal of the same grant/funding ts pending. 88 5 U41 RRO1685-03 Appendix to Instructions - DRR Scientific Classification AXIS I AXIS II Code Resource Material/Research Area Code Research Areas Nos. (Maximum 4 Codes) Nos. (Maximum 4 Codes) 1 Aninals: 30 Aging a. Vertebrates, Mammal 82 Anesthesiology bd. Vertebrates, Non-Mammal 34 Anthropology/Ethnography c. Invertebrates 36 Behavioral Sci/Psychology/Social Sci 2 Biological/Chemical Compounds 38 Bioethics 3 Biomaterials 40 Communication Science 4 Cells 2 Subcellular Material 42 Computer Science & Human Subjects 44 Congenital Defects or Malformations 6 Membrane/Tissue/Isolated Organ 46 Degenerative Disorders 7 Microoganisns: 48 Device Prothesis Intra/Extracorporea a. Bacteria 50 Drug Studies: b. Virus a. Toxic c. Orphan Drugs c. Parasites b. Other d. Other 52 Engineering/Bioengineering 8 Plants/Fungus 54 Environmental Sciences: 9 Technology/Technique Development {54 a. Toxic 10 Other (SPECIFY) b. Other 12 Clinical Trials: 56 Epidemiology a. Multicenter b. Single Center {58 Genetics, Including Metabolic Errors 60 Growth and Development ANATOMICAL SYSTEM/RESEARCH AREAS 62 Health Care Applications 64 Immunology and Allergy 13 Cardiovascular System 66 Infectious Diseases 14 Connective Tissue 68 Information Science i& Endocrine System 70 Instrument Development 16 Gastrointestinal System: 72 Mental Disorders/Psychiatry a. Esophagus 74 Metabolism and Transport: b. Gallbladder a. Carbohydrate c. Intestine b. Electrolyte & Water Balance ad. Liver ce. Enzymes e. Pancreas a. Gases 17 Hematological System e. Hormone 18 Integumentary Systen f. Lipid 19 Lymphatic and Recticulo- @. Nucleic Acid Endothelial System h. Protein & Amino Acid 20 Muscular System 76 Neoplasms/Oncology: 21 Nervous Systen a. Benign 22 Oral/Dental bd. Malignant 23 Reproductive Systen 78 Nutrition 24 Respiratory System 80 Radiology/Radiation Nuclear Medicine: 25 Sensory Systen: a. Ionizing (Xray, Nuclear Reactor) a. Ear b. Non-ionizing (Microwave, Radar) bz. Eye 82 Rehabilitation c. Taste/Smell/Touch 64 Statistics/Mathematics 26 Skeletal Systen 86 Surgery 27 Urinary Systen 68 Transplantation 28 Other (SPECIFY) 90 Trauma 92 Other (SPECIFY) 89 5 U41 RRO1685-03 BIONET™ User Agreement e The BIONET resource will not be used for any commercial purpose which is not specifically identified to and approved by BIONET’s National Advisory Committee (NAC). Any pertinent change in sponsorship, continuity of grant support, or use made of BIONET will be reported promptly to the BIONET Resource Manager. e The NAC will approve all access and will make the final judgment on applications that are questionable in nature, scope, or funding of research. e Standard DEC-2060 facilities for file protection will be available to protect the integrity of your data and programs. e Ownership of data and software developed on or contributed to the Resource will be subject to the guidelines of the Principal Investigator’s institution and granting agency, to which all questions on legal issues should be directed. The BIONET Resource will retain a non- exclusive, royalty-free right to use, by approved BIONET investigators, of the data and executable versions of the software on BIONET. e All investigators granted BIONET access must provide brief annual summaries of research results. The summaries must be included in our annual report of Resource activites to the NIH. Investigators will have sufficient advance notice to prepare the summaries. e All publications that involve use of the Resource must acknowledge the Resource by name and NIH grant number (e.g.: Computer resources used to carry out our atudies were provided by the BIONET'™ National Computer Resource for Molecular Biology, whose funding is provided by the Biomedical Research Technology Program, Division of Research Resources, National Institutes of Health, Grant #1 U41 RR-01685.) Investigators must send three (3) copies of these publications to the Resource Manager. e Access to BIONET will be granted to a Principal Investigator and designated members of his or her research group. Each group will be allocated a fixed amount of disk storage space distributed by the PI and designated associates. Class II users will be granted larger amounts of disk space. e We request that each PI limit access of his or her group to one login to BIONET at a time. Use of the Resource will be carefully monitored by the staff and the NAC. e The BIONET Resource provides only a computer facility and associated services. It does not provide research equipment. The Resource has a smal! fund for fostering collaborations and will use this fund, when no other means are available, to support an effort that will advance the goals of the Resource. 90 5 U41 RRO1685-03 I assume full responsibility for all users listed on this application form and will monitor their compliance to the conditions and restrictions for access to the BIONET Resource. | will inform the BIONET Consultant, (electronic mail address BIONET), by electronic mail, immediately about any changes in this group of users, i.e., departure of an existing user or addition of new staff qualified to use the Resource. | will inform new users of the above mentioned conditions and restrictions. As Principal Investigator of this grant to use the BIONET Resource, I agree, by signing this application, to adhere to all conditions and restrictions for use of the BIONET Resource, as described above and such further regulations as may be issued from time to time by the NIH or the NAC. Signature of Principal Investigator: Date: I have also furnished a copy of this application to the responsible grant administrative officer of my institution, whose name and signature are given below: Name of official: Signature: 91 5 U41 RRO1685-03 IX. ADVERTISMENT FOR BIONET TRAINING SESSION If you are a MOLECULAR BIOLOGIST you may be eligible to join BIONET“.. BIONET SEMINAR AT THE MIAMI WINTER SYMPO- SIUM, WEDNESDAY, FEBRUARY 11 AT SEVEN PM, FOR ALL MOLECULAR BIOLOGISTS. ONE SESSION ONLY. CLASS CONTENT AND LEVEL OF COMPLEXITY WILL BE OETERMINED BY YOUR RESPONSE, AND WILL INCLUDE HANDS-ON USE OF INTELLIGENETICS PROGRAMS. SEATING IS LIMITED. REGISTER EARLY TO AVOID DISAPPOINTMENT. MAIL YOUR REGIS- TRATION TODAY WITH CHECK OR MONEY ORDER FOR $30 TO: BIONETINTELLIGENETICS, 1975 EL CAMINO REAL WEST, MOUNTAIN VIEW, CA 94040-2216, OR CALL 415-965-5576 FOR MORE INFORMATION. BIONET IS THE NATIONAL COMPUTER RESOURCE FOR MOLECULAR BIOLOGY. FUNDED BY A GRANT FROM THE NATIONAL INSTITUTES OF HEALTH. Cp oF copy. and mail with check or money order for $30 10 BIONETAntethiGenetics. 1975 Ei Camino Real West Mountain View. CA 94040-2216 lwestigator Please check topics of interest to you institution D Managing large DNA sequencing projects " DO Restriction mapping toots Address BD Simulation and design of recombinant City State Zip DNA experiments O DNAor protein sequerice opted organization & search methods O New user 0 Advanced user oO DNA of protein sequence analysis D tam interested In hosting a session at my Institution 0 Sequence comparison methods DO YOU KNOW WHAT YOU'RE MISSING? SECOND YEAR OF BIONET IS GREAT SUCCESS!!! ITS TIME TO RENEW YOUR SUBSCRIPTION NOW! BIONET has ensered its third year stronger than ever. In keeping with its projected schedule, the BIONET staff has: © afided wore databases © brought im contributed software © expanded the Bulletin Boards o epgraded existing programs © Geubled the number of talecommeunice tion pers © qitablished « training regres and more... WITHOUT INCREASING THE SUBSCRIPTION FEE. We are excited about the growth and changes fa BIONET end hope yeu are too. For more information on any of the ebovw enhancements to the Reswurce, please give a a call .. or better yet, log-ts and check it ext yourse®l! 92 X. RENEWAL NEWSLETTER 5 U41 RRO1685-03 INTELLIGENETICS BECOMES JOINT VENTURE Equipped with » mouse end windows, Sirategene lets Chics ord oe eo IntelliGenetics from IntelliCorp, Inc., of Mountain View, making ImelliGenetics 8 venture jointly owned by the two companies. EntelliGenetics will continue t market its current line of wolecular biology programs and wil! maintain its editions! emphasis on customer suppori. This relationship with Amoco will provide resources for the development of new software. There are plans to add several pew programs to the software that russ on the SUN workstation , the VAX tminicomputer, the microVAX D , end hee timeshering oytem. ‘The most sophisticated sew software wil be Sureargene , @ geectc engineering workstation based oc @rtificial intelligence tchnology. Molecular biologist: at the Amoco Research Cooter have been working with knowledge engineers at lnieliCorp for the past two yea W& apply onelliCorp’s KEE™ , an The accessibility of DNA information apd the ease aod accuracy of simulations make it possible for scientists © experiment with a much larger number of vectors than Gey would ordinarily use. Swategene vses Al techniques to organize Knowlecige about DNA molecules. This knowledge encompasses both descriptive information and rules for feasoning about cloning experiments. The system contains s reference library of vectors and allows sescerchers enter and Setrieve information from individual and laboratory Woraries of constructions. Surategene is designed to @peraic in conjunction with IneelliGenetics' peckage of amelytic software. The synieto q@urrently rem on a Xerox 1386. ROROOROOROROARNRND es GET YOUR UPDATED INTRODUCTION TO BIONET FREE WITH YOUR SUBSCRIPTION RENEWAL. ACT NOW! ee gonenepananoenaney 93 INTELLIGENETICS ANNOUNCES. . . PC/GENE A Personal Computer Genetic Engineering Environment PC GENE is 2 comprehensive package of moleuclar biology software for microcompurens. ht ennteins almost fifty different programs for analyzing peptides and nucleic ecids. You can use POCGENE as the perfect companion to BIONET, or you can run it independently. Dats can be wansferred efficiently through a modem connection. This allows you t perform large Getabese seraches end snquence comperizcos on BIONET. At the sume time you cen take advantage of FC convenience and graphics capabilities t run a host of different analyses ip your own laboraicry. Tens, BIONET sebecribers can stil! get the speed aad memory of a large computer when they seed i. PC/GENE microcomputer software comes with the ‘same high leve} of eapport you heve come tm expect Supresentatives will offer the same degree of personal service for this new software package. PCIGENE allows biologists with litke computer expericnce to wee the programs productively in s matte: of minutes. The system preaezis a series of choices of analyses that arc expressed ip terms thal biologists use. It is necessary only to click the Mapuse oF press a single key w choose 2 sequence tO mmalyze, to define parameters, or to display the results in a variety of ways. 5 U41 RRO1685-03 ALL UNINET DIAL-UP PHONE NUMBERS CHANGING IN SEPTEMBER Uninet is being ebsorbed into GTE Televet w form US Sprint Telenet This means that the phowe sumbers to access the BIONET Ceuta! resource will change. Additionally, we regret that there will be a light change ip the procedure wsed once you dial up. This change will occur in September. Each Pi will receive a special mailing August with al) the details. Information will also be available on BIONET via the sign-oo banner. As 8 positive benefit of us combined petwork, US Sorin: Telenet will have access gutmbers ip over SO new local dialing areas. A database of access numbers is available ea Telenet to all users. ‘We are working to mrenge 4 two week overlap when both the old and the new ecceas methods will work. The Uninet dia)- wpe will be m service for 6 weeks following Gc change, but BIONET will no! be accessible through them. Please consider this if you will be out of touch with BIONET for 8 month or more. Starting in September, access via the olf UNINET dial-ups will produce an error anessage. We will try to make the transition as smooth as possible. In the event of any problems reaching BIONET electronically, you cap telephone the consultant at 415 324-GENE for sasistance. Some of the anatyaes that PC/GENE performs on peptides are: © Computing best oligonucleotide probe © Predicting entigenic determinants © Searching for peptide subsequences © Comparing sequences using the Needleman Wunech algorithm ¢ Aligning two sequences © Determining secondary struchre exing the Chou and Fassmen or the Garnier waethod « Predicting membrane ansociated alpha helices « Piloting local concentrations of amino acids * Calculating matistics of msage of di and @ipeptides ¢ Photting e protein's hydropethic index Some of the analyses that PC/GENE performs om nucielk ecids are: « Displsying (RNA io a clover leaf configuration * Translating sequences » Trapslating introns and exons using EMBL @atabese annotations Searching for subsequences in pucleic acids » Sewching for coding regions wsing both Ficket's and Shepherd's methods * Finding restriction cies, abering restriction eazymes lists, digesting sequences © Creating 8 restriction cite with a single mutation + Comparing sequences with the Pustell doi watrix method Searching for hairpm loops * Analyzing pucieic acid sequence statistics: codon waage, local bese concentrations, enriched maquences BIONET™ Nationa! Computer Resource for Molecular Biology is Sanded through 8 cooperative agreement with IntelliGenetics, inc. by the Biomedical Research T Research Resources, , Division of Nationa! Institutes of Heath. Grant Number BROS InteRiGenetics, Inc. is located at 1975 El Caminc Real West, Mountain View, CA 96040. Phone 415 965-5575 94 5 U41 RR01685-03 by Nancy Bigham aces’ COME ALIVE! , veer ef the BIONET Resowce is to masa romote collsboration and EVOLUTION, message #12). work of the bhoard leaders, sapid sharing of information ee the bboerds comtain more fanoeg 8 satiooa! comunity of BIONET community heve been eaciting and peninent actenticts. The BIONET galected to be bulletin board faformation than ever before bulletin boards (board) mect leaders. They will provide the Please take time to view the this gou! by providing « balietio boards with the most bboards in you field of Sty sees BRET sacent and pertinent interest and to contribute ——— oe information. Under this vew information w sy of oe ides with others CURRENT BULLETIN BOARD LEADERS oie . ieformation — about reading Gene-Exprension Wiliam Sofer board Dausing the Genomic-Organization Steve Harris oor _ pest 6 moeths, Larry Kedes eee the BIONET Molecular-Evolstion Dan Davison metsages, sce hes Doug Brutlag your Reworce Plant-Molecalar-Biology Renald Sedero! INTRODUCTION esuceurating co Politics Michelle Cimbala TO BIONET wpdaing and RNA-Folding Michael Zuker manual. improving the bulletin boards. ‘There are now 29 bulletin i i - ° beadership, 3 dynamic bulletin meng thon eviees of pe board community is being communications software to an rely barchonge of mee @ticle ebout Fast Fourier faformation, maintaining a NEW Transforms and related ital resource of community TRANSLATION algorithens for sequence mews, and archiving outdated mpalyzis (MOLECULAR- belletins, TABLES IN SEQ by Terry Friedemann VECTORBANK UPDATED By ElienHaukr VectorBank is ImelliGenetics’ collection of maps of common vectors designed for use in the CLONER program. It is easier w use bacause we have added two new files. VECTORBANKLST is a list of all vectors available in VectorBank, and VECTORBANK.TXT is a description of some important features of VectorBank As in previous releases of VectorBank, we have provided Rultiple maps for each vector. The following list of maps for PBR322 illustrates the file naming convention. Eile name Map Description: PBR322_6CUTMAP All six-cutter sites. PBR322_COH MAP All cohesive cutter sites. PBR322_COM.MAP All protwtype sites. PBR322_FLUSH MAP All flush cuter sites. PBR322_UNQ.MAP AB enique cutter siees. Please note that we heve replaced the hyphen in the all games with an enderscore, ie., PBR322-6CUTMAP is now PBR322_6CUT MAP. To obtain a vector map for we in CLONER, follow these procedures: 1. Find the vector you want by looking st VECTORBANKLST. This is a listing of al! the vectors in VectorBank. 2. Ener the CLONER program and LOAD your vector fom VecuxBank. ¥f you are unfamiliar with the CLONER program, work @arough the mtorial on CLONER in your BIONET training ‘manual. A eew additon w SEQ Provides two different ways to aher the codon tables used in SEQ for wansisting a DNA sequence. *You can directly edit the codon table that contains the standard genetic code with your particular codon changes and save those changes © Or, if you use the wanstation tables supplied by the program,you can fwenslaie sequences uring « different genet code without having t0 do my editing. To see exasaples- for yeast mitochondria codon changes with both the new editable eoden lable option amd one of the pew translation options, Jog om and send your request! Bese ee RES 95 EVOLUTIONARY ORIGIN OF HEPATITIS B VIRUS AND RETROVIRUSES by MaryJo Lawler Dr. William Robinson, s Professor at Sunford University and one of the Gent reeewchen w& join BIONET, end Roger Miller, 0 fey were also present ip type C revoviruses. As a result of further analyses in the lab, they gathered addivona) evidence which led them to guggesi thal HBV and vetoviruses have e common 5 U41 RR01685-03 DNA sequences by searching over the Genbank and EMBL Getabeses using the IFIND program. Additional gearching using IFIND and the SEARCH functionality of PEP seeding frame analysis, brydropathicity plou: and secondary structure Dr. Miller mode extensive ese of the electronic mail and bulletin board facilities on BIONET w trade unpublished hepadna virus sequences with several other labs on the sytem. COMPUTERS HELP ANALYZE SHOPE FIBROMA GENE _ by Mayio Lawier Gram McFadden and Chris Upton, working at the University of Alberta, have used the BIONET Resource extensively for their work op the molecule organization of the Shope fibroma gene. They have qubmnined for publication a paper entitled “DNA Sequence Homology between the Terminal Inverted Repeats of Shope Fibroma Virus and op Endogenous Cellular Plasmid Species.” PEP'S DIGEST OPTION DIGEST is a major new addition to the protein analysis fenctions in PEP which is designed w help you smudy proteins by rapid}y simulating the action of a peptide @igestion with proteases or by chemicals. The program provides s list of commonly and sot s0 commonly used proteases and cleavage chemicals. You can add & this bist, or you can create op optirely differen! list. When you add s new protease, DIGEST allows you to place the cleavage site before or after the recognition site. DIGEST also accommodates proteases that cleave si more than one site. Once you are satisfied with the list of proteases, you can feo DIGEST, choosing one or more proteases from the hist. ‘The resulting digestion simulation shows the location, length, and molecula weights of the fragments. ‘There are severa) additional options open to you. You can ask t© sce 0 map of the cleavage sites or see amino acid qremposition data for the fragments. You cam weal apy of the fragments os if they were independent peptides end feo malyze Gem with any of the PEP functions. You can also ask DIGEST to simulate a peptide fingerprint by asking the program w drew a plot of the molecular weight of the As i al) letelliGenctics programs, you can ask for on- Sine help. Before you begin the option, we recommend tha! you vead the introduction after the DIGEST: prompt. In their paper, Dr. McFadden and Dr. Upton Gacuss three research findings and suggest how they correlate: the presence of ap extrachromosomal g@atonomous DNA species, its hybridization to Shope fibroma virus (SFV), and the exclusively for thes computer aalysis. They used the GEL program extensively for sequence ery and assembly, Once wupplied with the sequence, Dr. Upton used the SEQ program & stuly the ipveried vepeat regions of the SFV DNA af & analyz the cytoplasmic: DNA molecules through restriction enzyme, weeding frame, and translation ampalyses. ‘The investigators’ homology comparisons werc done wing the SEARCH waing IFIND over the Genbank end EMBL databanes showed @o additiona) homologous @equences to the inveried Supeat region of SFV, but seveeled similarity between continued on page 6 96 5 U41 RRO1685-03 SIMPLE SEARCHES Predicting Experimental Resutts by Doug Brotleg and Alan When you perform s restriction digestioe of a sewly cloned Engelberg saquvence and electophorese the resulting fragments, CLONER Can bave a great deal of time by quickly and accuraniely The computer operating Prodicting the possible digestion patierns. Ip the following eyterms from which you use brief example we show how you can determine the orientation of inelliGenetics propams your clone. if you vector contains more than one potential provide severe) tools on the fasertico site, you cap use the same procedure to determine into DEC 2060 for rapidly which site you've cloned the insen. esarching unformatied CLONER: load pitecommsn =“ We: ill inser! our Gatabaces and txt files such Sragmen: into pBR322 @ sequence dats files Reading fe PBR322_COM.MAP ... Oo the DEC 2060 the 1. PBR322-COM (4363 N) C ; DEFINITION PLASMID PBR322 famest and simplest wo! is (E.COL] CLONING VECTOR) the FIND program which is CLONER: sew god for looking for one a 6 NEW allows us to ater the restriction information about the Sew patterns in a single file drsert. Yf we had sequence information, we could create a The FIND program has the vasriction map in SEQ and then lead that map into CLONER. SCOPE concept from QUEST in Name for sew map: Brag7 that it allows you w look for a Length of pew map: 1420 puiere fp s line, a peragreph, Topology: Jincat ers page. bi permits a himited Eoter os many sew lines of comments os desired; End with an amount of embiguity but only extra allows you t examine a single > Comains GeneZ fe af o time. Using FIND is 3 amabgous t looking in a Please enter each site same followed by its location. phone book for s person's Finish with a blank extry. same. Site name and cut position(s): pati 1 1480 ‘The simplest and most Site name and cut position(s): bambi 245 750 eommon application of FIND is Site name and cut position(s): geor, 1200 w type FIND WITHIN cline> Site name and cut position(s): IN , (are FragZ is map number 2. azample below) leaving out Losding editor help ext... all the other qualifiers. MapEsit: region continued on page 7 continued on page 6 actin in the word prolactin. FIND shows each line where a hil occurs. ACAACTI ;AMOEBA (A. CASTELLANITI) ACTIN GENE-I. : BOVACT] BOVINE ACTIN MRNA, 5 END. ( BOVACT2 ;BOVINE ACTIN MRNA, 3' END. i BOVPRL Bovine prolactin (pri) mRNA. Since FIND simply searching for a sequence of characters, it will report : hits when that sequence appears in @ larger sequence, iz., it finds a hit on : BOVPRLP] BOVINE PROLACTIN, 5S FLANK AND EXON }. i BOVPRLP2 jBOVINE PROLACTIN, S FLANK AND PARTIAL EXON 2. : CELACTI CAENORHABDITIS ELEGANS (NEMATODE) ACTIN 1 GENE 5' END. CELACTH CAENORHABDITIS ELEGANS (NEMATODE) ACTIN I GENE 5 END. CELACTM CAENORHABDITIS ELEGANS (NEMATODE) ACTIN III GENE 5 END. CELACTIV] SCAENORHABDITIS ELEGANS (NEMATODE) ACTIN IV GENE 5 ENIXSEG }). CELACTIV2 CAENORHABDITIS ELEGANS (NEMATODE) ACTIN IV GENE 5 ENIXSEG 2). CELMYH XCELEGANS MAJOR MYOSIN HEAVY CHAIN (UNC-54 I) GENE, 3' END. Here FIND reports a hit on the second patiern, myosin CELMYUNC ;CELEGANS MAJOR MYOSIN HEAVY CHAIN ISOZYME UNC-54 1 GENE. Uf you wave to ant the name of the file where the sequence is located, you Simply leave out “within line” in the FIND command line. The default scope &s peregraph and that makes is possible to set the file name. find ac in in nih *ACA.NIH >» ACAACTI 7AMOEBA (A. CASTELLANTI) ACTIN GENE-I. © ACARRSBS ;A.CASTELLANI (AMOEBA) 5.88 RIBOSOMAL RNA. ° ACARRSS =ACASTELLANT] (AMOEBA) SS RIBOSOMAL RNA. *BOV.NIH > BOVACT! BOVINE ACTIN MRNA, 5 END. © BOVACT2 OVINE ACTIN MRNA, 3 END. The pointer °>° indicates the line with the matching string of laners 97 PREDICTING continued fom page 5 Name for pew region _GencZ Region boundaries: 140 1320 WAU character. (=-) Polarity (<, Lor >): (e) 2 Region GeneZ is sow on feve) | MapEsit: quit CLONER: lis: 1. PBR322-COM (4363 N) C ; DEFINITION PLASMID PBR322 (E.COLI CLONING VECTOR) 2. PragZ (1480 N) L ; Comtains GeneZ CLONER: insert 2.1 pati We simulate the insertion of FragZ éxito pBR. Name for sew map: (=PBR322-COM-FragZ) Retain comments from PBR322-COM? (Y, N, D, ?, or “) (cCR>~Y) po Retain comments from PragZ? (Y, N, D, 7, or *) (eCR>eY) Roter os many new lines of comments os desired; Ead with an extra . Dh map is ProgZ, insereed_ into pRR ot the pail sie. fR> PBR3Z2-COM-PregZ is map sumber 3. CLONER: adit 2 MapEdit: Dip Flipping the map of the insert will allow us to Simulate 0 fragman: inserted with the veverse @riartation. Area t& iovertal) MapEdit quit CLONER: ingen. 2. L psi We repeat the same insertion but in this map the orientation of the inno is reverned. Nase for sew wap: (=PBR322-COM-FregZ) pixZ-fipped Retain comments from PBR322-COM? (Y, N, D, ?, of *) (=Y) po Retain commenu from FragZ? (Y, N, D, 7, of *) (=Y) po Enter as many sew kines of comments as desired; End with an extra COM-FRAGZ, ‘sR pixZ-flipped is map number 4. To dasermine the orientation of the insert, we simulate a digestion with the anzyme chosen to analyze the clones and see which digestion maiches the experimental! results We could have run simulated digestions with a number of enrymes to see which would give us the most distinct results. CLONER: digesi 3 bambi Enzo Sic Lengh Enonm Sik BanH! (376) 422 BamH! (3858) BanH (4363) 1856 Bank! =) BemH] (32858) 505 BanH! (4363) CLONER: diger 4 bambi Enzyme Sie Lengih Eon Six BemHi (376) 3968 BamHI (4344) Bam) (4945) = 1370 BanH! § (76) BamH) (4344) BamH] (4849) we asad only compare the gel patiern to the two sats of patterns hove to dasermine the eria@iation of the inseri. § U41 RRO1685-03 SHOPE continued from page 4 te extracelula DNA and a family of cellule protease fahdion. Is addition to having access to analytical programs, Dr. Upton is pleased with the opportunity © use electronic mai) communicate with other scientists. Like many othe: BIONET scientists, Dr. Upton compunities. Dr. Upion has weded codon usage lables with enveral other BIONET ecientists and has become one of the community's Macintosh authorities. He is currently working with an investigator fo New York, whom he met @rough interections on BIONET, apd they ae setting tp what they cal) s “personal petwork” for their collective apalytis needs. He foresees using BIONET even more extensively than in the pasl, 15 w 2OKB since he began work in the group. | DID YOU KNOW When you are editing a meld, you can seve your edils in twee different ways. SAVE gels” is the progam default. Editing changes you have made are retained for the Current session only and have ot been written in the pro file. Hf you lose your job either because the computer evaches or because there @e wlecommunications problems, then you will lose Ghose edits. "SAVE files” saves them permanenly . These edits eannc be lost. "Set sutossve on” will automatically save your @diting changes ip the -pro file if you type “Set autotave en” after the "Medit:” prompt after the “Medit:” promp. 5 U41 RRO1685-03 SUBSTRING search routine (compiled 11-Ju-80) ? far help M@o-wima 1k Files to search: i Piles to search: (continued) : x R> ‘Target 1) mynsincCR> Target 2) actincCR> Target 3) Equivalences: 1/ cf R> Gurrent expression: | V 2 Expression: Create PL files? NOV/cCR> Output goes to: * TTY: 7 You can search for more than one patiern. Yeon ie. Lor 2 ia aot oso hit This sends the output io your terminal. Type DEL er RUBOUT Wo abon any particular file search. Searching APE.NIH.&510 Searching GCR.NIH.ES10 Searching HUM.NIH.8510 Searching HUMA.NIH.8510 Searching HUMA) NIH.8510 Searching HUMAC.NIH.8510 When XSEARCH finds a match it displays the file in which the match is docated and the tine in which the hil occws. (HUMACNIH.E530 1.1) {actin} ; DEFINITION HUMAN BETA-ACTIN RELATED PSEUDOGENE H-BET, ‘A-AC-PSI-) SEND. (HUMAC.NIH.8510 1.4) {actin} ; KEYWORDS ACTIN; PROCESSED GENE. (HUMAC.NIH.8510 1.12) {actin} i TITLE STRUCTURE OF TWO HUMAN BETA-ACTIN-RELATED PROCESSED GENES ONE OF (HUMAC.NIH.8510 1.19) {actin} ; SITE 420-1540 += HOMOLOGOUS TO ACTIN READING FRAME (HUMAC.NIH.8510 2.1) {ectin} 3 DEFINITION HUMAN BETA-ACTIN RELATED PSEUDOGENE H-BETA-AC-PSI-] SEND. (HUMAC.NTH.B510 2.4) {actin} ; KEYWORDS ACTIN; PROCESSED GENE. Searching HUMMY Searching HUMTB NIH.8510 Searching HUMTR.NIH.8510 98 SEARCHES com, from page 5 ‘Whew: VectrBank contains a XSEARCH afte: the “@” Core bene Particular vector. You can Prompt XSEARCH is not as t FINDccr> Seach the fe vectorbenk tsi, @oavesient as QUEST in that Geacription of these & hist of ab) the vectobenk you campot COLLECT hits por qualifers). For example, maps. To see if pUC is a0 you contr! the output NIHLLST is @ Geobenk fule that preeent, type FIND WITHIN However, it searches databases Gonlains 8 one line extry for fine puc IN vactorbank Ist 10 tw SO times faster than ech Genbank quence On Many people at ming QUEST and is useful for an Mime appears the snquence QUEST 10 search for simple dmitial acteen if you don't wane and the fir line of enambigvous keys in ‘eed ambiguous bases. Once eommnrcis. Hf you were Sequences or in commenis. A XSEARCH has reported the gearching for a word or two fauch fester and simpler Semes of the files that conlain that you expected t appear in rogram called XSEARCH (sec enact hits you can use QUEST the definition line, you would example below) will allow you © march just these files and search the file NIH.LST as W search daisbases for keys ep COLLECT os print out the shown on page 4. with a0 ambiguous leners. FIND can also determine To run the program, type SEARCHES cone. page 8 M@XSEARCH (HUMTR NIH.B510 1.1) {myosin} Here is a his with another large: pattern ; DEFINITION HUMAN NON-MUSCLE (FIBROBLAST) TROPOMYOSIN GENE. Lines recognized = 190 - Sermg 9 Matches Unrecognized Matches 1) “myosin” 3 0 2) actin" 2S 0 Leter case ignored (‘Ab” = “aB’). Files with vo matches: APE.NIH.8510, GCR.NIH.8510, ... MNKR.NIH.8510. 68 files searched, 63 withou! matches. . ~DONE.. eontisve © start over 99 FINDING NEAR-RECOGNITION SITES WITH QUEST wy uncai QUEST s Nexidility makes i possible to search for 0 great variety of petierns in sequences. For example, QUEST can be weed to design a key to locate sequences of beers the! wre one base sway from being 8 restriction enryme site and that, if Changed, would sot alter the translation of the sequence. The Purpose of Gus search is to locate @ place to introduce 8 Sew recognilion site to easily identify positive clooes. Ie the keys described below we have developed patterns that search for a set of eser-EcoRI sites, bul the same procedure can be used to find any mear-restriction site tha! does sot alter the translation Since we don’ want tp alter the translation we mast ‘This key represeets the MET start codon inmnedistely followed by onc or more wiples. Be onder 201 t alter the wanciation of the sequeace when we alter the single bese that introduces the recognition site, we ‘amma take advantage of the degeneracy of the genetic code. The secognition site for EcoRI is GAATTC. We can make s change in tee third bese of 8 codon. The reading frame értermines Which bese is the degenerate one. The first key is: ATG & (..){1,} & GAGTTC. km this key, the reading frame is such that the fint Gs the firs: base of a codon. The third base was changed from A to G because both of these codons code for Glu. The next key is: ATG & (..){1,} & GAATTT. Jn this key the reading frame is the same as the one above except that we ae making a change im the second codon, Changing the codoe from TTC to TTT, since both code for Phe. However, there is 20 reason to assume this particular seeding frame with regard to the recognition site. Instead of (here being an even sei of triplets between the slart codon and the recognition site, there could be one or two additional bases. For one additions) base the key is: ATG & (..4{1,} & () & GAATCC. Ip this case the reading frame is shifted by one, 90 we want tm book for gaATCc instead of gsATTc, since both ATC and ATT onde for Ie. ‘The final key is: ATG & (...){1,} & (..) & GAACTC. This key assumes the frame shift is 2. For an example of the way to use this key, simply log on and and send your request to BIONET, using the Electronic Mail. EatclliGenetics also maintains a database of key patterns (hat you can use in QUEST to help identify various structural and consensus regions in nucleic acid and protein sequences. - ‘The files are located in the directory. ‘We have collected the following files. Lf you have written a weeful key, we would be delighted to include it in the KeyBank 5 U41 RRO1685-03 SEARCHES con. from page 7 CONTEXT around the hits. XSEARCH first prompy you for “Files to Seach” and you msy fespond with filenames Containing wildcards or indirect filenames (HUM®*.* or @NIH-PRIMATESFLS). Then it prompu you for Tergeu” which are just character stings to search for. If you specify more than one targe! XSEARCH then prompts you for a Boolean relationship Qetween the targets and the Gefaul: is t search for target 1 OR target 2 OR urget 3 OR ... XSEARCH sent asks you for equivalences and the default (ebizined by hitting carriage seurrn) is t equate upper and lower case leuers. If you wish t have XSEARCH search trough sequence information father then comments then you should type the letter A (with NO carriage return') at the “Equivaleoces:” prompt, @nd wheo it asks you for an “equivalence file,” type SEQUENCE.XSE. This file not @uly equates upper and lower Cane, it equates Ts and U's ad causes XSEARCH to ignore Carriage returns, tine feeds, tabs end other punctuation in sequences. Il is equivalen! to SEQUENCE SCOPE mm QUEST. Otherwise XSEARCH works exclusively in LINE SCOPE. Try XSEARCH. Type a ? (with NO carriage return) a! each romp! w find out much more abou: XSEARCH's capabilities and Limitations. rey AAKEY Wentifies codons for antigenic sites. AACOMP.KEY Mentifies codons of complementary strand for antigenic sites AMINOKEY Equates one-letier amino acid code with three-letter code. GENE KEY Identifies open reading frames. KEY). KEY Shows keys from Ques! Help Topic KEY)-EXAMPLE. KEY2.KEY Shows beys from Ques! Help Topic KEY2-EXAMPLE. KEY3.KEY Shows keys from Quest Help Topic KEY3-EXAMPLE. KEY4 KEY Shows keys from Quest Help Topic KEY4- EXAMPLE. NADKEY Mentifies dinucleotide-binding region for peptides. PROMOTER KEY Shows suggested consensus sequences for procaryotic promoters. REST KEY Identifies prototype restriction enzyme recognition sequences. SIGNALKEY Identifies consensus for leader pepude cleavage site. ZDNAKEY Shows poteotial Z-DNA purine-pyrimidine patern. 100 5 U41 RRO1685-03 XI. BULLETIN BOARD LEADER AD HOW TO SAVE $400 AND HELP BRING YOUR FIELD INTO THE COMPUTER AGE The BIONET-NEWS bulletin board has messages posted which describe the variety of uses for the bulletin board system and file transfer facilities. These uses range from having a continuous on-line scientific meeting in your research area to sending manuscripts to colleagues in distant labs. The list could undoubtedly be extended by creative people. (See also HELP MEETINGS.) To encourage expanded use of the communications facilities, particularly the bulletin board network, we are offering a FREE ONE YEAR BIONET SUBSCRIPTION to users who are willing to organize and lead a bulletin board. Bulletin board leaders should be actively engaged in research in the selected area of interest. Leading a bulletin board would involve contributing items of interest to the board, encouraging other people in the research field to participate (leaders should have plenty of contacts!), monitoring incoming messages, archiving dated material, and finally submitting a brief year-end report on the bulletin board activity to BIONET. Renewal of the position would be subject to a yearly review by BIONET. We estimate that the work involved would occupy only a few hours each month, but some responsibilities could be delegated to other lab members. Prospective leaders should submit a proposal via electronic mail to BIONET. The proposal should include a description of the suggested bulletin board along with an estimate of the number and potential activity of participants. The activity estimate could be gathered by e-mail contacts prior to submitting the proposal. The final selection of bulletin board topics and leaders will be made in conjunction with BIONET and its National Advisory Committee. Please contact your BIONET consultant at 415-324-4363 if you have any questions. A list of current bulletin board topics and names of leaders can be obtained by typing HELP BB-LIST after the prompt. Some of the current boards need leaders. However, new topics are especially encouraged. As more molecular biologists and biochemists become computer-literate, participation in the bulletin board system should accelerate. As activity increases, the leadership positions will grow in influence. This is your opportunity to get involved with a new communications medium at its inception! Somebody will eventually lead your research field into the computer age. Why not make it you?