WormBase release WS164 now online
"WormBase" <[email protected]> Sat, 30 Sep 2006 20:46:54 -0400
| Newsgroups | gmane.science.biology.wormbase.announce |
|---|---|
| Message-ID | <200610010046.k910ksZ1004442__40776.6450582587$1159663886$gmane$org@brie6.cshl.edu> |
This is an automatic announcement that WormBase
(http://www.wormbase.org) has just been updated. New releases occur
roughly every three weeks.
The text of the AceDB release notes, which contains highlights of the
new data is attached. You can download the full AceDB files from:
ftp://ftp.sanger.ac.uk/pub/wormbase/current_release/
or
ftp://ftp.wormbase.org/pub/wormbase/acedb/current_release
Additional information on the release, including any necessary patches or bug
fixes can be found on the WormBaseWiki:
http://www.wormbase.org/wiki/index.php/WSWS164
New release of WormBase WS164, Wormpep164 and Wormrna164 Wed Sep 6 13:39:15 BST 2006
WS164 was built by Gary Williams
======================================================================
This directory includes:
i) database.WS164.*.tar.gz - compressed data for new release
ii) models.wrm.WS164 - the latest database schema (also in above database files)
iii) CHROMOSOMES/subdir - contains 3 files (DNA, GFF & AGP per chromosome)
iv) WS164-WS163.dbcomp - log file reporting difference from last release
v) wormpep164.tar.gz - full Wormpep distribution corresponding to WS164
vi) wormrna164.tar.gz - latest WormRNA release containing non-coding RNA's in the genome
vii) confirmed_genes.WS164.gz - DNA sequences of all genes confirmed by EST &/or cDNA
viii) cDNA2orf.WS164.gz - Latest set of ORF connections to each cDNA (EST, OST, mRNA)
ix) gene_interpolated_map_positions.WS164.gz - Interpolated map positions for each coding/RNA gene
x) clone_interpolated_map_positions.WS164.gz - Interpolated map positions for each clone
xi) best_blastp_hits.WS164.gz - for each C. elegans WormPep protein, lists Best blastp match to
human, fly, yeast, C. briggsae, and SwissProt & TrEMBL proteins.
xii) best_blastp_hits_brigprot.WS164.gz - for each C. briggsae protein, lists Best blastp match to
human, fly, yeast, C. elegans, and SwissProt & TrEMBL proteins.
xiii) geneIDs.WS164.gz - list of all current gene identifiers with CGC & molecular names (when known)
xiv) PCR_product2gene.WS164.gz - Mappings between PCR products and overlapping Genes
Release notes on the web:
-------------------------
http://www.wormbase.org/wiki/index.php/Release_notes
Genome sequence composition:
----------------------------
WS164 WS163 change
----------------------------------------------
a 32365888 32365888 +0
c 17779857 17779857 +0
g 17756012 17756012 +0
t 32365687 32365687 +0
n 0 0 +0
Total 100267444 100267444 +0
Chromosomal Changes:
--------------------
There are no changes to the chromosome sequences in this release.
Gene data set (Live C.elegans genes 23787)
------------------------------------------
Molecular_info 22066 (92.8%)
Concise_description 4255 (17.9%)
Reference 6323 (26.6%)
CGC_approved Gene name 8851 (37.2%)
RNAi_result 19804 (83.3%)
Microarray_results 19123 (80.4%)
SAGE_transcript 19738 (83%)
Wormpep data set:
----------------------------
There are 20073 CDS in autoace, 23180 when counting 3107 alternate splice forms.
The 23180 sequences contain 10,177,179 base pairs in total.
Modified entries 7
Deleted entries 13
New entries 28
Reappeared entries 1
Net change +16
Status of entries: Confidence level of prediction (based on the amount of transcript evidence)
-------------------------------------------------
Confirmed 7791 (33.6%) Every base of every exon has transcription evidence (mRNA, EST etc.)
Partially_confirmed 10753 (46.4%) Some, but not all exon bases are covered by transcript evidence
Predicted 4636 (20.0%) No transcriptional evidence at all
Status of entries: Protein Accessions
-------------------------------------
UniProtKB/Swiss-Prot accessions 3270 (14.1%)
UniProtKB/TrEMBL accessions 19550 (84.3%)
Status of entries: Protein_ID's in EMBL
---------------------------------------
Protein_id 22820 (98.4%)
Gene <-> CDS,Transcript,Pseudogene connections (cgc-approved)
---------------------------------------------
Entries with CGC-approved Gene name 7185
GeneModel correction progress WS163 -> WS164
-----------------------------------------
Confirmed introns not in a CDS gene model;
+---------+--------+
| Introns | Change |
+---------+--------+
Cambridge | 19 | 2 |
St Louis | 10 | 0 |
+---------+--------+
Members of known repeat families that overlap predicted exons;
+---------+--------+
| Repeats | Change |
+---------+--------+
Cambridge | 6 | 0 |
St Louis | 6 | 0 |
+---------+--------+
Synchronisation with GenBank / EMBL:
------------------------------------
No synchronisation issues
There are no gaps remaining in the genome sequence
---------------
For more info mail [email protected]
-===================================================================================-
New Data:
---------
Mass spectrometry peptide data
Mass spectrometry peptide data from Lukas Reiter at Michael
Hengartner's laboratory in the University of Zürich has been added to
this release. Proteins from various fractions of C.elegans
preparations and from various stages of the life cycle have been
fragmented and the molecular weights of the fragments have then been
used to deduce the sequence of the fragments by comparing to a
database of C. elegans protein sequences using the package Sequest.
The results were statistically verified with PeptideProphet and
ProteinProphet (these are tools developed in the group of R.
Aebersold who is working at the ETH in Zürich)
The peptides were mapped to the genome using the known positions of
the peptides in the proteins and the known positions of the proteins'
genes on the genome.
About 7,500 C. elegans genes have one or more mass spectrometry
peptides mapped to them.
This data gives confirmatory evidence of many exons and (where they
span splice sites) some introns. It gives data on the stage of the
life cycle and location in the organism that proteins are expressed.
COMPARA data
New EnsEMBL COMPARA based predictions are included in WS164 and
replace the older orthologue predictions. As the EnsEMBL COMPARA is
now included in the regular Wormbase pipeline the predictions will be
updated during each following build.
The COMPARA orthologue prediction algorithm uses a combination of
bidirectional best hits and conserved gene order to determine
orthologueous genes and is available as part of the EnsEMBL codebase.
A dump (mySQL) of the EnsEMBL databases used for the predictions is
available from the Sanger ftp-site.
data file: <ftp.sanger.ac.uk/pub2/wormbase/data/compara_164.tar.bz2>
check-sum: <ftp.sanger.ac.uk/pub2/wormbase/data/compara_164.md5>
The homologous genes represent the best reciprocal BLAST hits for the
two species with additional pairs obtained by a combination of BLAST
and location information for more closely related species. These
homologues may therefore potentially represent orthologues. The
Compara gene orthology predictions pipeline works at the protein
sequence level and involves:
We only analyze protein-coding genes, skipping all pseudogene
predictions and non-coding RNA gene types. Due to alternate splicing
of exons, one gene can produce multiple transcripts and hence
multiple translations. In order to provide orthology at the gene
level, one protein sequence must be assigned to the gene. The Ensembl
compara pipeline picks the longest translation. This set of longest
gene translations is now the proteome for each genome, which is
analyzed.
For each genome, the set of longest gene translations are dumped into
FASTA files to be used as blastp databases. We use the Washington
University version 2.0 of BLASTP with the following command line
options: -filter none -span1 -postsw -V=20 -B=20 -sort_by_highscore -
warnings -cpus 1. The -postsw option will perform full Smith-Waterman
alignment of sequences and re-rank the database matches accordingly
prior to output. We run BLASTP with only one query sequence per job
resulting in a large number of BLASTP jobs managed by a processing
system based on autonomous workers the Ensembl hive. The BLASTP
output is then post-filtered so that only hits with an expectation
(E)-value less than 1e-10 are stored in the compara database.
The orthologue prediction algorithm is based on the concept of Best
Reciprocal Hits. The analysis is done on a genome pair basis, for
example a search for human - mouse orthologues. Each gene's longest
translation will likely hit the target genome in multiple locations.
The 'best' hit is the one with the highest BLASTP score, followed by
lowest E-value, followed by highest percent identity, followed by
highest percent positivity. Since the 'best' is a simple sort, the
score takes precedence over the other measures. Due to gene
duplications, although rare, a query translation may align with
identical score, E-value, percent identity, and percent positivity to
more than one target translation. This can then result in 'ties' for
the best position.
For closely related species (i.e. Caenorhabditae), where some gene
order conservation is expected, we identify additional orthologous
pairs obtained by a combination of reciprocal BLAST and location
information. This results in a reciprocal pair, where one direction
is the best hit, but the reverse hit is less than best. To classify
as orthologue the pair must also maintain synteny (conserved gene
order) within 1.5 MB of another orthologue pair.
ncRNA data
Included in this and subsequent WormBase releases is a set of 3672
predicted ncRNA genes that are viewable through the genome browser
<http://wormbase.sanger.ac.uk/db/seq/gbrowse/wormbase/>
This set of genes was predicted using the RNAz programme
<http://www.tbi.univie.ac.at/~wash/RNAz/> and published by K. Missal
et al. J Exp Zoolog B Mol Dev Evol. 2006 Jul 15;306(4):379-92.
<http://wormbase.sanger.ac.uk/db/misc/paper?name=WBPaper00027050;class=Paper>
The objects are not stored in the ACeDB release of the database but
this is a large set of predictions that overlaps with the WormBase
curated ncRNA gene set.
Genome sequence updates:
-----------------------
None.
New Fixes:
----------
None.
Known Problems:
---------------
None.
Other Changes:
--------------
The phenotypic term "No_abnormality_scored," has been made obsolete
effective as of this release (WS164). Instead, "Abnormal" will be
used with the "Not" qualifier instead of "No_abnormality_scored."
"Abnormal" will then be the root term for the entire phenotype
ontology.
Proposed Changes / Forthcoming Data:
-------------------------------------
Model Changes:
------------------------------------
Changed ?Paper.Reference.Year to DateType rather than Int
Existing data that is year as Int (eg 2006) will still be valid
-===================================================================================-
Quick installation guide for UNIX/Linux systems
-----------------------------------------------
1. Create a new directory to contain your copy of WormBase,
e.g. /users/yourname/wormbase
2. Unpack and untar all of the database.*.tar.gz files into
this directory. You will need approximately 2-3 Gb of disk space.
3. Obtain and install a suitable acedb binary for your system
(available from www.acedb.org).
4. Use the acedb 'xace' program to open your database, e.g.
type 'xace /users/yourname/wormbase' at the command prompt.
5. See the acedb website for more information about acedb and
using xace.
____________ END _____________