
Analysis of DNA-binding motis with CisFinderCisFinder integrates several tool Sfor the analysis of DNA-binding motifs of transcription factors:
CisFinder is based on clustering of over-represented short words (e.g., 8bp) in the DNA sequence. and then, significant elementary motifs are clustered into longer position frequency matrixes (PFMs). PFM is estimated directly from counts of 8-bp words with and without gaps; then PFM is extended over gaps and flanking regions. CisFinder works by several orders faster than MEME suite. All components of CisFinder can also be downloaded, compiled and executed at the command line. Guest login is available for a single-session task. However, registration allows users to keep their profile, uploaded files, and results of analysis for future sessions. Users can access all previous results. To register use the link on the home page. If you have a list of genome coordinates of TF binding sites, you can find DNA motifs over-represented in these regions. We recommend using 200-bp sequences centered at the expected binding site of the TF. If ChIP-chip has a lower spatial resolution, thus the length of sequences can be increased to 300 or even 400 bp. Upload your file using the "Upload New Data" menu. Control files help to reduce the number of false positives in results. Flanking sequences (e.g., 500-1000 bp away from binding sites) can be used as control. If control file is not selected, then a 3rd-order Markov process is used to generate a random control sequence. After files are selected, click the "Identify motifs" button from the 'My CisFinder' tools menu. When results are generated, two buttons "Show elementary motifs" and "Show clusters of motifs" are used to visualize elementary motifs (before clustering) and clusters. Clusters include an aligned set of member elementary motifs. Both files can be saved for future reference: select which file to save, rename it if needed and edit the description. Motifs generated as described above generally have high-quality PFMs. However, matrixes can be further fine-tuned using the iterative resampling method. Any two files with motifs can be compared with each other by similarity (correlation of PWMs). Select the first motif file (e.g., results of motif identification) and click the "Compare motifs" button from the My CisFinder menu. Then you can select another motif file for comparison. Prediction of individual TF binding sites is based on finding locations in the genome sequence that match the given binding motif represented by PFM or PWM. Select the sequence (fasta) file to be searched in the 1st line of data file selection menu and a motif file with PFMs in the 3rd line. Then click the "Search motifs in sequence" button from the My CisFinder menu. There are two ways to control the stringency of motif matching. To control the number of false positive matches per 10 Kb of sequence length. The default is 5 false positive matches per 10 Kb of random sequence. One of the objectives for developing CisFinder was to provide a repository of public data on ChIP-chip and ChIP-seq technologies in mammals. All files used in CisFinder are in plain text format and use tabs for delimiting fields within one line. There is a total of 4 main types of files in the CisFinder: sequences, motifs, search results, and repeats. Motifs can be submitted as PFMs and as degenerate consensus sequences, which are converted to PFM form. Detailed instructions are available in the help file. |