Download manual
Transcript
CHAPTER 4. READ MAPPING 25 clc_mapper ... -g 1 -e 3 ... i.e. first set both the deletion and insertion cost to 1 and the set deletion cost back to 3. 4.3.4 Affine gap cost If you believe that your data contains relatively few gaps, though they may be quite long, you can use the -G/--gapopen parameter to introduce an additional penalty for opening the gap in the first place. This would typically be used along with a low per-base gap cost, typically 1. In this scheme the total cost c(λ) of a gap of length λ is given by the following formula: c(λ) = G + gλ (4.1) where the open cost G and the extension cost g are specified using -G and -g, respectively. Note that the affine model reduces to the linear one, when the open cost, G, is set to zero. Using a combination of a relatively high open cost and a low extension (per-base) gap cost, indicates to the mapper that you expect few gaps, but that these may be quite long. You should be careful about using high open costs, if you have lots of sequencing errors in the form of deletions or insertions, as these may be penalized to the point, where you get lots of long unaligned ends. In such cases linear gap cost may be a better choice. Like linear gap costs, affine gap costs can be used asymmetrically. This is done by independently setting the deletion open cost using the -E/--deletionopen parameter. 4.3.5 Alignment mode The mapper supports two different modes of alignment: local (default) and global. The alignment mode is set using the -a/--alignmode parameter. In local mode, -a local, alignments are chosen to be locally optimal, i.e. the mapper looks for the highest scoring alignment of part of the read. That is, unaligned ends are not penalized in any way. In global mode2 , -a global, the mapper is forced to look for the highest scoring alignment of the entire read. That is, unaligned ends are no longer free. Generally, we recommend that you stick to local alignment, as there are several reasons, why you may not want to force the ends of a raw read sequence to align: 1. Read quality deteriorates towards the end of a read, so you may end up aligning noise. 2. If present, untrimmed adapter sequence cannot be aligned in any meaningful way. 3. Structural variation may be present in the sample, so the reference may not have corresponding sequence for all of the read. In general, unaligned ends express a lack of confidence in a portion of the read; information that can be very helpful for several, common downstream analyses. 2 This is sometimes called semi-global alignment, because the alignment is global with respect to the read, but remains local with respect to the reference.