The current GATK version is 3.7-0
Examples: Monday, today, last week, Mar 26, 3/26/04

Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!

Get notifications!


You can opt in to receive email notifications, for example when your questions get answered or when there are new announcements, by following the instructions given here.

Did you remember to?


1. Search using the upper-right search box, e.g. using the error message.
2. Try the latest version of tools.
3. Include tool and Java versions.
4. Tell us whether you are following GATK Best Practices.
5. Include relevant details, e.g. platform, DNA- or RNA-Seq, WES (+capture kit) or WGS (PCR-free or PCR+), paired- or single-end, read length, expected average coverage, somatic data, etc.
6. For tool errors, include the error stacktrace as well as the exact command.
7. For format issues, include the result of running ValidateSamFile for BAMs or ValidateVariants for VCFs.
8. For weird results, include an illustrative example, e.g. attach IGV screenshots according to Article#5484.
9. For a seeming variant that is uncalled, include results of following Article#1235.

Did we ask for a bug report?


Then follow instructions in Article#1894.

Formatting tip!


Wrap blocks of code, error messages and BAM/VCF snippets--especially content with hashes (#)--with lines with three backticks ( ``` ) each to make a code block as demonstrated here.

Jump to another community
Picard 2.9.0 is now available. Download and read release notes here.
GATK 3.7 is here! Be sure to read the Version Highlights and optionally the full Release Notes.

RealignerTargetCreator

shaky_dingoshaky_dingo VancouverPosts: 3

I am doing exome sequencing in 700 individuals from a species with a large genome and I would like to use GATK to realign around indels. I am using a reduced reference, which is still about 3Gb. I tested out the target creator, but it is taking 5 days for 12 individuals when each is done individually and this time frame is not feasible for all 700 individuals. I tried to run more in parallel (~30 individuals), but there are RAM limitations on our 250G server. I am currently testing out the program by running all 12 test samples as input for the same run and the time estimate is very long (on the order of several hundred weeks). Based on a preliminary run I have also included a vcf file with likely indels to try and speed the process. Can you suggest another way in which I can make the time frame for all 700 individuals more reasonable? Otherwise we will not be able to use this tool.

Answers

  • Geraldine_VdAuweraGeraldine_VdAuwera Cambridge, MAPosts: 11,736 admin

    Hi there,

    Can you tell me the version you are using and post your command line?

    Are you passing in the list of target intervals for your exomes with -L? If not that should significantly speed up the process. And are you using the multithreading options?

    Geraldine Van der Auwera, PhD

  • shaky_dingoshaky_dingo VancouverPosts: 3

    I am using GenomeAnalysisTK-2.4-9

    This is the command for the individual target creator:
    java -Xmx15g -jar $GATK_PATH/GenomeAnalysisTK.jar -T RealignerTargetCreator -R $REF -nt 10 -o $OUT_PATH/$1.bam.list -I $OUT_PATH/$1.mdup.sort.bam --known $OUT_PATH/INDELS.vcf >> $OUT_PATH/$1.log 2>&1

    I am also trying it with multiple input bams and a great deal more memory.

    We have already reduced the genome to contigs with blast hits to our target sequences (the genome is highly fragmented), however I could try using the locations of the aligned target sequences as the target area is much smaller than the reduced genome. However, we are expecting to recover areas adjacent to the target sequences that might have indels.

    Thanks for your help!

  • Geraldine_VdAuweraGeraldine_VdAuwera Cambridge, MAPosts: 11,736 admin

    If you are concerned about losing information in the areas adjacent to the target sequences, you can pad the intervals. At Broad we typically pad the intervals by 50 bp to capture that information.

    How many contigs do you have in your reference genome? We have seen that genomes with hundreds or more contigs (such as draft genomes) cause a huge inflation in runtimes.

    Geraldine Van der Auwera, PhD

  • shaky_dingoshaky_dingo VancouverPosts: 3

    The genome is a draft and there are hundreds of thousands of contigs in the reduced reference that we are aligning to, so perhaps that is contributing to the long run times. I will try padding the intervals. Thanks again!

  • Geraldine_VdAuweraGeraldine_VdAuwera Cambridge, MAPosts: 11,736 admin

    Ah yes, that explains it. GATK is not designed to process highly fragmented references.

    If using the intervals is not enough, you might try making pseudo-scaffolds of subsets of contigs to reduce the number of contigs the GATK has to handle. It will require that you keep track of the contigs within the scaffolds, but it will certainly help reduce runtimes.

    Geraldine Van der Auwera, PhD

Sign In or Register to comment.