Bug Bulletin: The GenomeLocPArser error in SplitNCigarReads has been fixed; if you encounter it, use the latest nightly build.

Include no-calls in vcf with only variant sites

caschcasch Posts: 14Member

Is there a way to include only variant sites and no-calls in your final vcf. I know during SNP calls you can only emit variants, or only confident sites or all. However is there a way to reduce your vcf in the end to only variant sites (vsqr passed) and places where no calls could be made. So the end vcfs have only variant sites and missing data - and everything that is not listed in the vcf file is reference. I need such a file for merging with other vcf files - so that every position that is not in the vcfs while merging can be called ref.

So far i have called snps with emit-all and done vsqr - I now want to reduce vcfs in size by excluding NO_VARINATION sites (but want to keep information on "missing" sites)

Best Answer

Answers

  • othoth OsloPosts: 2Member

    I am also interested in obtaining a vcf including only confident varants and sites with missing data. I could not find the recommended workflow for this; could you please direct me to it? I'm working with a haploid genome, and have therefore been using the UnifiedGenotyper followed by VariantFiltation. To obtain data for the missing regions I have up until now relied on grep of a vcf emitting all sites.

  • Geraldine_VdAuweraGeraldine_VdAuwera Posts: 6,176Administrator, GATK Developer admin

    I would recommend using DiagnoseTargets to identify sites that cannot be called.

    Geraldine Van der Auwera, PhD

  • othoth OsloPosts: 2Member

    Thank you for your recommendation. Since my goal is to create a consensus fasta file from the vcf; I am however concerned about the correspondence between missing data from DiagnoseTargets and sites with missing data in the UnifiedGenotyper vcf (I am correct to assume that sites with missing data are marked by ./. in the last (bar) column of the vcf?) I also tried running the DiagnoseTargets tool on my dataset, but got the following error:

    ERROR ------------------------------------------------------------------------------------------
    ERROR stack trace

    java.lang.NullPointerException at org.broadinstitute.sting.gatk.io.stubs.OutputStreamArgumentTypeDescriptor.parse(OutputStreamArgumentTypeDescriptor.java:89) at org.broadinstitute.sting.commandline.ArgumentTypeDescriptor.parse(ArgumentTypeDescriptor.java:129) at org.broadinstitute.sting.commandline.ArgumentSource.parse(ArgumentSource.java:119) at org.broadinstitute.sting.commandline.ParsingEngine.loadValueIntoObject(ParsingEngine.java:488) at org.broadinstitute.sting.commandline.ParsingEngine.loadArgumentsIntoObject(ParsingEngine.java:408) at org.broadinstitute.sting.commandline.ParsingEngine.loadArgumentsIntoObject(ParsingEngine.java:382) at org.broadinstitute.sting.commandline.CommandLineProgram.loadArgumentsIntoObject(CommandLineProgram.java:265) at org.broadinstitute.sting.gatk.CommandLineExecutable.execute(CommandLineExecutable.java:110) at org.broadinstitute.sting.commandline.CommandLineProgram.start(CommandLineProgram.java:248) at org.broadinstitute.sting.commandline.CommandLineProgram.start(CommandLineProgram.java:155) at org.broadinstitute.sting.gatk.CommandLineGATK.main(CommandLineGATK.java:107)

    ERROR ------------------------------------------------------------------------------------------
    ERROR A GATK RUNTIME ERROR has occurred (version 3.1-1-g07a4bf8):
    ERROR
    ERROR This might be a bug. Please check the documentation guide to see if this is a known problem.
    ERROR If not, please post the error message, with stack trace, to the GATK forum.
    ERROR Visit our website and forum for extensive documentation and answers to
    ERROR commonly asked questions http://www.broadinstitute.org/gatk
    ERROR
    ERROR MESSAGE: Code exception (see stack trace for error itself)
    ERROR ------------------------------------------------------------------------------------------
  • Geraldine_VdAuweraGeraldine_VdAuwera Posts: 6,176Administrator, GATK Developer admin

    I am correct to assume that sites with missing data are marked by ./. in the last (bar) column of the vcf?

    Yes that's correct.

    I also tried running the DiagnoseTargets tool on my dataset, but got the following error

    Can you please post the full log output including the starting command line?

    Geraldine Van der Auwera, PhD

Sign In or Register to comment.