If you happen to see a question you know the answer to, please do chime in and help your fellow community members. We encourage our fourm members to be more involved, jump in and help out your fellow researchers with their questions. GATK forum is a community forum and helping each other with using GATK tools and research is the cornerstone of our success as a genomics research community.We appreciate your help!

Test-drive the GATK tools and Best Practices pipelines on Terra

Check out this blog post to learn how you can get started with GATK and try out the pipelines in preconfigured workspaces (with a user-friendly interface!) without having to install anything.

variant QC graphs: bimodal QD

Will_GilksWill_Gilks University of Sussex, UKMember ✭✭


I'm trying to finalise good hard-filtering parameters. Does anyone know why the quality-by-depth distribution having has two peaks (See attached graphs, last column"QD").

This happens even after very strict filtering by basic metrics. There seems to be a lot of variants contributing to the two peaks so I'm guessing it's not due to a particular genomic region. (The graph lines are Drosophila chromosomes. Chr4 in blue is clearly poor. The Inbreeding coefficient and allele frequency are expected to be weird-looking due to our breeding design.)

Filter parameters are
(MQ > 61, MQ < 68, FS < 5, AN > 420, InbreedingCoeff > -1, DP < 10000, DP > 1000, ReadPosRankSum > -1, ReadPosRankSum < 1, ClippingRankSum > -.5, ClippingRankSum < .5, BaseQRankSum > -1, BaseQRankSum < 1, MQRankSum > -.5, MQRankSum < .5, EVENTLENGTH < 1, EVENTLENGTH > -1).




Best Answer


  • Will_GilksWill_Gilks University of Sussex, UKMember ✭✭

    Thanks @Geraldine_VdAuwera If I understand you correctly, one of the peaks should be full of variants which are mostly heterozygous, and the other peak is full of variants that are mostly homozygous. I shall try and plot the quality distribution by heterozygous counts. I feel as though my filtering parameters are very strict as they reduce the number of variants from 2m to 300k. I guess I have to see if these dropped variants are mostly around centromeres or other non-unique sequence to be confident that I'm not throwing out true and accurate ones.

  • Geraldine_VdAuweraGeraldine_VdAuwera Cambridge, MAMember, Administrator, Broadie admin

    That's right.

    What can be helpful if you have a set of known sites is to look at the subset of those sites where you have a variant call in your data, and plot their distribution relative to the overall distribution. That will give you an additional view on how appropriate or not your filters may be (if you're filtering out a lot of the known variants, you're being too strict).

Sign In or Register to comment.