The current GATK version is 3.6-0
Examples: Monday, today, last week, Mar 26, 3/26/04

Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!

Powered by Vanilla. Made with Bootstrap.
Register now for the upcoming GATK Best Practices workshop, Nov 7-8 at the Broad in Cambridge, MA. Open to all comers! More info and signup at

positive and negative set in model training by variantRecalibrator

ying_sheng_1ying_sheng_1 Posts: 64Member


Thanks for develop this tool set and share with others with good supports, it really contains a lot of wonderful tools.

I just try to understand VQSR more into detail. If I give the resources (dbsnp, hap map and 1kg omni data) recommended by the best practice with default settings (which one is in training, which one is the true set ...), Does the positive set contain all variants which recorded in the resources having train=TRUE, but how does the tool select negative set? Does it order the variants from high to low by the QUAL value, and pick up the 5% from the bottom (if the percentBad = 0.05)? Will there be some overlap between positive set and negative set? And is there any quality filtration on the data, e.g. one date point is more than a standard deviation away from average...


Best Answer


  • ying_sheng_1ying_sheng_1 Posts: 64Member

    Thanks Geraldine. I think you answer all my questions.

Sign In or Register to comment.