Gatk Variant CallingSAFE
The largest open-source medical AI skills library for OpenClaw🦞.
Overview
The largest open-source medical AI skills library for OpenClaw🦞.
29f31a89230cOBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
<!-- # COPYRIGHT NOTICE # This file is part of the "Universal Biomedical Skills" project. # Copyright (c) 2026 MD BABU MIA, PhD <[email protected]> # All Rights Reserved. # # This code is proprietary and confidential. # Unauthorized copying of this file, via any medium is strictly prohibited. # # Provenance: Authenticated by MD BABU MIA --> --- name: bio-gatk-variant-calling description: Variant calling with GATK HaplotypeCaller following best practices. Covers germline SNP/indel calling, GVCF workflow for cohorts, joint genotyping, and variant quality score recalibration (VQSR). Use when calling variants with GATK HaplotypeCaller. tool_type: cli primary_tool: gatk measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command --- # GATK Variant Calling GATK HaplotypeCaller is the gold standard for germline variant calling. This skill covers the GATK Best Practices workflow. ## Prerequisites BAM files should be preprocessed: 1. Mark duplicates 2. Base quality score recalibration (BQSR) - optional but recommended ## Single-Sample Calling ### Basic HaplotypeCaller ```bash gatk HaplotypeCaller \ -R reference.fa \ -I sample.bam \ -O sample.vcf.gz ``` ### With Standard Annotations ```bash gatk HaplotypeCaller \ -R reference.fa \ -I sample.bam \ -O sample.vcf.gz \ -A Coverage \ -A QualByDepth \ -A FisherStrand \ -A StrandOddsRatio \ -A MappingQualityRankSumTest \ -A ReadPosRankSumTest ``` ### Target Intervals (Exome/Panel) ```bash gatk HaplotypeCaller \ -R reference.fa \ -I sample.bam \ -L targets.interval_list \ -O sample.vcf.gz ``` ### Adjust Calling Confidence ```bash gatk HaplotypeCaller \ -R reference.fa \ -I sample.bam \ -O sample.vcf.gz \ --standard-min-confidence-threshold-for-calling 20 ``` ## GVCF Workflow (Recommended for Cohorts) The GVCF workflow enables joint genotyping across samples for better variant calls. ### Step 1: Generate GVCFs per Sample ```bash gatk HaplotypeCaller \ -R reference.fa \ -I sample.bam \ -O sample.g.vcf.gz \ -ERC GVCF ``` ### Step 2: Combine GVCFs (GenomicsDBImport) ```bash # Create sample map file # sample_map.txt: # sample1 /path/to/sample1.g.vcf.gz # sample2 /path/to/sample2.g.vcf.gz gatk GenomicsDBImport \ --genomicsdb-workspace-path genomicsdb \ --sample-name-map sample_map.txt \ -L intervals.interval_list ``` ### Alternative: CombineGVCFs (smaller cohorts) ```bash gatk CombineGVCFs \ -R reference.fa \ -V sample1.g.vcf.gz \ -V sample2.g.vcf.gz \ -V sample3.g.vcf.gz \ -O cohort.g.vcf.gz ``` ### Step 3: Joint Genotyping ```bash # From GenomicsDB gatk GenotypeGVCFs \ -R reference.fa \ -V gendb://genomicsdb \ -O cohort.vcf.gz # From combined GVCF gatk GenotypeGVCFs \ -R reference.fa \ -V cohort.g.vcf.gz \ -O cohort.vcf.gz ``` ## Variant Quality Score Recalibration (VQSR) Machine learning-based filtering using known variant sites. Requires many variants (WGS preferred). ### SNP Recalibration ```bash # Build SNP model gatk VariantRecalibrator \ -R reference.fa \ -V cohort.vcf.gz \ --resource:hapmap,known=false,training=true,truth=true,prior=15.0 hapmap.vcf.gz \ --resource:omni,known=false,training=true,truth=false,prior=12.0 omni.vcf.gz \ --resource:1000G,known=false,training=true,truth=false,prior=10.0 1000G.vcf.gz \ --resource:dbsnp,known=true,training=false,truth=false,prior=2.0 dbsnp.vcf.gz \ -an QD -an MQ -an MQRankSum -an ReadPosRankSum -an FS -an SOR \ -mode SNP \ -O snp.recal \ --tranches-file snp.tranches # Apply SNP filter gatk ApplyVQSR \ -R reference.fa \ -V cohort.vcf.gz \ -O cohort.snp_recal.vcf.gz \ --recal-file snp.recal \ --tranches-file snp.tranches \ --truth-sensitivity-filter-level 99.5 \ -mode SNP ``` ### Indel Recalibration ```bash # Build Indel model gatk VariantRecalibrator \ -R reference.fa \ -V cohort.snp_recal.vcf.gz \ --resource:mills,known=false,training=true,truth=true,prior=12.0 Mills.vcf.gz \ --resource:dbsnp,known=true,training=false,truth=false,prior=2.0 dbsnp.vcf.gz \ -an QD -an MQRankSum -an ReadPosRankSum -an FS -an SOR \ -mode INDEL \ --max-gaussians 4 \ -O indel.recal \ --tranches-file indel.tranches # Apply Indel filter gatk ApplyVQSR \ -R reference.fa \ -V cohort.snp_recal.vcf.gz \ -O cohort.vqsr.vcf.gz \ --recal-file indel.recal \ --tranches-file indel.tranches \ --truth-sensitivity-filter-level 99.0 \ -mode INDEL ``` ## Hard Filtering (When VQSR Not Suitable) For small datasets, exomes, or single samples where VQSR fails. ### Extract SNPs and Indels ```bash gatk SelectVariants \ -R reference.fa \ -V cohort.vcf.gz \ --select-type-to-include SNP \ -O snps.vcf.gz gatk SelectVariants \ -R reference.fa \ -V cohort.vcf.gz \ --select-type-to-include INDEL \ -O indels.vcf.gz ``` ### Apply Hard Filters ```bash # Filter SNPs gatk VariantFiltration \ -R reference.fa \ -V snps.vcf.gz \ -O snps.filtered.vcf.gz \ --filter-expression "QD < 2.0" --filter-name "QD2" \ --filter-expression "FS > 60.0" --filter-name "FS60" \ --filter-expression "MQ < 40.0" --filter-name "MQ40" \ --filter-expression "MQRankSum < -12.5" --filter-name "MQRankSum-12.5" \ --filter-expression "ReadPosRankSum < -8.0" --filter-name "ReadPosRankSum-8" \ --filter-expression "SOR > 3.0" --filter-name "SOR3" # Filter Indels gatk VariantFiltration \ -R reference.fa \ -V indels.vcf.gz \ -O indels.filtered.vcf.gz \ --filter-expression "QD < 2.0" --filter-name "QD2" \ --filter-expression "FS > 200.0" --filter-name "FS200" \ --filter-expression "ReadPosRankSum < -20.0" --filter-name "ReadPosRankSum-20" \ --filter-expression "SOR > 10.0" --filter-name "SOR10" ``` #
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (1)
Gates applied: no_behavioural_pass.
29f31a89230cfull audit observations/trust-audit/skill/freedomintelligence__gatk-variant-calling.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 29f31a89230c | SAFE | B | 89 | first audit |
Questions
What does the Gatk Variant Calling skill do?
The largest open-source medical AI skills library for OpenClaw🦞.
Is Gatk Variant Calling safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Gatk Variant Calling access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (29f31a89230c), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.