Filtering Best PracticesSAFE
The largest open-source medical AI skills library for OpenClaw🦞.
Overview
The largest open-source medical AI skills library for OpenClaw🦞.
29f31a89230cOBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
<!-- # COPYRIGHT NOTICE # This file is part of the "Universal Biomedical Skills" project. # Copyright (c) 2026 MD BABU MIA, PhD <[email protected]> # All Rights Reserved. # # This code is proprietary and confidential. # Unauthorized copying of this file, via any medium is strictly prohibited. # # Provenance: Authenticated by MD BABU MIA --> --- name: bio-variant-calling-filtering-best-practices description: Comprehensive variant filtering including GATK VQSR, hard filters, bcftools expressions, and quality metric interpretation for SNPs and indels. Use when filtering variants using GATK best practices. tool_type: mixed primary_tool: bcftools measurable_outcome: Execute skill workflow successfully with valid output within 15 minutes. allowed-tools: - read_file - run_shell_command --- # Variant Filtering Best Practices ## Filter Selection Decision Tree ``` Is dataset large enough for VQSR? (>30 exomes or WGS) ├── Yes → Use VQSR (machine learning) └── No → Use hard filters ├── Germline → GATK recommended thresholds └── Somatic → Caller-specific filters + manual review ``` ## GATK Hard Filter Thresholds ```bash # SNPs gatk VariantFiltration \ -R reference.fa \ -V raw_snps.vcf \ -O filtered_snps.vcf \ --filter-expression "QD < 2.0" --filter-name "QD2" \ --filter-expression "FS > 60.0" --filter-name "FS60" \ --filter-expression "MQ < 40.0" --filter-name "MQ40" \ --filter-expression "MQRankSum < -12.5" --filter-name "MQRankSum-12.5" \ --filter-expression "ReadPosRankSum < -8.0" --filter-name "ReadPosRankSum-8" \ --filter-expression "SOR > 3.0" --filter-name "SOR3" # Indels gatk VariantFiltration \ -R reference.fa \ -V raw_indels.vcf \ -O filtered_indels.vcf \ --filter-expression "QD < 2.0" --filter-name "QD2" \ --filter-expression "FS > 200.0" --filter-name "FS200" \ --filter-expression "ReadPosRankSum < -20.0" --filter-name "ReadPosRankSum-20" \ --filter-expression "SOR > 10.0" --filter-name "SOR10" ``` ## Understanding Quality Metrics | Metric | Meaning | Good Value | |--------|---------|------------| | QD | Quality by Depth | >2 (variant quality normalized by depth) | | FS | Fisher Strand | <60 SNP, <200 indel (strand bias) | | MQ | Mapping Quality | >40 (RMS mapping quality) | | MQRankSum | MQ Rank Sum | >-12.5 (ref vs alt mapping quality) | | ReadPosRankSum | Read Position | >-8 (position in read bias) | | SOR | Strand Odds Ratio | <3 SNP, <10 indel (strand bias) | | DP | Depth | Sample-specific, avoid extremes | | GQ | Genotype Quality | >20 (confidence in genotype) | ## bcftools filter ### Soft vs Hard Filtering ```bash # Hard filter (remove variants) bcftools filter -e 'QUAL<30' input.vcf.gz -o filtered.vcf # Soft filter (mark, don't remove) bcftools filter -s 'LowQual' -e 'QUAL<30' input.vcf.gz -o marked.vcf # Variants failing filter get "LowQual" in FILTER column # Include instead of exclude bcftools filter -i 'QUAL>=30' input.vcf.gz -o filtered.vcf ``` ### Expression Syntax | Operator | Meaning | |----------|---------| | `<`, `<=`, `>`, `>=` | Comparison | | `=`, `==` | Equals | | `!=` | Not equals | | `&&`, `\|\|` | AND, OR | | `!` | NOT | ### Aggregate Functions | Function | Description | |----------|-------------| | `MIN(x)` | Minimum across samples | | `MAX(x)` | Maximum across samples | | `AVG(x)` | Average across samples | | `SUM(x)` | Sum across samples | ### Common bcftools Filters ```bash # Basic quality filter bcftools filter -i 'QUAL>30 && DP>10' input.vcf -o filtered.vcf # Complex filter with multiple metrics bcftools filter -i 'QUAL>30 && INFO/DP>10 && INFO/DP<500 && \ (INFO/FS<60 || INFO/FS=".") && INFO/MQ>40' input.vcf -o filtered.vcf # Genotype-level filters bcftools filter -i 'FMT/DP>10 && FMT/GQ>20' input.vcf -o filtered.vcf # Remove filtered sites bcftools view -f PASS input.vcf -o passed.vcf # Keep only biallelic SNPs bcftools view -m2 -M2 -v snps input.vcf -o biallelic_snps.vcf # Check for missing values bcftools filter -e 'QUAL="."' input.vcf.gz # Exclude missing QUAL bcftools filter -i 'INFO/DP!="."' input.vcf.gz # Include only if DP exists ``` ## bcftools view Filtering ### Filter by Variant Type ```bash # SNPs only bcftools view -v snps input.vcf.gz -o snps.vcf.gz # Indels only bcftools view -v indels input.vcf.gz -o indels.vcf.gz # Exclude SNPs bcftools view -V snps input.vcf.gz -o no_snps.vcf.gz ``` ### Filter by Region ```bash bcftools view -r chr1:1000000-2000000 input.vcf.gz -o region.vcf.gz # Multiple regions bcftools view -r chr1:1000-2000,chr2:3000-4000 input.vcf.gz ``` ### Filter by Samples ```bash # Include samples bcftools view -s sample1,sample2 input.vcf.gz -o subset.vcf.gz # Exclude samples bcftools view -s ^sample3,sample4 input.vcf.gz -o subset.vcf.gz ``` ## Depth Filtering ```bash # Calculate depth percentiles bcftools query -f '%DP\n' input.vcf | \ sort -n | \ awk '{a[NR]=$1} END {print "5th:", a[int(NR*0.05)], "95th:", a[int(NR*0.95)]}' # Filter to middle 90% of depth distribution bcftools filter -i 'INFO/DP>10 && INFO/DP<200' input.vcf -o depth_filtered.vcf ``` ## Allele Frequency Filters ```bash # Minor allele frequency filter (population data) bcftools filter -i 'INFO/AF>0.01 && INFO/AF<0.99' input.vcf -o maf_filtered.vcf # Allele balance for heterozygotes bcftools filter -i 'GT="het" -> (AD[1]/(AD[0]+AD[1]) > 0.2 && AD[1]/(AD[0]+AD[1]) < 0.8)' \ input.vcf -o ab_filtered.vcf ``` ## Region-Based Filtering ```bash # Exclude problematic regions bcftools view -T ^blacklist.bed input.vcf -o filtered.vcf # Keep only exonic regions bcftools view -R exons.bed input.vcf -o exonic.vcf ``` ## Sample-Level Filtering ```bash # Missing genotype rate per sample bcftools stats -s - input.vcf | grep ^PSC | cut -f3,14 # Filter samples with >10% missing bcftools view -S good_samples.txt input.vcf -o sample_filtered.vcf # Filter sites with >5% missing genotypes bcftools filter -i 'F_MISSING<0.05' input.vcf -o sit
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (1)
Gates applied: no_behavioural_pass.
29f31a89230cfull audit observations/trust-audit/skill/freedomintelligence__filtering-best-practices.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 29f31a89230c | SAFE | B | 89 | first audit |
Questions
What does the Filtering Best Practices skill do?
The largest open-source medical AI skills library for OpenClaw🦞.
Is Filtering Best Practices safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Filtering Best Practices access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (29f31a89230c), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.