Atlas / Skills / freedomintelligence / Tooluniverse Variant Analysis

Tooluniverse Variant AnalysisSAFE

skills/freedomintelligence/tooluniverse-variant-analysis

The largest open-source medical AI skills library for OpenClaw🦞.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
—
Stars
3,053
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

Production-ready VCF processing, variant annotation, and mutation analysis for bioinformatics.

Overview

This skill provides comprehensive variant analysis capabilities:

  • VCF Parsing: Pure Python + cyvcf2 parsers supporting VCF 4.x, gzipped, multi-sample files
  • Mutation Classification: Automatic mapping of SnpEff, VEP, and GATK annotations to standardized mutation types
  • Flexible Filtering: VAF, depth, quality, consequence, population frequency, chromosome
  • ToolUniverse Integration: Annotation via MyVariant.info (ClinVar, dbSNP, gnomAD, CADD, SIFT, PolyPhen)
  • BixBench Support: Designed to answer bioinformatics analysis questions about VCF data

Installation

# Required
pip install pandas

# Recommended (faster VCF parsing)
pip install cyvcf2 pysam

# For ToolUniverse annotation
pip install tooluniverse

Quick Start

from python_implementation import (
parse_vcf,
filter_variants,
compute_variant_statistics,
answer_vaf_mutation_fraction,
FilterCriteria,
)

# Parse VCF
vcf_data = parse_vcf("variants.vcf")
print(f"Variants: {len(vcf_data.variants)}, Samples: {vcf_data.samples}")

# Filter
criteria = FilterCriteria(min_vaf=0.1, min_depth=20, pass_only=True)
passing, failing = filter_variants(vcf_data.variants, criteria)

# Statistics
stats = compute_variant_statistics(passing)
print(f"Ti/Tv: {stats['ti_tv_ratio']}")
print(f"Missense: {stats['mutation_types'].get('missense', 0)}")

# Answer BixBench question
result = answer_vaf_mutation_fraction("variants.vcf", max_vaf=0.3, mutation_type="missense")
print(f"Fraction of VAF<0.3 that are missense: {result['fraction']:.4f}")

Files

Read from source at commit 29f31a89230cOBSERVED · 2026-10-08
02

Install

Commands as the repository documents them. They are shown, not run.

pip install pandas
pip install cyvcf2 pysam
pip install tooluniverse
03

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: tooluniverse-variant-analysis
description: Production-ready VCF processing, variant annotation, mutation analysis, and structural variant (SV/CNV) interpretation for bioinformatics questions. Parses VCF files (streaming, large files), classifies mutation types (missense, nonsense, synonymous, frameshift, splice, intronic, intergenic) and structural variants (deletions, duplications, inversions, translocations), applies VAF/depth/quality/consequence filters, annotates with ClinVar/dbSNP/gnomAD/CADD via ToolUniverse, interprets SV/CNV clinical significance using ClinGen dosage sensitivity scores, computes variant statistics, and generates reports. Solves questions like "What fraction of variants with VAF < 0.3 are missense?", "How many non-reference variants remain after filtering intronic/intergenic?", "What is the pathogenicity of this deletion affecting BRCA1?", or "Which dosage-sensitive genes overlap this CNV?". Use when processing VCF files, annotating variants, filtering by VAF/depth/consequence, classifying mutations, interpreting structural variants, assessing CNV pathogenicity, comparing cohorts, or answering variant analysis questions.
---

# Variant Analysis and Annotation

Production-ready VCF processing and variant annotation skill combining local bioinformatics computation with ToolUniverse database integration. Designed to answer bioinformatics analysis questions about VCF data, mutation classification, variant filtering, and clinical annotation.

## When to Use This Skill

**Triggers**:
- User provides a VCF file (SNV/indel or SV) and asks questions about its contents
- Questions about variant allele frequency (VAF) filtering
- Mutation type classification queries (missense, nonsense, synonymous, etc.)
- Structural variant interpretation requests (deletions, duplications, CNVs)
- Variant annotation requests (ClinVar, gnomAD, CADD, dbSNP)
- CNV pathogenicity assessment using ClinGen dosage sensitivity
- Cohort comparison questions
- Population frequency filtering (SNVs or SVs)
- Intronic/intergenic variant filtering
- Gene dosage sensitivity queries

**Example Questions**:
- "What fraction of variants with VAF < 0.3 are annotated as missense mutations?"
- "After filtering intronic/intergenic variants, how many non-reference variants remain?"
- "What is the clinical significance of this deletion affecting BRCA1?"
- "Which dosage-sensitive genes overlap this 500kb duplication on chr17?"
- "How many variants have clinical significance annotations?"
- "Compare variant counts between samples"

---

## Core Capabilities

| Capability | Description |
|-----------|-------------|
| **VCF Parsing** | Pure Python + cyvcf2 parsers. VCF 4.x, gzipped, multi-sample, SNV/indel/SV |
| **Mutation Classification** | Maps SO terms, SnpEff ANN, VEP CSQ, GATK Funcotator to standard types |
| **VAF Extraction** | Handles AF, AD, AO/RO, NR/NV, INFO AF formats |
| **Filtering** | VAF, depth, quality, PASS, variant type, mutation type, consequence, chromosome, SV size |
| **Statistics** | Ti/Tv ratio, per-sample VAF/depth stats, mutation type distribution, SV size distribution |
| **Annotation** | MyVariant.info (aggregates ClinVar, dbSNP, gnomAD, CADD, SIFT, PolyPhen) |
| **SV/CNV Analysis** | gnomAD SV population frequencies, DGVa/dbVar known SVs, ClinGen dosage sensitivity |
| **Clinical Interpretation** | ACMG/ClinGen CNV pathogenicity classification using haploinsufficiency/triplosensitivity scores |
| **DataFrame** | Convert to pandas for advanced analytics |
| **Reporting** | Markdown reports with tables and statistics, SV clinical reports |

---

## Workflow Overview

```
Input VCF File (SNVs/indels or SVs)
    |
    v
Phase 1: Parse VCF
    |-- Pure Python parser (any VCF 4.x)
    |-- cyvcf2 parser (faster, C-based)
    |-- Extract: CHROM, POS, REF, ALT, QUAL, FILTER, INFO, FORMAT, samples
    |-- Extract per-sample: GT, VAF, depth
    |-- Extract annotations from INFO (ANN, CSQ, FUNCOTATION)
    |-- Detect variant class: SNV/indel vs SV/CNV
    |
    v
Phase 2: Classify Variants
    |-- Variant type: SNV, INS, DEL, MNV, COMPLEX, SV
    |-- Mutation type: missense, nonsense, synonymous, frameshift, splice, etc.
    |-- Impact: HIGH, MODERATE, LOW, MODIFIER
    |-- SV type: DEL, DUP, INV, BND, CNV (if structural variant)
    |
    v
Phase 3: Apply Filters
    |-- VAF range (min/max)
    |-- Read depth minimum
    |-- Quality threshold
    |-- PASS only
    |-- Variant/mutation type inclusion/exclusion
    |-- Consequence exclusion (intronic, intergenic)
    |-- Population frequency range
    |-- Chromosome selection
    |-- SV size range (for structural variants)
    |
    v
Phase 4: Compute Statistics
    |-- Variant type distribution
    |-- Mutation type distribution
    |-- Impact distribution
    |-- Chromosome distribution
    |-- Ti/Tv ratio (for SNVs)
    |-- Per-sample VAF/depth stats
    |-- Gene mutation counts
    |-- SV size distribution (for structural variants)
    |
    v
Phase 5: Annotate with ToolUniverse (optional)
    |-- MyVariant.info: ClinVar, dbSNP, gnomAD, CADD, SIFT, PolyPhen
    |-- dbSNP: Population frequencies, gene associations
    |-- gnomAD: Population allele frequencies
    |-- Ensembl VEP: Consequence prediction
    |
    v
Phase 6: Generate Report / Answer Question
    |-- Markdown report with tables
    |-- Direct answer to specific question
    |-- DataFrame for downstream analysis
    |
    v
Phase 7: Structural Variant & CNV Analysis (if SV/CNV detected)
    |-- Annotate with gnomAD SV population frequencies
    |-- Query DGVa/dbVar for known SVs (Ensembl)
    |-- Identify affected genes
    |-- Query ClinGen dosage sensitivity (HI/TS scores)
    |-- Classify pathogenicity (Pathogenic/Likely Pathogenic/VUS/Benign)
    |-- Generate SV clinical report with ACMG/ClinGen guidelines
```

---

## Phase Summaries

### Phase 1: VCF Parsing

**Use pandas for**:
- Reading VCF as structured data
- Quick exploratory analysis
- When you need to manipulate 
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (1)

LOWInventory / provenance · inv.hidden_file · CWE-1104
.env.template
.env.template
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 29f31a89230cfull audit observations/trust-audit/skill/freedomintelligence__tooluniverse-variant-analysis.json · Report an issue / request a re-scan
05

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-0829f31a89230cSAFEB89first audit
06

Questions

What does the Tooluniverse Variant Analysis skill do?

The largest open-source medical AI skills library for OpenClaw🦞.

Is Tooluniverse Variant Analysis safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Tooluniverse Variant Analysis access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (29f31a89230c), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement