cogneeBLOCK
Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory with small models for free
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
Reproducible, one-command memory-quality benchmarking for cognee. The harness chains four steps — corpus building → question answering → evaluation → dashboard — for a single deterministic config.
The harness ships with cognee, but its heavier dependencies — the HTML dashboard, the DeepEval engine, and some dataset downloads — live in the optional `eval` extra. Core cognee never imports the harness.
Install
pip install "cognee[eval]"
This one extra installs the analysis/dashboard dependencies plus the DeepEval engine — everything needed to run a benchmark end to end. Core cognee never imports the harness, so a plain pip install cognee is unaffected. The extra itself is only needed for the HTML dashboard (plotly), the DeepEval engine (deepeval), and downloading some benchmark datasets (e.g. Musique via gdown): a --engine direct_llm --no-dashboard run works without it. When the dashboard is enabled but its dependencies are missing, the runner fails fast before any pipeline work with an actionable error.
Run a benchmark in one command
CLI:
cognee eval --benchmark HotPotQA --engine direct_llm --limit 5
Or as a module:
python -m cognee.eval_framework --benchmark HotPotQA --engine direct_llm --limit 5
Key flags:
d3d09ecddf62OBSERVED · 2026-09-30Install
Commands as the repository documents them. They are shown, not run.
pip install "cognee[eval]"
uv run python -m cognee.eval_framework.beam.preprocessing.preprocess \
uv run python -m cognee.eval_framework.beam.local_ingest \
uv run python -m cognee.eval_framework.beam.eval.run_sweep \
uv run python -m cognee.eval_framework.beam.eval.aggregate_cross_run \
uv sync --dev --all-extras
Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| cursor | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: cognee
description: >
Use this skill whenever the user asks about Cognee, AI memory, persistent agent memory,
self-improving agents, agents learning from feedback, knowledge graphs, graph-based RAG,
long-term memory for agents, short-term memory for agents, personalization, personas,
temporal search, temporal knowledge graphs, ontology-based extraction, ontology grounding,
feedback, Cypher search, natural-language graph search, chunk search, RAG search, cross-session memory,
session feedback, feedback loops, session based memory, redis based memory, knowledge promotion.
Also use when the user describes the workflow such as:
"turn documents into a knowledge graph", "build memory from files", "search my graph",
"extract entities and relations", "sync data into a graph", "update graph memory",
"store memories for an agent", "help my agent learn over time", "visualize a knowledge
graph built from documents", "let the agent learn", "adaptive agents", "personalized agents",
"session based personalization", "find important ontologies", "find custom pydantic models",
"isolate agentic behaviour", "add permission control to retrieval", "reduce context bloating".
---
# Cognee
Use this skill for **Cognee-specific Python API help** and for mapping user goals to the right
Cognee workflow.
## When to apply this skill
Apply this skill whenever the user wants to do any of the following with Cognee:
- ingest text, files, URLs, repos, or datasets
- build or rebuild a knowledge graph
- search documents, chunks, summaries, triplets, or graph context
- choose a `SearchType`
- enrich an existing graph with `improve`
- define custom graph extraction models or `DataPoint` types
- run custom task pipelines
- configure LLM, graph DB, vector DB, or storage settings
- tag and scope memory with `node_set` / NodeSets
- build persistent memory for agents across sessions
- create feedback loops or self-improving agent workflows
- work with temporal extraction, ontologies, Cypher, or natural-language graph queries
- manage datasets, sessions, feedback, deletion, updates, or visualization
If the user's intent is "store information in memory and query it later", prefer Cognee's
memory API: **remember -> recall** (plus **improve** to enrich and **forget** to delete).
## Core workflow
```python
import cognee
# Store: ingest + build the graph (+ improve, because self_improvement=True by default)
await cognee.remember(
"Your text, file path, URL, or list of inputs",
dataset_name="main",
)
# Query: recall picks a search strategy automatically (rule-based, no LLM call)
results = await cognee.recall("What are the key insights?", datasets=["main"])
for r in results:
print(r.source, r) # each result is tagged "graph" / "session" / ...
```
`remember()` runs `add()` + `cognify()` and then `improve()` underneath. `recall()` wraps
`search()` and adds routing, session memory as a source, and `source`-tagged results.
## Default guidance
When helping with Cognee:
1. Start with the **simplest working path** unless the user explicitly asks for advanced
configuration.
2. Prefer the memory API:
- `remember(...)` to store (ingest + graph build)
- `recall(...)` to query
- `improve(...)` to enrich or index an existing graph, and to bridge sessions into it
- `forget(...)` to delete
3. Treat Cognee APIs as **async**.
4. Use `dataset_name` / `datasets` to keep work organized when the user has multiple sources.
5. Use `node_set` when the user wants lightweight tagging, project scoping, per-user memory
buckets, or subgraph filtering.
6. Pass `query_type=SearchType.X` to `recall()` only when the user needs a specific strategy;
otherwise let it route.
7. Recommend advanced features only when they match the task:
- `session_id` for fast short-term memory and conversation continuity
- `graph_model=` for schema-shaped extraction, `DataPoint` types for direct insertion
- `extractor="gliner_demo"` for LLM-free graph extraction
- custom pipelines for non-default task orchestration
- feedback loops for retrieval improvement
- visualization tools for graph inspection
## When to drop to the low-level operations
`add()`, `cognify()`, `search()` and `memify()` still ship; the memory API calls them. Reach for
them only when `remember`/`recall` cannot express the job:
- `cognify(datasets=["a", "b"])` or `datasets=None` processes several datasets at once;
`remember()` always targets exactly one dataset.
- Rebuilding a graph over data already in the DB (after `forget(memory_only=True)`, or with a
new `graph_model` / ontology) is `cognify()` only; `remember()` always runs `add()` first.
- `search()` takes the agentic extras as first-class parameters (`skills`, `tools`, `max_iter`,
`code_query`, `node_type`) and returns raw `SearchResult` objects instead of tagged dicts.
- `add()` is a staging area: use it to ingest now and `cognify()` later.
- `prune.prune_system(metadata=True)` also drops the relational DB (users, ACLs, dataset
registry); `forget()` never touches those. Use prune only for full test teardown.
`memify()` and `improve()` are the same enrichment pipeline; recommend `improve()`.
`cognee.delete()` is deprecated in favour of `forget()`. Full comparison of the two query
functions: `docs/recall-vs-search.md`.
## Common tasks
### Store data
`remember()` accepts text, file paths, URLs, directories, git repository URLs, binary streams,
or lists of these.
```python
await cognee.remember("notes.md", dataset_name="research")
await cognee.remember("https://example.com", dataset_name="research")
await cognee.remember(["paper.pdf", "summary.txt"], dataset_name="research")
```
Use `node_set` when the user wants data grouped into logical memory buckets.
```python
await cognee.remember(
"Customer prefers concise weekly summaries and Slack delivery.",
dataset_name="customer_success",
node_set=["preferences", "customer_123", "weekly_reports"],Trust audit
BLOCKgrade F · trust 46/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (6 observation(s))
- Network
- declared (7 observation(s))
- Shell
- declared (10 observation(s))
- Dependencies
- pinned
- Secrets in source
- found
Findings (25)
exec(result, mod.__dict__) # noqa: S102 - runs the generated model module source on purpose
connection_string="postgresql://ro_user:pw@host:5432/analytics",
# like postgresql://user:pw@host/db -- the form .env.template itself uses --
beam_existing_ingestion_metrics_conv1_hybrid_completion_20_20_qa_v1_run0.json.gz
beam_existing_ingestion_metrics_conv1_hybrid_completion_20_20_qa_v1_run1.json.gz
beam_existing_ingestion_metrics_conv1_hybrid_completion_20_20_qa_v1_run2.json.gz
beam_existing_ingestion_metrics_conv1_hybrid_completion_20_20_qa_v1_run3.json.gz
beam_json_sessions_metrics_conv0_routed_by_question_type_run0.json.gz
.claude/skills
module = __import__(module_path, fromlist=[class_name])
importlib.import_module(entrypoint)
f"importlib.import_module({entrypoint!r})\n"module = importlib.import_module(module_path)
module = importlib.import_module(_LEGACY_IMPORTS[name])
return tuple(sorted(modes[: random.randint(1, len(modes))], key=lambda mode: mode.label))
logger.info("User %s has forgot their password. Reset token: %s", user.id, token)logger.info("Verification requested for user %s. Verification token: %s", user.id, token)f"✓ Cognee MCP server starting on http://127.0.0.1:{mcp_port}/sse ({mode_info})"logo_url = "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAw8AAAHnCAYAAAD+VGEQAACAAElEQVR4Xuy9B5gsR3X2fyWRMTkbYzCO2MbYBozDB7ZxBoMNNjhgw/+zH38gsHK6ygFlCSQhoYgiSCgghBKSkAhKKCeUkJCEEtLqXkk3zPak3dn691s13d
'{"analytics": {"connection_string": "postgresql://u:p@h/db",'password="feedback_pipeline_noaccess_password",
h = hashlib.md5()
"dlt_manifest_" + hashlib.md5(manifest_text.encode()).hexdigest() + ".txt"
hash_contents = hashlib.md5(data_contents).hexdigest()
hash_contents = hashlib.md5(data_contents).hexdigest()
Gates applied: no_behavioural_pass.
d3d09ecddf62full audit observations/trust-audit/skill/topoteretes__cognee.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-09-30 | d3d09ecddf62 | BLOCK | F | 46 | first audit |
Questions
What does the cognee skill do?
Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory with small models for free
Is cognee safe to install?
No — not without reading the findings first. The audit graded it F (46/100) and found 3 critical or high issues in the source. Each one is listed on this page with the file and line it is on.
What can cognee access on my machine?
The audit observed that it reaches the network, runs shell commands and reads or writes files. Each of those is consistent with what it says it does. Secrets in the source: found — see the findings.
Which assistants does cognee work with?
Its documentation mentions cursor. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (d3d09ecddf62), read on 2026-09-30. The repository is watched, and a new audit runs when it changes — this is the first audit.