Knowledge Graph ConstructionSAFE
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Overview
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
e1ba289846fdOBSERVED · 2026-10-08Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| openclaw | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: knowledge-graph-construction
description: "Build research knowledge graphs for literature synthesis and RAG systems"
metadata:
openclaw:
emoji: "🗺"
category: "tools"
subcategory: "knowledge-graph"
keywords: ["knowledge graph", "knowledge modeling", "ontology", "RAG", "retrieval augmented generation"]
source: "N/A"
---
# Knowledge Graph Construction Guide
## Overview
Knowledge graphs (KGs) organize information as networks of entities and relationships, making them powerful tools for research synthesis, literature exploration, and AI-augmented retrieval. In academic contexts, knowledge graphs can represent relationships between papers, authors, methods, datasets, findings, and concepts -- enabling queries like "Which methods have been applied to dataset X?" or "What are the common limitations reported across studies of Y?"
This guide covers building knowledge graphs for research applications: defining schemas (ontologies), extracting entities and relations from text, storing and querying graph data, and integrating knowledge graphs with Retrieval Augmented Generation (RAG) systems for AI-powered research assistants.
Whether you are building a personal research knowledge base, constructing a domain-specific literature graph, or developing a RAG system for an academic chatbot, these patterns provide a solid foundation.
## Knowledge Graph Fundamentals
### Core Components
| Component | Definition | Research Example |
|-----------|-----------|-----------------|
| Entity (Node) | A distinct concept or object | Paper, Author, Method, Dataset |
| Relation (Edge) | A typed connection between entities | "cites", "uses_method", "evaluates_on" |
| Property | An attribute of an entity or relation | Paper.year, Author.affiliation |
| Ontology/Schema | Formal definition of entity and relation types | Research ontology defining valid types |
### Designing a Research Ontology
```yaml
# research_ontology.yaml
entities:
Paper:
properties: [title, year, doi, abstract, venue]
Author:
properties: [name, affiliation, orcid]
Method:
properties: [name, description, category]
Dataset:
properties: [name, domain, size, url]
Finding:
properties: [description, metric, value, significance]
Concept:
properties: [name, definition, domain]
relations:
CITES:
from: Paper
to: Paper
AUTHORED_BY:
from: Paper
to: Author
USES_METHOD:
from: Paper
to: Method
EVALUATES_ON:
from: Paper
to: Dataset
REPORTS_FINDING:
from: Paper
to: Finding
RELATED_TO:
from: Concept
to: Concept
INTRODUCES:
from: Paper
to: Method
```
## Entity and Relation Extraction
### LLM-Based Extraction
Using a large language model to extract structured knowledge from paper abstracts:
```python
import json
from openai import OpenAI
client = OpenAI()
EXTRACTION_PROMPT = """Extract entities and relationships from this research paper abstract.
Return JSON with:
- entities: list of {type, name, properties}
- relations: list of {source, relation, target}
Entity types: Paper, Method, Dataset, Finding, Concept
Relation types: USES_METHOD, EVALUATES_ON, REPORTS_FINDING, RELATED_TO, INTRODUCES
Abstract: {abstract}
Respond ONLY with valid JSON."""
def extract_from_abstract(abstract, paper_title):
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are a research knowledge extraction system."},
{"role": "user", "content": EXTRACTION_PROMPT.format(abstract=abstract)}
],
response_format={"type": "json_object"},
temperature=0
)
result = json.loads(response.choices[0].message.content)
# Add the paper itself as an entity
result['entities'].insert(0, {
'type': 'Paper',
'name': paper_title,
'properties': {'abstract': abstract[:200]}
})
return result
```
### SpaCy + Custom NER for Domain-Specific Extraction
```python
import spacy
from spacy.tokens import Span
nlp = spacy.load("en_core_web_trf")
# Register custom entity types
@spacy.Language.component("research_entities")
def research_entity_component(doc):
# Pattern-based recognition for methods
method_patterns = [
"random forest", "gradient boosting", "neural network",
"transformer", "attention mechanism", "BERT", "GPT",
"convolutional", "recurrent", "GAN"
]
new_ents = list(doc.ents)
for token in doc:
for pattern in method_patterns:
if pattern.lower() in doc[token.i:token.i+3].text.lower():
span = doc.char_span(token.idx, token.idx + len(pattern),
label="METHOD")
if span and span not in new_ents:
new_ents.append(span)
doc.ents = spacy.util.filter_spans(new_ents)
return doc
nlp.add_pipe("research_entities", after="ner")
```
## Graph Storage and Querying
### Neo4j (Production)
```python
from neo4j import GraphDatabase
class ResearchGraph:
def __init__(self, uri, user, password):
self.driver = GraphDatabase.driver(uri, auth=(user, password))
def add_paper(self, paper):
with self.driver.session() as session:
session.run("""
MERGE (p:Paper {doi: $doi})
SET p.title = $title, p.year = $year, p.abstract = $abstract
""", **paper)
def add_citation(self, citing_doi, cited_doi):
with self.driver.session() as session:
session.run("""
MATCH (a:Paper {doi: $citing})
MATCH (b:Paper {doi: $cited})
MERGE (a)-[:CITES]->(b)
""", citing=citing_doi, cited=cited_doi)
def add_method_usage(self, paper_doi, method_name):
with self.driver.session() as session:
session.run("""
MATCH (p:Paper {doi: $doi})
MERGE (m:Method {name: $method})
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__knowledge-graph-construction.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | e1ba289846fd | SAFE | B | 89 | first audit |
Questions
What does the Knowledge Graph Construction skill do?
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Is Knowledge Graph Construction safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Knowledge Graph Construction access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Knowledge Graph Construction work with?
Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.