Web Scraping Ethics GuideSAFE
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Overview
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
e1ba289846fdOBSERVED · 2026-10-08Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| openclaw | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: web-scraping-ethics-guide
description: "Scrape web data ethically and legally for research purposes"
metadata:
openclaw:
emoji: "🌐"
category: "tools"
subcategory: "scraping"
keywords: ["web scraping", "ethical scraping", "robots.txt", "rate limiting", "research data collection", "crawling"]
source: "wentor-research-plugins"
---
# Ethical Web Scraping for Research
A skill for collecting web data ethically and legally for research purposes. Covers robots.txt compliance, rate limiting, legal frameworks, data privacy considerations, and practical scraping techniques that respect website operators and comply with institutional review requirements.
## Ethical Framework
### Principles of Ethical Scraping
```
1. Respect robots.txt and Terms of Service
- Check robots.txt before scraping any site
- Review the site's ToS for explicit prohibitions
- When in doubt, contact the site operator
2. Minimize server impact
- Use rate limiting (1-2 requests per second maximum)
- Scrape during off-peak hours when possible
- Cache responses to avoid redundant requests
- Use conditional requests (If-Modified-Since headers)
3. Collect only what you need
- Define your data requirements before scraping
- Do not scrape personal data without ethical justification
- Anonymize or pseudonymize personal information
4. Attribution and transparency
- Set a descriptive User-Agent header with contact info
- Be prepared to identify yourself if contacted
- Credit data sources in publications
5. Institutional compliance
- Check if your IRB/ethics board requires approval for web data
- Follow your institution's acceptable use policy
- Consider data protection regulations (GDPR, CCPA)
```
## Checking robots.txt
### Parsing Robots.txt
```python
import urllib.request
import urllib.robotparser
def check_robots_txt(base_url: str, target_path: str,
user_agent: str = "*") -> dict:
"""
Check if a URL is allowed by robots.txt.
Args:
base_url: The website's base URL (e.g., 'https://example.com')
target_path: The path you want to scrape (e.g., '/data/papers')
user_agent: Your bot's user agent string
"""
robots_url = f"{base_url}/robots.txt"
rp = urllib.robotparser.RobotFileParser()
rp.set_url(robots_url)
try:
rp.read()
except Exception as e:
return {
"robots_txt_found": False,
"error": str(e),
"recommendation": "Proceed with caution; use conservative rate limiting"
}
full_url = f"{base_url}{target_path}"
allowed = rp.can_fetch(user_agent, full_url)
crawl_delay = rp.crawl_delay(user_agent)
return {
"robots_txt_found": True,
"url_checked": full_url,
"allowed": allowed,
"crawl_delay": crawl_delay or "Not specified (use 1-2 seconds)",
"recommendation": (
"Proceed with specified crawl delay"
if allowed
else "Do NOT scrape this path -- it is disallowed"
)
}
```
## Rate-Limited Scraping
### Respectful Request Pattern
```python
import time
import urllib.request
def scrape_with_rate_limit(urls: list[str],
delay: float = 1.0,
user_agent: str = None) -> list[dict]:
"""
Scrape a list of URLs with rate limiting and proper headers.
Args:
urls: List of URLs to fetch
delay: Seconds to wait between requests
user_agent: Custom user agent string
"""
if user_agent is None:
user_agent = (
"ResearchBot/1.0 (Academic research; "
"contact: [email protected])"
)
results = []
for i, url in enumerate(urls):
try:
req = urllib.request.Request(url, headers={
"User-Agent": user_agent,
"Accept": "text/html",
})
response = urllib.request.urlopen(req, timeout=30)
content = response.read().decode("utf-8", errors="replace")
results.append({
"url": url,
"status": response.status,
"content_length": len(content),
"success": True
})
except Exception as e:
results.append({
"url": url,
"error": str(e),
"success": False
})
# Rate limiting
if i < len(urls) - 1:
time.sleep(delay)
return results
```
## Legal Considerations
### Key Legal Frameworks
```
United States:
- CFAA (Computer Fraud and Abuse Act): Unauthorized access is illegal
- hiQ v. LinkedIn (2022): Scraping public data is generally permissible
- Key question: Is the data publicly accessible without authentication?
European Union:
- GDPR: Personal data requires legal basis for processing
- Database Directive: Protects substantial investment in databases
- Text and Data Mining exception (DSM Directive, Art. 3-4):
Research organizations can mine lawfully accessible content
General guidance:
- Public data is more defensible than data behind login walls
- Scraping that circumvents technical measures is riskier
- Academic fair use / research exceptions vary by jurisdiction
- When in doubt, consult your institution's legal counsel
```
## Research-Specific Considerations
### IRB and Ethics Approval
```python
def assess_irb_requirements(data_type: str,
contains_pii: bool) -> dict:
"""
Assess whether web scraping requires IRB review.
Args:
data_type: Type of data being collected
contains_pii: Whether data includes personally identifiable information
"""
if contains_pii:
return {
"irb_required": "Likely yes",
"rationale": (
"Data that identifies or can re-identify individuals "
"generallyTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__web-scraping-ethics-guide.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | e1ba289846fd | SAFE | B | 89 | first audit |
Questions
What does the Web Scraping Ethics Guide skill do?
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Is Web Scraping Ethics Guide safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Web Scraping Ethics Guide access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Web Scraping Ethics Guide work with?
Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.