# TOOLS.md - Local Notes

Skills define _how_ tools work. This file is for _your_ specifics — the stuff that's unique to your setup.

## What Goes Here

Things like:

- Camera names and locations
- SSH hosts and aliases
- Preferred voices for TTS
- Speaker/room names
- Device nicknames
- Anything environment-specific

## Examples

```markdown
### Cameras

- living-room → Main area, 180° wide angle
- front-door → Entrance, motion-triggered

### SSH

- home-server → 192.168.1.100, user: admin

### TTS

- Preferred voice: "Nova" (warm, slightly British)
- Default speaker: Kitchen HomePod
```

## Why Separate?

Skills are shared. Your setup is yours. Keeping them apart means you can update skills without losing your notes, and share skills without leaking your infrastructure.

---

## Python / Miniconda

- **Miniconda 安装路径:** `/root/miniconda3`
- **数据采集专用环境:** `data-collector` (Python 3.11.15)
- **激活方式:** `source /root/miniconda3/bin/activate && conda activate data-collector`
- **直接调用 Python:** `/root/miniconda3/envs/data-collector/bin/python`
- **直接调用 Pip:** `/root/miniconda3/envs/data-collector/bin/pip`

**已安装核心库**
- HTTP/网络: requests, httpx, aiohttp, lxml
- 解析: beautifulsoupser4, lxml
- 浏览器自动化: playwright (chromium), selenium, undetected-chromedriver
- 爬虫框架: scrapy, scrapy-playwright
- 工具: fake-useragent, tenacity, retrying, python-dotenv, pyyaml
- 数据处理: pandas, openpyxl
- **综合爬虫框架:** scrapling 0.4.2 (全功能安装，含 fetchers, MCP, shell)
- **AI 驱动爬虫:** crawl4ai 0.8.6 (LLM 友好，RAG/知识库专用)
- **高级反检测爬虫:** botasaurus 4.0.97 (绕过 Cloudflare WAF/Turnstile)
- **Twitter/X 爬虫:** twscrape 0.17.0 (GraphQL + Search API)
- **OSINT 工具库:** OSINT-TOOLS-2025 (GitHub 仓库，已克隆到 tools/ 目录)
- **标准化 OSINT 资源库:** OSINT-for-countries-V2.0 (24+国家标准化资源，已克隆到 tools/ 目录)

## 技术栈说明

### lxml 版本兼容性
- **当前版本:** lxml 6.0.2
- **scrapling 要求:** >=6.0.2 ✅
- **crawl4ai 要求:** ~=5.3 ⚠️ (实际 6.0.2 可正常工作)

## Scrapling CLI

- **安装路径:** `/usr/local/bin/scrapling`
- **版本:** 0.4.2
- **主要命令:**
  - `scrapling extract get` - 简单 GET 请求
  - `scrapling extract fetch` - 浏览器动态获取
  - `scrapling extract stealthy-fetch` - 隐蔽模式（绕过 Cloudflare）
  - `scrapling shell` - 交互式爬虫控制台

## Crawl4AI

- **Python 路径:** `/root/miniconda3/envs/data-collector/bin/python -c "import crawl4ai"`
- **版本:** 0.8.6
- **核心特性:**
  - LLM 就绪的 Markdown 输出
  - 自动 3 层反机器人检测
  - Shadow DOM 展平
  - 专用: RAG、知识库、Agent 数据管道

## Botasaurus

- **Python 路径:** `/root/miniconda3/envs/data-collector/bin/python -c "import botasaurus"`
- **版本:** 4.0.97
- **核心特性:**
  - 绕过 Cloudflare WAF/Turnstile
  - 真实人类鼠标运动模拟
  - 自动代理轮换
  - 专用: 高反爬目标、模拟人类行为

## twscrape

- **Python 路径:** `/root/miniconda3/envs/data-collector/bin/python -c "import twscrape"`
- **版本:** 0.17.0
- **核心特性:**
  - 支持 Search + GraphQL Twitter API
  - 异步并行处理
  - 自动账户切换（速率限制）
  - 专用: X/Twitter 数据采集、舆情监测

---

Add whatever helps you do your job. This is your cheat sheet.
