c49fa88b80
* refactor(skills): shipped-set slim — 15 skills to optional, github six-way merge, pdf absorbs OCR+nano-pdf, channel-gated teams pipeline
Maintainer-directed shipped-skills curation (skills index 1,900 -> ~1,400
tok/call on desktop; every session pays the index, so this is a per-call
diet on all installs):
- optional-skills moves (installable via skills hub, history preserved):
creative comfyui/ascii-art/excalidraw/pretext/sketch/touchdesigner-mcp;
ALL of mlops (huggingface-hub, llama-cpp, serving-llms-vllm,
weights-and-biases, evaluating-llms-harness — subcategory structure
kept); research-paper-writing (55 supporting files, 17.3K-tok load);
openhue; blogwatcher (first taught the cronjob monitor-field watch
pattern + web_extract instead of pre-cron manual workflows)
- DELETED session-librarian (Aug-12 'inspired by Perplexity Computer'
port, never maintainer-intended; session_search covers discovery)
- github: six skills (auth, issues, pr-workflow, issue-to-pr,
code-review, repo-management) merged into ONE software-development/
github skill — routing body + complete per-workflow references;
benbarclay authorship credited; codebase-inspection rides along;
discipline pins from test_github_issue_to_pr_skill.py preserved
against the reference body in the new test_github_skill.py
- pdf absorbs ocr-and-documents + nano-pdf as references/ + scripts
(extract_pymupdf, extract_marker converted to the argparse house
standard its contract test enforces)
- NEW session_platforms frontmatter gate (metadata.hermes): hides a
skill from the index on gateway channels it is not for; fail-open on
unknown platform; teams-meeting-pipeline gated to [teams, cron]
- blocked-page-recovery: research -> new web category; trigger-first
description ('Use when a fetch fails: 403/429, paywall, WAF, bot
wall.') so the model actually reaches for it on blocked fetches
- docs regenerated via generate-skill-docs.py (195 pages); related_skills
swept repo-wide; tests: 1672 passed (2 openclaw failures pre-existing
on clean main, Windows-local)
* chore: ignore .skills_prompt_snapshot.json (local index cache, accidentally committed)
147 lines
6.1 KiB
Markdown
147 lines
6.1 KiB
Markdown
---
|
|
name: blogwatcher
|
|
description: "Monitor blogs and RSS/Atom feeds via blogwatcher-cli tool."
|
|
version: 2.0.0
|
|
author: JulienTant (fork of Hyaxia/blogwatcher)
|
|
license: MIT
|
|
platforms: [linux, macos, windows]
|
|
metadata:
|
|
hermes:
|
|
tags: [RSS, Blogs, Feed-Reader, Monitoring]
|
|
homepage: https://github.com/JulienTant/blogwatcher-cli
|
|
prerequisites:
|
|
commands: [blogwatcher-cli]
|
|
---
|
|
|
|
# Blogwatcher
|
|
|
|
Track blog and RSS/Atom feed updates with the `blogwatcher-cli` tool. Supports automatic feed discovery, HTML scraping fallback, OPML import, and read/unread article management.
|
|
|
|
## Working with Hermes tools (read this first)
|
|
|
|
`blogwatcher-cli` is the feed database; Hermes tools do the automation around it:
|
|
|
|
- **Recurring watch — use the cronjob tool's `monitor` field, not a bare schedule.** `monitor` runs a script each tick and only wakes the agent when output changes: set it to a script that runs `blogwatcher-cli scan >/dev/null 2>&1 && blogwatcher-cli articles` (deterministic output; new articles = changed output = agent wakes with the diff injected). Unchanged ticks cost zero LLM calls. Set `deliver` to route digests to a chat/channel; add `continuity: true` so consecutive digests can dedupe.
|
|
- **Reading an article the user asks about**: `web_extract([url])` on the article URL from `blogwatcher-cli articles` — do not re-scrape by hand.
|
|
- **One-off "watch this page for changes" without feed semantics**: skip this skill; the cronjob tool's `monitor` field accepts an http(s) URL directly.
|
|
- **Company/competitor tracking with analysis and citations**: prefer the `competitor-news-monitor` skill; blogwatcher is the lighter raw-feed layer it can sit on.
|
|
|
|
## Installation
|
|
|
|
Pick one method:
|
|
|
|
- **Go:** `go install github.com/JulienTant/blogwatcher-cli/cmd/blogwatcher-cli@latest`
|
|
- **Docker:** `docker run --rm -v blogwatcher-cli:/data ghcr.io/julientant/blogwatcher-cli`
|
|
- **Binary (Linux amd64):** `curl -sL https://github.com/JulienTant/blogwatcher-cli/releases/latest/download/blogwatcher-cli_linux_amd64.tar.gz | tar xz -C /usr/local/bin blogwatcher-cli`
|
|
- **Binary (Linux arm64):** `curl -sL https://github.com/JulienTant/blogwatcher-cli/releases/latest/download/blogwatcher-cli_linux_arm64.tar.gz | tar xz -C /usr/local/bin blogwatcher-cli`
|
|
- **Binary (macOS Apple Silicon):** `curl -sL https://github.com/JulienTant/blogwatcher-cli/releases/latest/download/blogwatcher-cli_darwin_arm64.tar.gz | tar xz -C /usr/local/bin blogwatcher-cli`
|
|
- **Binary (macOS Intel):** `curl -sL https://github.com/JulienTant/blogwatcher-cli/releases/latest/download/blogwatcher-cli_darwin_amd64.tar.gz | tar xz -C /usr/local/bin blogwatcher-cli`
|
|
|
|
All releases: https://github.com/JulienTant/blogwatcher-cli/releases
|
|
|
|
### Docker with persistent storage
|
|
|
|
By default the database lives at `~/.blogwatcher-cli/blogwatcher-cli.db`. In Docker this is lost on container restart. Use `BLOGWATCHER_DB` or a volume mount to persist it:
|
|
|
|
```bash
|
|
# Named volume (simplest)
|
|
docker run --rm -v blogwatcher-cli:/data -e BLOGWATCHER_DB=/data/blogwatcher-cli.db ghcr.io/julientant/blogwatcher-cli scan
|
|
|
|
# Host bind mount
|
|
docker run --rm -v /path/on/host:/data -e BLOGWATCHER_DB=/data/blogwatcher-cli.db ghcr.io/julientant/blogwatcher-cli scan
|
|
```
|
|
|
|
### Migrating from the original blogwatcher
|
|
|
|
If upgrading from `Hyaxia/blogwatcher`, move your database:
|
|
|
|
```bash
|
|
mv ~/.blogwatcher/blogwatcher.db ~/.blogwatcher-cli/blogwatcher-cli.db
|
|
```
|
|
|
|
The binary name changed from `blogwatcher` to `blogwatcher-cli`.
|
|
|
|
## Common Commands
|
|
|
|
### Managing blogs
|
|
|
|
- Add a blog: `blogwatcher-cli add "My Blog" https://example.com`
|
|
- Add with explicit feed: `blogwatcher-cli add "My Blog" https://example.com --feed-url https://example.com/feed.xml`
|
|
- Add with HTML scraping: `blogwatcher-cli add "My Blog" https://example.com --scrape-selector "article h2 a"`
|
|
- List tracked blogs: `blogwatcher-cli blogs`
|
|
- Remove a blog: `blogwatcher-cli remove "My Blog" --yes`
|
|
- Import from OPML: `blogwatcher-cli import subscriptions.opml`
|
|
|
|
### Scanning and reading
|
|
|
|
- Scan all blogs: `blogwatcher-cli scan`
|
|
- Scan one blog: `blogwatcher-cli scan "My Blog"`
|
|
- List unread articles: `blogwatcher-cli articles`
|
|
- List all articles: `blogwatcher-cli articles --all`
|
|
- Filter by blog: `blogwatcher-cli articles --blog "My Blog"`
|
|
- Filter by category: `blogwatcher-cli articles --category "Engineering"`
|
|
- Mark article read: `blogwatcher-cli read 1`
|
|
- Mark article unread: `blogwatcher-cli unread 1`
|
|
- Mark all read: `blogwatcher-cli read-all`
|
|
- Mark all read for a blog: `blogwatcher-cli read-all --blog "My Blog" --yes`
|
|
|
|
## Environment Variables
|
|
|
|
All flags can be set via environment variables with the `BLOGWATCHER_` prefix:
|
|
|
|
| Variable | Description |
|
|
|---|---|
|
|
| `BLOGWATCHER_DB` | Path to SQLite database file |
|
|
| `BLOGWATCHER_WORKERS` | Number of concurrent scan workers (default: 8) |
|
|
| `BLOGWATCHER_SILENT` | Only output "scan done" when scanning |
|
|
| `BLOGWATCHER_YES` | Skip confirmation prompts |
|
|
| `BLOGWATCHER_CATEGORY` | Default filter for articles by category |
|
|
|
|
## Example Output
|
|
|
|
```
|
|
$ blogwatcher-cli blogs
|
|
Tracked blogs (1):
|
|
|
|
xkcd
|
|
URL: https://xkcd.com
|
|
Feed: https://xkcd.com/atom.xml
|
|
Last scanned: 2026-04-03 10:30
|
|
```
|
|
|
|
```
|
|
$ blogwatcher-cli scan
|
|
Scanning 1 blog(s)...
|
|
|
|
xkcd
|
|
Source: RSS | Found: 4 | New: 4
|
|
|
|
Found 4 new article(s) total!
|
|
```
|
|
|
|
```
|
|
$ blogwatcher-cli articles
|
|
Unread articles (2):
|
|
|
|
[1] [new] Barrel - Part 13
|
|
Blog: xkcd
|
|
URL: https://xkcd.com/3095/
|
|
Published: 2026-04-02
|
|
Categories: Comics, Science
|
|
|
|
[2] [new] Volcano Fact
|
|
Blog: xkcd
|
|
URL: https://xkcd.com/3094/
|
|
Published: 2026-04-01
|
|
Categories: Comics
|
|
```
|
|
|
|
## Notes
|
|
|
|
- Auto-discovers RSS/Atom feeds from blog homepages when no `--feed-url` is provided.
|
|
- Falls back to HTML scraping if RSS fails and `--scrape-selector` is configured.
|
|
- Categories from RSS/Atom feeds are stored and can be used to filter articles.
|
|
- Import blogs in bulk from OPML files exported by Feedly, Inoreader, NewsBlur, etc.
|
|
- Database stored at `~/.blogwatcher-cli/blogwatcher-cli.db` by default (override with `--db` or `BLOGWATCHER_DB`).
|
|
- Use `blogwatcher-cli <command> --help` to discover all flags and options.
|