Files
hermes-agent/skills/productivity/pdf/scripts/pdf_secure.py
T
Teknium 51570f4da7 feat: replace Anthropic office document skills with clean-room MIT implementations
The bundled docx, xlsx, powerpoint, and pdf skills were adapted from
Anthropic's document skills and carried their proprietary LICENSE.txt
(no derivatives, no redistribution). Flagged as critical license
findings by the SkillEvaluator Tier 1 scan of our skill tree.

This replaces all four with clean-room rewrites:

- Authored from scratch against library knowledge only (python-docx,
  openpyxl, python-pptx, pypdf/reportlab/pdfplumber — all MIT/BSD) by
  isolated subagents given functional specs, with an explicit
  prohibition on reading the prior skill content or anthropics/skills;
  session transcripts retained as provenance evidence.
- MIT licensed (LICENSE file per skill), author: Nous Research.
- Each skill: SKILL.md to house standards + argparse helper scripts
  with UTF-8-explicit I/O + its own e2e pytest suite (fixtures built
  on the fly, non-ASCII round-trips run under LC_ALL=C).
- All four pass SkillEvaluator Tier 1 pii+unicode+lint 3/3.

tests/skills/test_office_document_skills.py rewritten against the new
contracts: MIT/no-Anthropic-text invariants, scripts documented in
SKILL.md, argparse CLI shape, and a no-locale-default-open() check
(which caught and fixed a real gap: pdfplumber text reads are fine,
but the invariant scan now guards every future script).

Docs pages regenerated for the four skills (scoped; unrelated
generator drift excluded).

Honest capability deltas vs the old versions are documented per
SKILL.md (e.g. tracked-changes accept/reject and OOXML XSD validation
are not reimplemented; form flattening limits stated).
2026-08-08 10:46:20 -07:00

72 lines
2.6 KiB
Python

#!/usr/bin/env python3
"""Encrypt or decrypt a PDF with passwords (AES-256 via pypdf).
Note: permission flags set at encryption time are advisory — viewers may honor
them, but any PDF library can strip them. Only the user password gates content.
"""
from __future__ import annotations
import argparse
import json
import sys
def main() -> int:
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(encoding="utf-8")
except Exception:
pass
parser = argparse.ArgumentParser(description="Encrypt/decrypt PDFs (pypdf, AES-256).")
parser.add_argument("pdf", help="Input PDF path")
parser.add_argument("-o", "--output", required=True, help="Output PDF path")
mode = parser.add_mutually_exclusive_group(required=True)
mode.add_argument("--encrypt", action="store_true", help="Encrypt the PDF")
mode.add_argument("--decrypt", action="store_true", help="Remove encryption (password required)")
parser.add_argument("--user-password", help="User (open) password for --encrypt")
parser.add_argument("--owner-password", help="Owner password for --encrypt (defaults to user password)")
parser.add_argument("--password", help="Known password for --decrypt")
args = parser.parse_args()
try:
from pypdf import PdfReader, PdfWriter
except ImportError:
print("Missing dependency: install with 'python3 -m pip install pypdf'", file=sys.stderr)
return 2
reader = PdfReader(args.pdf)
if args.encrypt:
if not args.user_password:
print("Error: --encrypt requires --user-password", file=sys.stderr)
return 2
if reader.is_encrypted:
print("Error: input already encrypted; decrypt first", file=sys.stderr)
return 3
writer = PdfWriter()
writer.append(reader)
writer.encrypt(
user_password=args.user_password,
owner_password=args.owner_password or args.user_password,
algorithm="AES-256",
)
action = "encrypted"
else:
if not reader.is_encrypted:
print("Error: input is not encrypted", file=sys.stderr)
return 3
if args.password is None or not reader.decrypt(args.password):
print("Error: wrong or missing --password", file=sys.stderr)
return 4
writer = PdfWriter()
writer.append(reader)
action = "decrypted"
with open(args.output, "wb") as fh:
writer.write(fh)
print(json.dumps({"output": args.output, "action": action, "page_count": len(reader.pages)}))
return 0
if __name__ == "__main__":
sys.exit(main())