Commit Graph

6 Commits

Author SHA1 Message Date
Teknium fad88cf130 feat: extend clean-room office skills toward full parity
Same clean-room discipline as the initial rewrite (isolated subagents,
functional specs only, predecessor content banned including via git
history; transcripts retained). All additions test-proven.

docx (13->29 tests):
- docx_revisions.py: tracked changes list/accept/reject (all or by id),
  incl. tables and headers/footers, via direct oxml manipulation
- docx_comments.py: list/add/delete comments (native python-docx >=1.2
  API with XML fallback), anchored-text extraction
- docx_validate.py: package health check (rels, images, styles, CRC)
  with JSON severity report — explicitly not XSD validation
- docx_edit.py: run normalization; TOC + PAGE/NUMPAGES field insertion

xlsx (5->12 tests):
- xlsx_restructure.py: reference-aware insert/delete rows/cols —
  rewrites formulas on all sheets (absolute refs, ranges, cross-sheet,
  quoted names), shifts merges/autofilter/freeze/validation/CF ranges,
  tables, defined names; JSON report incl. honest not_shifted list
- native Excel tables, named ranges, hyperlinks, cell notes,
  sheet protection (documented as strippable, not security)
- xlsx_recalc.py: headless LibreOffice recalc with graceful degrade

powerpoint (11->21 tests):
- pptx_render.py: all slides -> PNGs (soffice + pdftoppm/pdftocairo),
  wired to vision_analyze review loop in SKILL.md
- run-merge normalize before replace (identical-format splits lossless)
- surgical chart ops (series/category/title) wrapping replace_data
- slide duplication with rel remap (clean refusal on chart slides)
- backgrounds, hyperlinks, slide numbers, footers, notes editing

pdf (8->21 tests):
- pdf_make_form.py: JSON spec -> AcroForm (text/checkbox/radio/dropdown)
- pdf_form_layout.py: pre-build layout lint (bounds/overlap/pairing)
  + rendered box overlay for vision_analyze review
- pdf_page_image.py + shared _raster.py: pypdfium2 -> pdftoppm chain,
  graceful degrade; connected to scanned-PDF triage flow
- pdf_stamp.py: text/image stamps at coordinates (rotation/opacity)
- pdf_meta.py: DocInfo metadata + attachments round-trip

Gates re-verified independently: 83 skill tests green under LC_ALL=C,
repo invariant suite 29/29, SkillEvaluator pii+unicode+lint 3/3 x4.
2026-08-08 10:46:20 -07:00
Teknium 51570f4da7 feat: replace Anthropic office document skills with clean-room MIT implementations
The bundled docx, xlsx, powerpoint, and pdf skills were adapted from
Anthropic's document skills and carried their proprietary LICENSE.txt
(no derivatives, no redistribution). Flagged as critical license
findings by the SkillEvaluator Tier 1 scan of our skill tree.

This replaces all four with clean-room rewrites:

- Authored from scratch against library knowledge only (python-docx,
  openpyxl, python-pptx, pypdf/reportlab/pdfplumber — all MIT/BSD) by
  isolated subagents given functional specs, with an explicit
  prohibition on reading the prior skill content or anthropics/skills;
  session transcripts retained as provenance evidence.
- MIT licensed (LICENSE file per skill), author: Nous Research.
- Each skill: SKILL.md to house standards + argparse helper scripts
  with UTF-8-explicit I/O + its own e2e pytest suite (fixtures built
  on the fly, non-ASCII round-trips run under LC_ALL=C).
- All four pass SkillEvaluator Tier 1 pii+unicode+lint 3/3.

tests/skills/test_office_document_skills.py rewritten against the new
contracts: MIT/no-Anthropic-text invariants, scripts documented in
SKILL.md, argparse CLI shape, and a no-locale-default-open() check
(which caught and fixed a real gap: pdfplumber text reads are fine,
but the invariant scan now guards every future script).

Docs pages regenerated for the four skills (scoped; unrelated
generator drift excluded).

Honest capability deltas vs the old versions are documented per
SKILL.md (e.g. tracked-changes accept/reject and OOXML XSD validation
are not reimplemented; form flattening limits stated).
2026-08-08 10:46:20 -07:00
teknium1 75e0d52034 fix(windows): sweep remaining bare read_text/write_text sites + linter rule
AST-driven pass over every Path.read_text()/write_text() without an
explicit encoding= across non-test code: 71 sites in 34 files
(skills_hub, hermes_cli/main+profiles+service_manager+container_boot,
mem0/hindsight/honcho plugins, achievements dashboard, release/CI
scripts, productivity+comfyui skill helpers, agent/*). Verified zero
positional-encoding collisions before insertion; per-file compile()
check after.

Adds a check-windows-footguns rule flagging bare single-line
read_text/write_text (multi-line forms stay covered by the AST guard
test from #38985). Together with the salvaged contributor commits this
retires the ~169-site bare file-I/O class (#37423's long tail).
2026-07-24 17:10:39 -07:00
solyanviktor-star adecb0d1a9 fix(skills): read OOXML parts as bytes and form JSON as UTF-8 in office skill scripts
The bundled office skills (#68595) read user documents and agent-authored
payloads with the locale-default codec:

- docx/powerpoint validators/base.py opened OOXML part XML in text mode
  before handing it to lxml. On Windows (cp1251/GBK) the bytes decode to
  mojibake that lxml then parses, so validation runs against silently
  corrupted document text; on locales where the UTF-8 bytes don't decode
  the validator crashes with UnicodeDecodeError instead of validating.
  Opening as bytes lets lxml honor the encoding declared in the XML prolog.

- The pdf form scripts (fill_fillable_fields, fill_pdf_form_with_annotations,
  create_validation_image, check_bounding_boxes) read the fields JSON the
  agent authors — UTF-8 by construction — with the locale codec, so
  non-ASCII form values (any Cyrillic/CJK/accented input) either crash or
  get written into the user's PDF as mojibake. The json.dump writers use
  ensure_ascii=True and were already safe; only the readers needed pinning.

Adds a contract test asserting every document/payload reader is
locale-independent, plus a live regression test that runs
check_bounding_boxes.py on a non-ASCII fields.json under a forced
non-UTF-8 locale — it fails without the fix on both POSIX (C locale)
and Windows (cp1251 chokes on the 0x98 byte of U+2018).
2026-07-24 17:10:39 -07:00
teknium1 d4b867cf9f fix(windows): sweep remaining unguarded text-mode subprocess sites codebase-wide
AST-driven pass over every subprocess.run/Popen/check_output/check_call/call
with text=True (or universal_newlines=True) and no explicit encoding=:
append encoding='utf-8', errors='replace' at the kwarg site. 136 call
sites across 28 files (cli.py, hermes_cli/main.py, tools_config.py,
environments, computer_use, gateway, scripts, skills helpers, agent/*).

Together with the salvaged #55339/#60741 commits this closes out issue
#53428's bug class; the salvaged #60751 linter rule in
check-windows-footguns.py now enforces it repo-wide (verified: 807 files
scanned, zero findings).
2026-07-24 11:45:57 -07:00
Teknium afb7bf6a5a feat(skills): bundle docx, xlsx, and pdf office skills; refresh powerpoint (#68595)
Non-technical users asking for Word docs, spreadsheets, or PDF work had
no bundled skill coverage — docx/xlsx creation required discovering and
installing hub skills, and PDF manipulation had no skill at all beyond
OCR extraction and nano-pdf edits.

- skills/productivity/docx: create (docx-js), edit (unzip -> XML -> zip),
  tracked changes, comments, validation. Adapted from anthropics/skills.
- skills/productivity/xlsx: openpyxl creation/editing, mandatory
  LibreOffice recalc gate, formula-compatibility rules, financial-model
  conventions. Points at optional excel-author for finance-grade work.
- skills/productivity/pdf: merge/split/rotate/watermark/encrypt, form
  filling (AcroForm + flat overlay scripts), text/table extraction,
  reportlab creation, forms.md + reference.md companions.
- skills/productivity/powerpoint: synced to current upstream pptx skill —
  richer pptxgenjs corruption footguns, template workflow, validate.py +
  validators + thumbnail.py, font-substitution QA guidance; drops the
  stale pack.py/editing.md/pptxgenjs.md workflow files.
- Cross-linked ocr-and-documents, nano-pdf, excel-author via
  related_skills so each office skill routes to its siblings.
- deliverable-mode docs mention the new skills; regenerated per-skill
  docs pages, catalogs, and sidebar.
- tests/skills/test_office_document_skills.py: frontmatter contracts,
  referenced-script existence, schema-map integrity, cross-link
  resolution, script compilation.

E2E validated: docx create->render->edit->validate, xlsx recalc
(SUM + _xlfn.TEXTJOIN evaluate correctly), pdf create->merge->extract,
pptx generate->validate->thumbnail.
2026-07-21 05:39:23 -07:00