Commit Graph

4 Commits

Author SHA1 Message Date
Teknium d4cec15b47 refactor(tools): first-wave simplification of tools/ (file ops split, lazy_deps, code_exec, approval, browser, delegate, mcp, skills, terminal, voice, media)
Behavior-neutral structural pass over tools/*: god-file extractions into
sibling modules (file_operations_common/lint/search, file_tools_paths/
read_tracking/write, code_execution_env/rpc, tool_search_catalog/names/
validation, tts_command_provider, ...), duplicate helper unification,
if/elif -> dispatch tables, dead-code removal, docstring compaction.
Tool schemas (get_tool_definitions) verified byte-identical to base.
2026-09-02 14:43:45 -07:00
Teknium 3360590115 fix: address second-round SkillEvaluator review feedback
Review feedback from NVIDIA (Nir Paz), minus the LLM items (declined
on the thread: cost-by-default + prompt-injection surface; static-only
also keeps the timeout moot at ~1.5s vs the 120s ceiling):

- Incomplete-validator findings are now PRESERVED as partial evidence;
  only the validator's pass/fail verdict is excluded from the advisory
  verdict. A report with findings from an incomplete check no longer
  reads as clean.
- Clean-report wording is now "no findings from completed checks"
  whenever any validator was incomplete.
- Pinned both scanner binaries to known releases in code comments,
  config guidance, and docs: SkillEvaluator v0.1.0, SkillSpector v2.9.5.
- Tests: 29 (was 28) — partial-evidence preservation flips the old
  discard-pinning test, plus the completed-checks wording case.
2026-08-17 17:04:40 -07:00
Teknium 2c2697b52e feat: widen Tier 1 advisory scan to license + security checks
Review feedback from NVIDIA (Nir Paz): run the full deterministic
Tier 1 surface, not just pii,unicode,lint.

- TIER1_CHECKS now pii,unicode,lint,license,security. License is pure
  static (no measurable cost); security invokes NVIDIA SkillSpector in
  its keyless static-rules mode (~+1.2s per install). schema/quality
  stay excluded: hygiene signal ("author not specified" is
  high-severity upstream), wrong noise for an install prompt.
- SkillSpector is a second optional binary, pinned separately. Absent
  or failing, the security check reports status="incomplete" and the
  adapter treats it as "no opinion" — surfaced as a dim "(not run: ...)"
  note, never as a failure.
- _parse_report derives the verdict from COMPLETED validators only.
  This also absorbs a live upstream inconsistency: SkillEvaluator's
  anti-tamper cross-check on SkillSpector's risk score currently trips
  on moderate-finding skills (fail verdict with zero findings, e.g.
  github-pr-workflow at 15 MEDIUM issues / score 35). Reported to
  NVIDIA separately; either way an evidence-free fail must not render
  as an unexplained failure at install time.
- Dashboard tier1 block gains incomplete_checks.
- Docs: SkillSpector install command + not-run semantics.
- Tests: 28 (was 24) — incomplete-status exclusion, verdict derivation,
  not-run formatting.

E2E against real binaries: clean skill (no findings), skill tripping
the upstream consistency check (passed, "(not run: Security Scan)"),
seeded dirty skill (2 findings, SECRETS row). Full scan cost measured
at ~1.4-1.5s per skill, install-time only.
2026-08-17 17:04:40 -07:00
Teknium 183f18d530 feat: advisory NVIDIA SkillEvaluator Tier 1 scan on skill installs
Adds an optional, advisory second-opinion scan to the skills hub install
path using NVIDIA SkillEvaluator's deterministic, keyless Tier 1 checks
(PII, unicode smuggling, script lint).

- tools/skillevaluator_scan.py: subprocess adapter — runs the scanner
  over the quarantined bundle, parses the JSON report, classifies
  secrets-class findings (private keys, tokens, credentialed connection
  strings) apart from advisory PII findings. Every failure mode
  (binary missing, timeout, crash, bad JSON) degrades to a no-op.
- hermes_cli/skills_hub.py: prints the advisory panel after the built-in
  guard's policy decision and before the install confirmation. Findings
  are shown with file:line; secrets-class findings render red with a
  loud warning. Warn-and-continue by design — the built-in skills guard
  remains the only enforcement layer, because the upstream PII scanner
  has known false-positive classes (git@github.com, docs example
  emails, op:// references).
- hermes_cli/web_routers/skills.py: the dashboard Browse-hub scan
  endpoint returns the same advisory data in a new `tier1` field.
- config: skills.tier1_advisory (default true; no-op without the
  optional scanner binary on PATH).
- docs: user-guide/features/skills.md section with install command and
  config toggle.

Scanner install (optional):
  uv tool install --python 3.13 \
    "skillevaluator @ git+https://github.com/NVIDIA/SkillEvaluator.git"

E2E-validated against the real scanner binary: clean bundled skill (no
findings, "no findings" line), seeded dirty skill (email + credentialed
connection string -> yellow/red panel, install continues), config
disable via real config.yaml (silence). Real scan cost: ~0.2s per skill.
2026-08-17 17:04:40 -07:00