Commit Graph

287 Commits

Author SHA1 Message Date
doryani-agent 2a980fbcbd fix(update): let systemd clients outwait legitimate unit transactions
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.

Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-07 08:20:09 -07:00
Teknium bbbcde8173 fix(process): stop late forks escaping deadline tree cleanup 2026-09-07 08:19:32 -07:00
Teknium a9ef4a7625 fix(codex): keep transport echoes out of durable user history
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.

Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.

Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698

Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
2026-09-07 08:09:57 -07:00
Teknium 026e3e84ea fix: reject unavailable desktop profile and session targets 2026-09-07 06:26:22 -07:00
Teknium 12891ea3bd fix(desktop): scope tool changes and commit rebuild ownership together 2026-09-07 06:26:22 -07:00
Teknium d79575cb89 fix(slack): preserve explicit suspension after timer removal 2026-09-07 06:10:54 -07:00
Teknium 34e512ae58 fix(sessions): remove remaining timer migration and guidance 2026-09-07 06:10:54 -07:00
Teknium 1481a0de96 docs(gateway): describe persistent sessions without reset timers 2026-09-07 06:10:54 -07:00
Teknium 42f8389987 fix: restore persisted assistant replies alongside tools and reasoning 2026-09-07 04:56:22 -07:00
Ahmad Al-Faqih 9d1eec0ce0 fix: retain Responses assistant replies in RPC session history
Carve the history-only implementation and regression from #104754
(b4bfa76facb43111a863b1256c89875e3f01fde0) by PLASMA-FR; omit
the unrelated locale-picker changes. Complement merged #104523.

Real serve + WebSocket reconnect: SQLite retains two rows; before,
resume/activate/history return only the user; after, both survive.
Plain-content control retains both rows in both arms. Native macOS
sleep and the full desktop symptom are not established by this probe.

Refs #68321
2026-09-07 04:56:22 -07:00
Teknium 94ff4fe8f9 test: verify profile rebuild persistence through live serve 2026-09-07 04:54:49 -07:00
Teknium 9da8df8d26 fix(prompt): preserve shared project prefixes across worktrees
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.

Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
2026-09-07 04:45:48 -07:00
Teknium 08b140d14e fix: gpt-6 Astra on Codex OAuth gets the 85% compaction autoraise
Codex OAuth caps gpt-6-astra at the same 272K window as gpt-5.4/5.5/5.6,
so the global 50% trigger compacted at ~136K. Extend the existing
codex_gpt55_autoraise gate to any slug containing "astra" (minus the
opt-in -900k picker variants, which already unlock the wider window).
Other routes (OpenAI direct, OpenRouter) keep the user threshold.
2026-09-06 23:08:45 -07:00
Teknium be58c276ee feat(compression): per-image token cost learned from the provider's own usage (#70328, supersedes #70463)
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).

The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.

- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
  anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
  ~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
  read the same bound value, so trigger and walk agree; the per-message memo now caches text
  tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.

evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.

Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
2026-09-06 14:19:42 -07:00
Joey 684a2cfbd7 fix(plugins): give dual-kind memory hooks a single owner 2026-09-06 13:36:12 -07:00
Joey d932fa5929 fix(memory): spill oversized external prefetch 2026-09-06 13:25:48 -07:00
Teknium 1d06f7a95a docs+evals(compression): token-accounting replay harness and developer-guide section (#104462)
evals/token_accounting/replay_gates.py runs the REAL AIAgent turn loop against a local fake
chat-completions server with scripted usage.prompt_tokens, in three shapes (CLI same-object
history, gateway JSON-reloaded history, SessionDB close/reopen + fresh agent) x two arms
(transcript 3x threshold by bytes/4 while real usage is under; transcript tiny while real usage is
over). Acceptance: no gate fires when real usage is under threshold regardless of estimate
inflation; every gate fires once real usage is over.

origin/main 5f406d88ea: cli/gateway/restore inflated all FAIL (1 local compaction each on the
estimate). This branch: 6/6 PASS, anchor restored in the fresh process.
2026-09-06 13:21:17 -07:00
Teknium f1ccf436a2 feat(web): Perplexity Search API as a web_search + web_extract backend
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:

- search: POST https://api.perplexity.ai/search (documented Search API),
  search_context_size=low so `snippet` stays description-sized;
  results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
  route behind `pplx content snippets` (the CLI's `content fetch` is
  deprecated upstream). web_extract has no query, so the URLs' path words
  serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
  backend set + credential ladder + availability probe (web_tools),
  registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
  dump key lists, nous_subscription direct-credential detection, setup
  summary, test conftests, docs.

Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
2026-09-04 07:17:00 -07:00
Teknium 4441a2a28d docs(agents): split AGENTS.md into root + per-area files (≤8k each, the subdirectory-hint cap)
Root AGENTS.md 100,797 → 29,295 chars: what applies everywhere (invariants, rubric, footprint ladder, layout + shape rules, commit/PR, testing) plus a routing table. Area rules move to agent/, hermes_cli/, gateway/, tools/, plugins/, tui_gateway/, web/, skills/, cron/, apps/desktop/src/ AGENTS.md (3–9k each; ceiling is now 32k after d61cff60e3, target ~8k). Long-form process-identity and skin key tables go to website/docs/developer-guide/cli-internals.md. Zero rule loss; map in /tmp/rf/agents_md_zero_loss.md. Stale Bot Mode test paths corrected to apps/desktop/src/plugins/hermes-bots/*.test.ts.
2026-09-04 02:12:35 -07:00
Teknium 0a5164cebe compat(plugins): tell users which installed plugins break on 2026-09-14, and stop loading them after
hermes_cli/plugin_compat.py is now the single source of truth for the compat window:
  COMPAT_REMOVAL_DATE = 2026-09-14; scan_plugin() statically finds `from F import n`, `import F` + `F.n`,
  alias forms and string targets against compat_manifest.json; compat_report() aggregates over the user's
  ENABLED external (non-bundled) plugins; disable_reason() decides the loader's skip.

Surfaces (all read from that one report):
  * CLI: yellow block under the banner naming plugins + date + `hermes plugins compat` (red + DISABLED after)
  * `hermes plugins compat [--json] [path]`: file:line, old -> new per hit; exit 1 while anything remains;
    `path` lets a plugin author scan their own checkout
  * `hermes doctor`: "Plugin import paths (removed Sep 14, 2026)" section next to the xAI retirement check
  * `hermes update`: post-update notice alongside the FTS/curator notices
  * Desktop: compat_report() writes HERMES_HOME/.plugin-compat-report.json (deleted when clean); Electron
    shows ONE warning dialog per distinct report after the backend is up and persists the dismissal in
    userData/plugin-compat-dismissed.json. A new affected plugin, or the date passing, is a new report.

From the date, PluginManager skips a hitting external plugin before importing it, with the reason in
LoadedPlugin.error ("uses N import path(s) removed on 2026-09-14; run `hermes plugins compat` ...") — the
same path a plugin with a broken register() takes, so nothing else is affected. Escape hatch:
plugins.allow_deprecated_imports: true (config_defaults), which only helps until the compat commit is
actually reverted.

Docs: COMPAT_MANIFEST.md (removal date, what-happens table, author instructions), plugin dev guide section.
Tests: tests/test_plugin_compat_notice.py (scanner forms, report scope, date gate + escape hatch, summary
text, report file lifecycle, loader skip via a real PluginManager), electron/plugin-compat-notice.test.ts
(show once, re-show on a different set or on the date passing, malformed file ignored).

Live A/B on this box with a demo plugin on old paths: before the date it loads and the banner/doctor/report
name it; with today=2026-09-14 it is skipped with the reason and the banner turns red; with the escape
hatch it loads again.
2026-09-04 01:28:31 -07:00
Teknium 4ab96371d2 docs: sync developer-guide tooling/gateway/CLI +
user-guide with the facade/siblings layout (#102117)
2026-09-04 00:14:22 -07:00
Teknium 4fb332b427 docs: sync developer-guide agent-core docs with the facade/siblings layout (#102117) 2026-09-04 00:07:14 -07:00
Teknium 2b55ded1ac perf(state): keep delegate-child transcripts out of the trigram FTS index (schema v30)
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.

Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.

The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
2026-09-03 02:35:37 -07:00
Teknium 0cbc6e37ac test/docs: trim seam tests to invariants, document create_client and external-process fields
Cuts the 41 contributor tests down to 8 pinning the before/after contracts
(out-of-tree provider resolves end to end, copilot-acp unchanged, broken
plugin falls through, flat-install discovery + non-provider kinds untouched).
Adds the create_client hook and process_* fields to the model-provider
plugin developer guide.
2026-09-02 09:57:39 -07:00
Hudson db639e1023 fix(gateway): bind every adapter to the runner at the _create_adapter boundary
Built-in adapters (Signal, WhatsApp Cloud, Weixin, MSGraph, BlueBubbles, ...)
were returned from the if/elif factory without `gateway_runner`, so
`build_source` never consulted `profile_routes` for them — routed inbound
events landed in the default profile's agent:main namespace. Only the
plugin-registry branch and api_server/webhook set the back-reference.

Split the factory: `_instantiate_adapter` builds, `_create_adapter` binds
the runner on every non-None result. All lifecycle callers (primary
startup, reconnect, secondary-profile startup) already go through
`_create_adapter`, so this covers every path with one seam instead of
per-branch assignments.

Salvaged from #70831 (Hudson). First reported in #68332.
2026-09-02 06:47:45 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium df4b3733ba fix(cron): every last_status consumer renders delivery_failed explicitly (dashboard badge, Desktop inspector, /cron list, docs)
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):

- web dashboard CronPage: last_status was never rendered at all — a
  delivery_failed job showed a green 'scheduled' badge and only a small red
  'delivery: ...' line. New pure cronLastResult() helper maps the closed
  literal set to tones (ok=success, delivery_failed/blocked_config=warning,
  error/unknown=destructive) and the card now shows an amber
  'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
  literal; routineLastResult() spells out each one ('Ran, but delivery
  failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
  appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
  on this branch; no consumer compared == 'ok' for success apart from the
  cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
  detail field carries the reason.

Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
2026-09-02 00:52:58 -07:00
Teknium 6879a621b1 fix(gateway): live foreign token lock at startup exits 78 instead of retry-queueing forever
BasePlatformAdapter._acquire_platform_lock emits `{scope}_lock` with
retryable=True on purpose (#54167): a MID-RUN reconnect must be able to
recover once the live holder exits or a stale record is cleared. The
startup router keyed solely off that flag, so a live foreign holder of the
bot token at zero-connected startup landed in `_failed_platforms` with
gateway_state=running — alive, deaf, and retry-storming the token every
backoff — instead of the exit-78 (EX_CONFIG / startup_failed) contract
that #51228 established for single-writer conflicts.

Minimal class fix, salvaged from #83183 (@alexgunsberg) against current
main:

- gateway/restart.py: `is_global_startup_conflict(error_code)` — matches
  the `*_lock` / `lock_conflict` code families every adapter emits for
  scoped-lock and identity conflicts. Code only, never message text.
- gateway/run.py primary startup routing: a lock-conflict failure is
  routed as non-retryable (parked `fatal`, not queued). Nothing else
  connected → exit 78; alongside a transient peer → NS-609 mixed mode,
  gateway stays alive and only the peer retries.
- gateway/run.py `_schedule_secondary_profile_startup_reconnect`: the same
  contract for multiplex secondaries — park `<profile>:<platform>` fatal
  like `duplicate_credential` instead of scheduling a reconnect storm.
- Mid-run behavior is untouched: `_handle_adapter_fatal_error_impl` and
  the reconnect watcher still treat `*_lock` as retryable (#54167).

Not carried over from #83183 (superseded on main or out of scope): the
`degraded` lifecycle write only fires on the all-retryable path and the
runner immediately overwrites it with `running` (so busy/drain already
see `running`); the secondary retry bridge landed separately in
96489f3c1b (#92064); Buzz/IRC/LINE lock-tuple unpack and the reconnect
ownership registry are separate class fixes.

Live repro (real GatewayRunner.start(), isolated HERMES_HOME + lock dir,
live holder subprocess owning the lock via production
acquire_scoped_lock): before — exit_code=None, gateway_state=running,
telegram `retrying`, queued in _failed_platforms; after — exit_code=78,
gateway_state=startup_failed, telegram `fatal`, _failed_platforms={}.

Co-authored-by: alexgunsberg <alex@gunsberg.fi>
2026-09-02 00:17:54 -07:00
Teknium 5dfd1e77a8 fix(recovery): register lazy state.db tables in one schema map; cover .recover lane and count-mismatch loss
Follow-up to the salvaged #100350 commits: replace the per-table
'if table == "delivery_obligations"' branches in session_recovery.py and
session_lost_and_found.py with a single _AUXILIARY_TABLE_SCHEMAS registry
(table -> destination DDL initializer) that both the SQL-level and the
lost_and_found lanes consume, so the next lazily-created state.db table is
one entry, not three code paths. The .recover lane now iterates
_CANONICAL_TABLES + _AUXILIARY_TABLES instead of a duplicated literal list.

Tests: the .recover direct-copy lane creates the missing ledger on the
destination; a source-vs-destination obligation count mismatch fails
verification (complete=False) instead of reporting a clean salvage.
Docs: state.db table inventory lists delivery_obligations.

Addresses #100313
2026-09-02 00:00:55 -07:00
Matthias Reso d7e92ab7e3 feat(image_gen): add Meta Model API (muse-image) provider plugin
Adds a bundled image-generation backend for the Meta Model API
(https://api.meta.ai/v1), which is OpenAI-compatible. Exposes the
muse-image-1.0 model via the standard image_generate tool. This is the
image-gen companion to the already-bundled meta-ai chat provider
(plugins/model-providers/meta-ai, PR #88565).

- plugins/image_gen/meta-ai/ — provider registered as `meta-ai`, matching
  the chat provider's id. Reuses the openai SDK pointed at Meta's base URL.
- Auth mirrors the chat provider: MODEL_API_KEY (Meta's documented var),
  with META_API_KEY / META_MODEL_API_KEY aliases and a META_BASE_URL
  override.
- Text-to-image only for now (capabilities gated); base64 (WebP) and URL
  responses both handled and saved under $HERMES_HOME/cache/images/.
- Auto-loads as `kind: backend` and appears in `hermes tools` with no
  central list edits, matching the other bundled providers.
- tests/plugins/image_gen/test_meta_ai_provider.py — 27 tests (metadata,
  auth-alias resolution, base-url override, model resolution, generate
  paths incl. b64 save, aspect mapping, URL caching, error handling).
- docs: image-generation feature page + provider-plugin built-in list.
2026-09-01 22:17:08 -07:00
Lakshya Agarwal 428e084dcd feat(web): add Tavily web search and extract provider
This commit re-introduces the Tavily provider, which supports both search and content extraction capabilities, which was removed in #99199.
2026-09-01 10:56:49 -07:00
Teknium 8b73720fa7 fix(memory): tolerate bare-signature v2 providers when forwarding checkpoint requirement
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
2026-08-31 09:57:22 -07:00
Teknium d6773cf26f refactor: remove the Tavily web backend; keyless ring is exa/parallel/firecrawl/keenable
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
  ring entry removed from keyless_mcp, legacy backend set / credential
  ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
  nous_subscription surfaces. The tvly- redaction pattern stays --
  legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
  the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
  environment-variables, tools-reference, web-dashboard, provider
  plugin dev guide.

Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
2026-08-31 00:56:41 -07:00
Teknium 4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
Marco Fernstaedt a2af8405d1 fix(compression): derive native threshold from local trigger 2026-08-30 05:15:37 -07:00
nftpoetrist 0610291b5a fix(prompt): sync DEFAULT_SOUL_MD with the #95681 identity rewrite
DEFAULT_AGENT_IDENTITY was rewritten in agent/prompt_builder.py (behavior
spec, exploration-thrift line deliberately removed) but the actual seed
written to disk on first run, hermes_cli/default_soul.py's
DEFAULT_SOUL_MD, was never updated. ensure_hermes_home() writes
DEFAULT_SOUL_MD into SOUL.md on every fresh install before the agent's
first turn, so virtually all real users end up as "SOUL.md users" seeded
with the pre-rewrite text -- including the exact "targeted and efficient
exploration" line the rewrite explicitly banned -- while the new
DEFAULT_AGENT_IDENTITY fallback essentially never serves the "fresh
install" audience its own PR body named as the target.

- DEFAULT_SOUL_MD now matches DEFAULT_AGENT_IDENTITY exactly.
- The pre-rewrite text is added to _LEGACY_TEMPLATE_SOULS so installs
  already seeded with it self-heal via the existing upgrade-in-place
  mechanism (same guarantee as the comment-only scaffold entries: the
  string carries zero user intent, so it's safe to replace).
- Synced the other places install.sh's own comment says "MUST match
  DEFAULT_SOUL_MD": scripts/install.sh, scripts/install.ps1,
  docker/SOUL.md, and the docs/i18n pages that quote the fallback text
  verbatim.
2026-08-29 18:10:47 -07:00
Neel Patel ca3060b70f docs: chat-completions is now a compat shim on Router, not a 404
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
2026-08-29 20:29:04 +05:30
Neel Patel 804f8b4732 feat(providers): add Ramp Router (router.com) provider plugin
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:

- plugins/model-providers/router/: RouterProfile plugin —
  api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
  RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
  GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
  Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
  runtime_provider._detect_api_mode_for_url: api.router.com ->
  codex_responses. The host is Responses-only — POST /v1/chat/completions
  does not exist and 404s — so this is a genuine host mandate (exact
  hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
  hook (tri-state: None=defer, ()=model takes no reasoning params,
  tuple=clamp set). Router validates reasoning.effort per model and
  returns HTTP 400 invalid-argument on levels outside the model's
  published vocabulary, and 400 unsupported_parameter when a
  non-reasoning model receives any reasoning field (both verified live).
  The profile answers from a cached copy of the catalog's
  router.capabilities.reasoning block: cache-only on the hot path,
  seeded for free by fetch_models(), disk-mirrored across processes
  (/cache/router_catalog.json), background-warmed when cold
  — same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
  the generic effort-clamp branch (xai/actual/github branches untouched;
  profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
  document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
  rejection, profile registration + auth auto-registry wiring, catalog
  parsing, and transport clamp/suppression/fallback paths.

Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
2026-08-29 20:29:04 +05:30
teknium1 272f4e4abe feat(plugins): generalize native platform handler registration to every gateway platform
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.

- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
  invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
  api_server/msgraph_webhook wire before their dispatch tables freeze;
  the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
  as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
  calling the hook.
2026-08-27 07:51:37 -07:00
teknium1 c96f830252 feat(plugins): let plugins register Telegram PTB handlers via ctx.register_telegram_handler
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.

Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
2026-08-27 07:51:37 -07:00
Teknium 6e5413844e feat(compression): lean tail retention is the default — compaction keeps 10-25K verbatim, not 100-240K
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.

Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.

Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).

Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).

E2E counterfactual (real imports, 1M window @ 0.85):
  main default:  legacy, tail 170,000 (ceiling 255,000)
  head default:  lean,   tail  25,000 (ceiling  37,500)
  head legacy:   170,000 (opt-out intact)
  update_model to 400K: 10,000 (lean preserved across switch)
2026-08-26 07:16:04 -07:00
Teknium 1ee524f77d fix(memory): bind the checkpoint gate to post-turn micro-compaction too
Independent review caught a compaction authority the gate missed:
post-turn micro-compaction (turn_finalizer -> _micro_compact) absorbs the
oldest exchanges into a rolling summary with no pre-compress checkpoint
hook in its path, and both compression.checkpoint_required and
compression.micro_compact could be enabled together — assistant evidence
could vanish into a summary the checkpoint filter later excludes, without
ever reaching the durable provider.

- agent_init: checkpoint_required forces micro-compaction off (warned),
  mirroring the native-compaction suppression
- turn_finalizer: defense-in-depth guard at the call site (attribute is
  plain mutable state a future path could flip on a live agent)
- behavioral regression test with a sabotage control (gate off proves the
  harness reaches the call site; gate armed proves zero calls)
- docs + config example mention the suppression; stale v1 test header fixed
2026-08-25 03:55:55 -07:00
Teknium 9e551d2931 refactor(memory): renumber checkpoint API — v1 is the implicit historical contract, v2 opts into fail-closed checkpoints
Per review: existing providers should not be retroactively re-versioned or
handed a changed payload. Version 1 is now the implicit historical
on_pre_compress() contract (best-effort, raw message list) that every
pre-existing provider is already on; the fail-closed checkpoint contract
becomes version 2. MemoryManager routes the raw transcript to v1 providers
unchanged and hands the host-normalized evidence list only to v2+ checkpoint
providers, so the plugin surface contract for shipped providers is
byte-identical with the gate off.
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 8cc379b528 fix(memory): bind the checkpoint gate to every compaction authority
The fail-closed gate lived only in compress_context(), but two native
lossy owners compact without ever crossing it (review on #93996):

- codex app-server: in "native"/"off" auto-compaction mode (native is the
  default) Hermes preflight is skipped and the codex agent compacts its
  own thread inside run_turn() — the compress_context() rejection was
  unreachable. init_agent now refuses checkpoint_required together with
  api_mode=codex_app_server (BLOCKED_MISSING_PREREQUISITE, extracted as a
  testable guard), and run_codex_app_server_turn() fails closed as
  defense in depth before a turn can reach the codex-owned boundary.
- Responses server-side native compaction:
  native_compaction_context_management() now returns None while the gate
  is armed, so context_management never goes on the wire and the
  checkpoint-aware Hermes compressor stays authoritative. The suppression
  is logged once per process, not silently applied.

Regressions: checkpoint_required + app-server raises before run_turn()
(the session is never created); checkpoint_required keeps
context_management off the wire while the plain configuration still
produces it; the init guard refuses exactly the incompatible pair. Docs
and cli-config.yaml.example describe both bindings.

Refs #93986
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 70d0b1fffb fix(memory): review follow-ups for the pre-compress checkpoint contract
Addresses the review on #93996:

- gateway: hygiene and manual /compress load the memory provider only when
  compression.checkpoint_required is enabled (skip_memory=not required).
  The historical fast path — no provider init, no best-effort hook — is
  back for everyone who did not opt in, so default behavior is truly
  unchanged.
- conversation_compression: assistant messages carrying both prose and
  tool_calls keep their prose in the checkpoint evidence (the tool_calls
  payload is stripped, the original message is not mutated); pure
  tool-call wrappers without prose are still dropped.
- tests: legacy-database regression proving the _compressed_summary column
  is added by the declarative _reconcile_columns() path on a plain reopen
  (no version-gated migration needed — append_message works right after),
  plus coverage for the prose-preserving filter.
- docs: providers must implement idempotent, content-keyed checkpoint
  writes — a fail-closed block means the next attempt re-runs
  on_pre_compress over largely the same transcript.

Refs #93986
2026-08-25 03:55:55 -07:00
Jan-Stefan Janetzky 3c31c44880 docs(memory): document the pre-compress checkpoint contract
Adds a 'Pre-Compress Checkpoints (fail-closed)' section to the memory
provider plugin guide: the versioned opt-in attribute, the operator-side
compression.checkpoint_required gate, fail-closed semantics, and the
normalized evidence contract including the persistent summary marker.

Refs #93986
2026-08-25 03:55:55 -07:00
Teknium 0484910787 feat(terminal): pluggable terminal environment backends via plugin registry
Third-party sandbox vendors can now ship a terminal backend as a standalone
plugin instead of landing in core. Adds the five-piece pluggable-subsystem
pattern for terminal environments:

- agent/terminal_env_provider.py — TerminalEnvironmentProvider ABC with
  declarative classification flags (is_remote, is_container,
  skip_container_guards, cache_path_base, strip_env_keys,
  session_isolated_when_nonpersistent) so every historical
  frozenset-of-names classification site consults the registry instead
- agent/terminal_env_registry.py — thread-safe scoped registry; built-in
  backend names are reserved and unregistrable
- PluginContext.register_terminal_environment_provider() mirroring
  register_browser_provider
- _create_environment falls through to registered providers; unknown-backend
  errors list plugin names
- Classification sites wired: approval guard skip, container path/cwd
  handling (terminal/file/code-exec), prompt-builder env hints + probe,
  host env probe suppression, skills remote-env note, cache path
  translation, subprocess secret stripping (both spawn paths),
  per-session isolation for name-resumed sandboxes
- Surfaces: hermes setup picker + doctor + status rows, dashboard
  terminal-backend picker rows/probe/validation, terminal.backend schema
  options recomputed per request
- Docs: developer-guide/terminal-environment-plugin.md + sidebar + plugins
  capability table
2026-08-24 20:10:44 -07:00
Teknium 933c209e96 fix(compression): -900k Codex variants keep the global 50% threshold; 85% autoraise stays on 272K base slugs
The 85% compaction autoraise exists to stop wasting the small advertised
272K Codex window. -900k large-context picker variants (#92797) run at
~900K, where the global compression.threshold (default 50%, ~450K) is the
right behavior — autoraising them to 85% (~765K) would delay compaction
far past what the user configured.

- _is_codex_gpt54_or_gpt55() excludes valid -900k variants, so both the
  85% override and the one-time autoraise notice skip those sessions.
- Base slugs are unchanged: 272K window + 85% autoraise.
- Tests: variant/base threshold pairs incl. namespaced ids; docs note in
  the -900k section.
2026-08-23 02:54:30 -07:00
Teknium 63a9c26fbe feat(model): Codex GPT slugs default back to 272K; explicit -900k picker variants opt into the verified large window
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).

- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
  advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
  gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
  the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
  gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
  the wire (main transport + auxiliary Responses adapter), and pricing
  aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
2026-08-23 02:14:35 -07:00
Teknium a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00