Commit Graph

24179 Commits

Author SHA1 Message Date
liuhao1024 118dbe871f fix(cli): map kitty CSI-u lock-bit variants so key combos survive NumLock
kitty and ghostty OR the CapsLock (64) / NumLock (128) state into the
CSI-u modifier parameter. With NumLock on, Ctrl+C arrives as
ESC[99;133u (5 + 128) instead of ESC[99;5u; the alias table had no
entry for it, so every key combo leaked as literal text like
[127;133u (#89651). Install every CSI-u alias with the lock-bit
variants (+64/+128/+192); the xterm modifyOtherKeys encoding never
carries lock bits, so the ESC[27;N;CP~ form is left untouched. The
Esc-key registration now covers modifier 1 as well (1+128=129 is a
lone Esc with NumLock on).
2026-08-20 12:16:20 +05:30
Teknium 79e6d3e6d7 Revert "ci: parse-cache buster (zero-job dispatch, new blob forces re-parse)"
This reverts commit 87944ad80b.
2026-08-19 23:42:05 -07:00
Teknium 87944ad80b ci: parse-cache buster (zero-job dispatch, new blob forces re-parse) 2026-08-19 23:40:20 -07:00
Teknium e8c6370740 ci: retrigger wave 3 — zero-job load-shedding 2026-08-19 23:38:51 -07:00
Teknium b8b9156827 ci: retrigger wave 2 — zero-job load-shedding 2026-08-19 23:37:30 -07:00
Teknium ce47a92c2d ci: retrigger wave 1 — zero-job load-shedding 2026-08-19 23:36:08 -07:00
Teknium 1faf409407 chore: retrigger CI (zero-job dispatch failure, auto-heal) 2026-08-19 23:31:56 -07:00
Teknium 5837714be5 chore: retrigger CI (zero-job dispatch failure, auto-heal) 2026-08-19 23:30:49 -07:00
Teknium f098602a34 chore: retrigger CI (zero-job dispatch failure, auto-heal) 2026-08-19 23:29:42 -07:00
Teknium db4b840b72 ci: retrigger — zero-job dispatch on 5a3b1c2de 2026-08-19 23:28:42 -07:00
Teknium dbbd8937ae docs(computer-use): document requesting the actual screenshot on chat surfaces
Follow-up to PR #90183 — computer_use now saves a bounded shareable copy of
image captures, so attachment-capable surfaces (Telegram, Discord, Desktop)
can deliver the real screenshot when the user asks. Documents the behavior,
the 20-file cache bound, and the no-automatic-send rule.
2026-08-19 23:23:56 -07:00
Teknium 5a3b1c2de8 chore: retrigger CI (zero-job dispatch failure, auto-heal) 2026-08-19 23:22:04 -07:00
Teknium 01c3bd4c81 test: explicit keyless Firecrawl selection asserts the keyless cloud route
The salvaged #50659 behavior makes 'firecrawl selected, no creds' a
WORKING keyless state, so the old expectation (hard error naming
FIRECRAWL_API_KEY) is stale. The test now mocks httpx and asserts the
request routes to api.firecrawl.dev with results returned — still
proving keyless Tavily can't silently take over, which was the test's
point. Also stops the test making a real network call in CI.
2026-08-19 23:20:52 -07:00
Teknium 8436e0d142 chore: map contributor email for #87427 salvage 2026-08-19 23:15:01 -07:00
Teknium a92412ede1 test: mock load_config_readonly in memory status gate tests
check_memory_requirements() reads the readonly config loader; the salvaged
tests only patched load_config, so the gate saw the real config.
2026-08-19 23:15:01 -07:00
Cassie Gray b38c40319d fix(cli): align hermes memory status and docs with memory tool gate 2026-08-19 23:15:01 -07:00
Brooklyn Nicholson b8850e17b5 fix(desktop): floating sidebar overlays keep an opaque background under glass 2026-08-20 01:13:51 -05:00
kshitijk4poor 45f11263bd fix(tui): skip the kitty protocol push for Ghostty in the Ink TUI too
Widen the cli.py Ghostty exception to the sibling sites the review found: the Ink TUI pushes CSI >1u at raw-mode entry (App.tsx), on alt-screen exit, and on the extended-keys re-assert path (ink.tsx) for every EXTENDED_KEYS_TERMINALS entry including ghostty - same Alt-stripping bug. New skipKittyKeyboardProtocol() helper in terminal.ts gates the ENABLE push at all 3 sites; the DISABLE (pop) stays unconditional since popping an empty stack is a spec no-op. Also fix the cli.py comment citing the modifyOtherKeys encoding where the kitty CSI-u form (ESC[127;3u) is what the broken path expected, dedupe the quadruplicated Ghostty comment, and update the stale 'mirroring the Ink TUI' docstring. 7 new vitest cases.
2026-08-20 11:39:03 +05:30
kshitij 1a8fea3ce2 fix(cli): skip Kitty keyboard protocol push for Ghostty, use modifyOtherKeys only
Ghostty's Kitty disambiguate-mode implementation strips the Alt modifier
from the Backspace key — Option+Backspace arrives as bare \x7f instead of
the expected \x1b[27;3;127~, breaking backward-kill-word.  This was a
regression introduced when PR #87630 re-added the CSI >1u Kitty protocol
push for all allowlisted terminals including Ghostty.

Under modifyOtherKeys mode (CSI >4;2m), Ghostty correctly sends
\x1b[27;3;127~ for Option+Backspace, which the alias table in
pt_input_extras already maps to (Escape, ControlH) = backward-kill-word.

Fix: for Ghostty only, push just modifyOtherKeys and skip the Kitty
protocol push.  All other terminals (iTerm2, WezTerm, kitty, tmux, VS Code)
still get the full dual-protocol push.

Ghostty upstream tracking: discussion #9560, issue #9895 (cmd+backspace
variant of the same root cause).
2026-08-20 11:39:03 +05:30
kshitijk4poor 67029492d4 chore: map salvaged contributor emails (Dhruv7201, ekinnee) 2026-08-20 11:37:01 +05:30
kshitijk4poor b7e12decc6 fix(agent): route relay-wrapped output-cap 429s into the output-cap handler
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
  provider fallback (a deterministic request-shape failure that failover
  cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
  clamp+compress recovery as plain output-cap 400s.

Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
2026-08-20 11:37:01 +05:30
ekinnee 99c980f466 fix(model-metadata): parse 'exceeds model maximum output tokens' cap errors
Recognizes the DeepSeek/OpenAI-compatible relay wording
  max_tokens (98304) exceeds model's maximum output tokens (65536)
in both parse_available_output_tokens_from_error (returns the cap) and
is_output_cap_error (keeps the 400 out of the compression death-loop).

Salvaged from PR #72283; the conversation_loop early-clamp block was
dropped in favor of routing through the existing output-cap handler
(follow-up commit).
2026-08-20 11:37:01 +05:30
Dhruv Modi 7f2733b71c fix(model-metadata): converge output-cap retry on vLLM
Fixes the retry loop that spins forever when a vLLM server rejects a
request for having a max_tokens too big for what is left of the context
window.

The catch is that vLLM does not tell you how big your prompt actually is
in that situation. It works the number backwards from the constraint it
just failed, so you get:

    "requested 65536 output tokens and your prompt contains at least
     36865 input tokens, for a total of at least 102401 tokens"

That 36865 is just window + 1 - requested, and the total is always
exactly window + 1. Subtracting it from the window hands back
requested - 1 every single time, whatever the real prompt size is.

parse_available_output_tokens_from_error believed it and returned
requested - 1. conversation_loop then takes off its 64 token safety
margin and retries, which walks the cap down 65 tokens at a time while
the reported input walks up by the same 65:

    65536 -> 65471 -> 65406 -> 65341

Three attempts is the default budget, so the session gives up with
"Context length exceeded" having closed 195 tokens of a roughly 28000
token gap. Compression cannot save it either, because the input was
never the problem, which is why the compressor keeps refusing with
"summary would have GROWN".

This is also what is behind the unexplained "input-token drift" in
issue #61761. The input is not drifting. It is a derived number, and it
moves because we moved max_tokens.

So when that shape shows up (the "at least" wording, plus a budget that
works out to exactly requested - 1), halve the requested cap instead. It
is still guaranteed to sit under whatever was just rejected, and it
converges on the first retry: 65536 -> 32768, which next to a real 36865
token prompt comes to 69633 against a 102400 window.

Nothing else moves. A measured input is still trusted, and a genuine
input overflow still returns None so the caller falls through to
compression the way it always did.

The existing test asserted the bogus 65535, so it is updated. Added
tests for the measured input path, and for the retry actually
converging.
2026-08-20 11:37:01 +05:30
kshitijk4poor 0596ccdeb3 fix(compression): salvage follow-up — todo snapshot last-resort, reuse prune helpers
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
  reduced only as a LAST resort after reasoning/tool/summary shrink ops,
  and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
  _PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
  and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
  truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
2026-08-20 11:36:54 +05:30
MindDragonLabs 5c03fbedc6 fix(config): recognize memory nudge interval 2026-08-20 11:36:54 +05:30
MindDragonLabs fb96247eaf fix(compression): salvage grown candidates before refusal 2026-08-20 11:36:54 +05:30
liuhao1024 62016a1b0a fix(compression): count a would-grow refusal as an ineffective strike
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).

Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.

Fixes #88568
2026-08-20 11:36:54 +05:30
Ben Barclay 5a17b1f41d Merge pull request #89584 from victor-kyriazakos/relay-ws-hardening
fix(relay): rc.4 relay transport + inbound fixes — dedupe replays, fail pending on drop, fail fast mid-redial, WAN keepalive
2026-08-20 16:06:23 +10:00
Gille 9ae7247616 fix(desktop): skip source restore in auxiliary windows 2026-08-20 01:04:42 -05:00
Gille c25aa35e1a fix(desktop): keep peer windows on shared gateway 2026-08-20 01:04:42 -05:00
Teknium 797bc4bf9b Merge remote-tracking branch 'origin/main' into feat/keyless-tavily-firecrawl-failover 2026-08-19 23:04:36 -07:00
Teknium c492379497 chore: map contributor email for Tavily salvage 2026-08-19 23:04:20 -07:00
Slobaka 6ff341c4d6 fix(doctor): web readiness reflects the selected provider's real state (#78412)
Salvaged from #78434 by @Slobaka (also the issue reporter; earlier than
the competing #78436). hermes doctor no longer paints a green web check
when the explicitly selected provider cannot initialize — web splits
into per-capability rows (web search / web extract) resolved through
the same registry resolvers the dispatchers use, with readiness from a
true availability probe (_provider_is_ready).

Keyless-tier integration on top of the salvage:
- _provider_is_ready counts is_keyless_available() as ready — keyless
  mode is a working state, not a misconfiguration (zero-config installs
  and selected-keyless Tavily/Firecrawl show ok, not warn)
- Tavily/Firecrawl gain is_keyless_available() (True only when
  explicitly selected — they stay out of the zero-config fallback)
- doctor triggers plugin discovery before reading the registry (fresh
  doctor processes saw an empty registry and warned on everything)

E2E: searxng-selected-without-URL warns (the #78412 repro);
zero-config, tavily-keyless, firecrawl-keyless all read ok;
parallel pinned paid without a key warns.
2026-08-19 23:03:58 -07:00
Brooklyn Nicholson f983828077 fix(desktop): stamp the focused profile on projects RPCs and refresh a remote scan
Remote mode used to return before asking the host for repos, and never sent
profile, so the sidebar stayed on the launch list. Ask discover_repos to scan,
forward the focused profile on every projects call, and drop late responses
from a profile the user already left.

Co-authored-by: Chen Jin <Enough1122@users.noreply.github.com>
Co-authored-by: Ryan Weddle <weddle@gmail.com>
Co-authored-by: izumi0uu <izumi0uu@gmail.com>
Co-authored-by: webtoolbox <1911826+webtoolbox@users.noreply.github.com>
2026-08-20 01:02:17 -05:00
Brooklyn Nicholson 4dcefed089 fix(tui_gateway): scan remote git roots and scope projects.* to the focused profile
A remote desktop cannot crawl the host disk, and projects.* always read the
launch profile's stores, so switching profiles left the wrong tree on screen.
Bind the requested profile's HERMES_HOME and session db for the whole family,
and add scan:true so repos with no Hermes sessions still appear.

Co-authored-by: Chen Jin <Enough1122@users.noreply.github.com>
Co-authored-by: Ryan Weddle <weddle@gmail.com>
Co-authored-by: izumi0uu <izumi0uu@gmail.com>
Co-authored-by: webtoolbox <1911826+webtoolbox@users.noreply.github.com>
2026-08-20 01:02:17 -05:00
Teknium 481bc9391e fix(memory): profile-only config gets narrow USER_PROFILE_GUIDANCE instead of the full memory block
With memory_enabled: false but user_profile_enabled: true, the memory tool
stays (it backs USER.md) but the full MEMORY_GUIDANCE told the model to save
notes to a MEMORY.md store that does not exist. Split the guidance: a
profile-only block is injected for that configuration, directing writes to
target='user' only.
2026-08-19 22:59:17 -07:00
HexLab98 a969c5a93d test(memory): cover the disabled built-in memory surface
Walks the real resolution chain -- config.yaml on a temp HERMES_HOME ->
check_memory_requirements -> get_tool_definitions -- rather than mocking
the availability check, since the bug was in how the flags reach the
schema. Covers both flags off, either one alone, no config file at all,
and a config read that raises (must fail open).

Also asserts the external provider's tools survive with the built-in tool
gone, so the fix cannot regress into taking Hindsight/Mem0 down with it,
while disabled_toolsets keeps its documented "hide everything" meaning.

The existing MEMORY_GUIDANCE test built a skip_memory agent whose flags
were both false, so it was asserting the old tool-presence-only behavior;
it now states its precondition and gains the false-case mirror.
2026-08-19 22:59:17 -07:00
HexLab98 d5cddae187 fix(memory): drop dead memory tool and guidance when built-in stores are off
With memory.memory_enabled and memory.user_profile_enabled both false,
agent_init never builds a MemoryStore -- but check_memory_requirements()
returned True unconditionally and MEMORY_GUIDANCE was gated only on the
tool being present in valid_tool_names. So the tool shipped in every
request's schema while answering "Memory is not available" on every call,
and the system prompt still told the model to save durable facts there.

Gate both on the config flags, using the store predicate for the tool and
the already-resolved agent state for the guidance (config is not re-read
mid-conversation, so the prompt stays byte-stable). Either flag alone
still backs the tool, so only turning both off removes it.

This lets a user running a third-party provider (Hindsight, Mem0, ...)
turn the built-in files off without paying for the dead surface on every
API call. The provider's own tools are unaffected: hiding the built-in
tool moves the decision onto the toolset gate, and listing memory under
agent.disabled_toolsets remains the only switch that takes those down.
2026-08-19 22:59:17 -07:00
kshitij 37fa4a7c63 chore: remove unused import pytest from test file
Follow-up cleanup from simplify-code review on PR #90521 salvage.
2026-08-20 11:28:26 +05:30
liuhao1024 fbca706789 fix(telegram): log the first confirmed getUpdates progress per generation
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).

Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.

Fixes #90504
2026-08-20 11:28:26 +05:30
Teknium 2eb5217af9 feat: keyless free-tier failover + Tavily/Firecrawl salvage integration
- Cross-vendor failover: when Exa's or Parallel's keyless free tier
  returns a rate-limit-shaped error, the request retries once on the
  other vendor's free endpoint (search + whole-batch extract). Result
  notes served_by; a peer pinned to its paid tier is never used;
  non-throttle errors never fail over.
- Docs: failover note + Tavily/Firecrawl keyless-when-selected rows.
- Firecrawl keyless test expectations aligned with the keyless tier.
2026-08-19 22:58:05 -07:00
Gille 188d47919a feat(computer-use): expose screenshots for chat delivery 2026-08-19 22:58:02 -07:00
Teknium 2d59cb4386 fix(api-server): 'max' and 'ultra' reasoning efforts are no longer silently ignored on API/browser requests
_request_reasoning_config() whitelisted none..xhigh, so a client sending
max or ultra (valid /reasoning + config.yaml levels) fell through to the
default effort with no error. The server now accepts the full internal
ladder (hermes_constants.VALID_REASONING_EFFORTS); per-provider wire
clamping happens downstream via agent.reasoning_effort, same as every
other entry surface. Salvages the api_server hunk of #78216 (credit
@snowzlmbot); the un-clamping half of that PR was rejected separately.
2026-08-19 22:56:14 -07:00
Teknium 4999af5cc1 fix: K3 plan-variant slugs (k3-256k) now get K3's effort vocabulary
kimi_supported_efforts() used exact/prefix matching and missed Kimi
Coding plan variants like k3-256k, which fell back to the K2-era
low/medium/high set and mistranslated efforts on a K3 wire. Replaced
with the boundary-token regex from #76427 (credit @ruizanthony), which
matches k3/k3-256k/kimi-k3* without matching kimi-k2.6 or mk3000.
2026-08-19 22:56:04 -07:00
Ben Barclay 141af4febf Merge branch 'main' into relay-ws-hardening
#85796 (live-cards gateway half) landed on main and touches the same
two files. Resolutions:

- adapter.py __init__: union — the dedupe seen-set and the live-cards
  draft/seal caches are independent sibling attributes.
- ws_transport.py _request_response: keep main's ambiguous-timeout
  contract AND this branch's raising-write catch, composed: a raise
  from the WRITE itself means the frame never reached the wire
  (definite non-delivery, no flag), while a failure surfaced after the
  frame was sent carries ambiguous=True like the timeout — tracked via
  a frame_sent marker.
2026-08-20 15:54:43 +10:00
LeonSGP43 f51e61136a fix(web): explicit Firecrawl selection works keyless against the public cloud API
Salvaged from #50659 by @LeonSGP43 onto current main (the client
resolver was rewritten for strict-selection semantics since the PR;
reapplied the keyless mode as a third client_mode inside the new
resolver). An explicit firecrawl selection with no FIRECRAWL_API_KEY /
FIRECRAWL_API_URL now routes through a minimal REST client (v2 search +
scrape, no Authorization header) instead of erroring. Unconfigured
installs never route here — the keyless path requires the explicit
selection. Fixes #49912.
2026-08-19 22:54:28 -07:00
hermes-seaeye[bot] bbb7b607b4 fmt(js): npm run fix on merge (#90552)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-08-20 05:49:39 +00:00
Lakshya Agarwal 6bf4375755 feat(onboarding): enhance Tavily backend support for keyless access
- Added support for keyless Tavily integration in the onboarding flow, allowing it to be recognized as available without an API key.
2026-08-19 22:44:11 -07:00
Lakshya Agarwal ee37f3d897 feat(tavily): update Tavily integration to support keyless access
- Updated the Tavily API key description to clarify that it is optional and keyless access is supported.
- Modified the Tavily plugin and provider to handle requests with or without an API key, using Bearer authentication when the key is provided.
- Enhanced documentation to reflect the new keyless functionality and updated environment variable descriptions.
- Added tests to ensure correct behavior for both keyed and keyless requests.
2026-08-19 22:42:58 -07:00
Brooklyn Nicholson bf4f5e17f8 feat(desktop): show unread count on the sessions sidebar toggle
A small overlay on the left sidebar icon (right when panes are flipped) so a
closed sessions list still reports unfinished-unread chats.
2026-08-20 00:42:21 -05:00