5cf6122e8709fc07d2e007496ae1d6902002fa8f
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
847e401b74 |
feat(computer_use): align cua-driver 0.9 contracts
Salvaged from PR #67807 by @f-trycua onto current main. - Foreground gate: discover delivery_mode support from the live tools/list inputSchema.properties (fail closed), not the never-shipped input.delivery_mode capability token - bring_to_front: standalone strict-schema MCP tool (inject_session=False), separate approval scope, requires foreground - Verdict precedence: confirmed > unverifiable (verify before retry) > suspected_noop/refusal (escalate); surfaced as explicit verdict field - Typed cua_browser_* route inside computer_use (browser_route.py) with exact-binding, adapter-injected session, snapshot-scoped refs - Per-Hermes-session backend isolation + release_computer_use_session seam wired into AIAgent.close() - Recorded 0.9 tools/list fixture replaces fabricated capability tokens |
||
|
|
9d6d772837 |
feat(computer_use): follow cua-driver's verify → escalate ladder (#67123)
Hermes' computer_use wrapper dropped cua-driver's structured action verdicts, exposed no delivery_mode, and injected background-only guidance — so the agent reported unverified no-ops as success and concluded cua-driver 'cannot drive' Electron/Chromium surfaces (observed live on tldraw offline). Fixes #67052. Phase A — preserve the result contract: - ActionResult carries verified/effect/escalation/path/degraded/code/delivery_mode - CuaDriverBackend._action() reads structuredContent (was data-only); a helper normalizes it, additive and None-safe on old drivers - _text_response surfaces the fields additively (ok stays transport-only) Phase B — bounded, model-reachable foreground: - delivery_mode (background|foreground) + bring_to_front on the schema, dispatcher, ABC, and all input methods - foreground is capability-gated (input.delivery_mode); old drivers get a structured foreground_unsupported refusal, never a silent background downgrade - no automatic/hidden foreground retry — the model selects it from the signal Phase C — guidance + isolation: - system prompt (prompt_builder) and bundled skills/computer-use/SKILL.md go from background-ONLY to background-FIRST, teaching the AX→PX→foreground ladder driven by returned effect/escalation, not predicted from the app being Electron - foreground approval scoped by (action, delivery_mode): a background approval never silently authorizes foreground - approval state keyed per session_id so concurrent gateway runs don't leak unlocks Tests: tests/tools/test_computer_use_delivery_ladder.py (15) cover confirmed/ unverifiable/suspected_noop/degraded/old-driver verdicts, delivery_mode gating + foreground_unsupported, and session-scoped foreground approval. Existing 265 computer_use tests still green. Live E2E (real cua-driver 0.8.3 + tldraw offline on Linux/X11): a background click returned effect='unverifiable'/path='ax' (no fabricated success), and a foreground request returned code='foreground_unsupported' — correct on a driver that predates the input.delivery_mode capability. |
||
|
|
2b6897f982 |
fix(computer-use): target Linux app windows reliably (#63725)
* fix(computer-use): target Linux app windows reliably Resolve app filters through the canonical cua-driver MCP app metadata and join running app PIDs back to windows. Preserve an exact selected window across capture_after, support direct capture by pid/window_id, and send the active window ID for coordinate pointer actions on Linux. Co-authored-by: annguyenNous <annguyenNous@users.noreply.github.com> Co-authored-by: grimmjoww578 <willies578@gmail.com> Co-authored-by: ai-ag2026 <261867348+ai-ag2026@users.noreply.github.com> * fix(computer-use): address review on PR #63725 Address three review comments from @f-trycua: 1. type_text, press_key, and hotkey now carry _active_window_id and fail closed when it is missing. Previously they sent only the PID, so CUA Driver fell back to the first window for that PID — input could reach the wrong window in multi-window apps. 2. Coordinate scroll x/y are now capability-gated behind input.scroll.coordinates. CUA Driver 0.7.1 Linux schema rejects x/y on scroll; omitting them when the driver doesn't advertise support avoids the schema rejection while still routing via window_id. 3. Windows are sorted by z_index descending (higher = front, per CUA Driver semantics) instead of ascending. Null z_index (Wayland) is coerced to 0 in _ingest_windows so it doesn't crash the sort and sorts to the back instead of being selected as the capture target. --------- Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com> Co-authored-by: annguyenNous <annguyenNous@users.noreply.github.com> Co-authored-by: grimmjoww578 <willies578@gmail.com> Co-authored-by: ai-ag2026 <261867348+ai-ag2026@users.noreply.github.com> |
||
|
|
f2e37549c6 |
feat(computer_use): cross-platform cua-driver (macOS/Windows/Linux)
Make the computer_use toolset platform-agnostic by driving cua-driver on macOS, Windows, and Linux. Consumes the 8 cua-driver decoupling surfaces (capability discovery, structuredContent AX tree, opaque element_token, click button enum, explicit mimeType, machine-readable manifest, structured list_windows, structured health_report), each degrading gracefully on older drivers. Adds `hermes computer-use doctor` (drives cua-driver health_report with a per-OS check matrix and an exit 0/1/2 ok/degraded/blocked contract), full typed wrappers for the previously-uncovered cua-driver tools plus a generic call_tool escape hatch, per-session agent-cursor lifecycle, platform-aware system-prompt guidance (host-deterministic, cache-safe), and honors HERMES_CUA_DRIVER_CMD end-to-end. Replaces the macOS-only skills/apple/macos-computer-use skill with a cross-platform skills/computer-use skill, and refreshes the EN + zh-Hans docs. Supersedes #44221 (Windows-enablement salvage of #30660). Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com> |
||
|
|
c52cd48e25 |
fix(computer-use): add set_value to ComputerUseBackend ABC and _NoopBackend stub
_dispatch() routes action="set_value" to backend.set_value(), but: - ComputerUseBackend did not declare set_value as @abstractmethod, so subclasses could silently omit it without a TypeError at class load time. - _NoopBackend (the test/CI stub) had no set_value method at all, causing AttributeError in any test that exercises the set_value action path. Fix: - Add set_value as @abstractmethod to ComputerUseBackend in backend.py. - Add a recording stub in _NoopBackend in tool.py. - Add two TestDispatch cases: one verifying the call reaches the backend, one verifying the missing-value guard returns a clean error. |
||
|
|
850413f120 |
feat(computer-use): cua-driver backend, universal any-model schema
Background macOS desktop control via cua-driver MCP — does NOT steal the user's cursor or keyboard focus, works with any tool-capable model. Replaces the Anthropic-native `computer_20251124` approach from the abandoned #4562 with a generic OpenAI function-calling schema plus SOM (set-of-mark) captures so Claude, GPT, Gemini, and open models can all drive the desktop via numbered element indices. - `tools/computer_use/` package — swappable ComputerUseBackend ABC + CuaDriverBackend (stdio MCP client to trycua/cua's cua-driver binary). - Universal `computer_use` tool with one schema for all providers. Actions: capture (som/vision/ax), click, double_click, right_click, middle_click, drag, scroll, type, key, wait, list_apps, focus_app. - Multimodal tool-result envelope (`_multimodal=True`, OpenAI-style `content: [text, image_url]` parts) that flows through handle_function_call into the tool message. Anthropic adapter converts into native `tool_result` image blocks; OpenAI-compatible providers get the parts list directly. - Image eviction in convert_messages_to_anthropic: only the 3 most recent screenshots carry real image data; older ones become text placeholders to cap per-turn token cost. - Context compressor image pruning: old multimodal tool results have their image parts stripped instead of being skipped. - Image-aware token estimation: each image counts as a flat 1500 tokens instead of its base64 char length (~1MB would have registered as ~250K tokens before). - COMPUTER_USE_GUIDANCE system-prompt block — injected when the toolset is active. - Session DB persistence strips base64 from multimodal tool messages. - Trajectory saver normalises multimodal messages to text-only. - `hermes tools` post-setup installs cua-driver via the upstream script and prints permission-grant instructions. - CLI approval callback wired so destructive computer_use actions go through the same prompt_toolkit approval dialog as terminal commands. - Hard safety guards at the tool level: blocked type patterns (curl|bash, sudo rm -rf, fork bomb), blocked key combos (empty trash, force delete, lock screen, log out). - Skill `apple/macos-computer-use/SKILL.md` — universal (model-agnostic) workflow guide. - Docs: `user-guide/features/computer-use.md` plus reference catalog entries. 44 new tests in tests/tools/test_computer_use.py covering schema shape (universal, not Anthropic-native), dispatch routing, safety guards, multimodal envelope, Anthropic adapter conversion, screenshot eviction, context compressor pruning, image-aware token estimation, run_agent helpers, and universality guarantees. 469/469 pass across tests/tools/test_computer_use.py + the affected agent/ test suites. - `model_tools.py` provider-gating: the tool is available to every provider. Providers without multi-part tool message support will see text-only tool results (graceful degradation via `text_summary`). - Anthropic server-side `clear_tool_uses_20250919` — deferred; client-side eviction + compressor pruning cover the same cost ceiling without a beta header. - macOS only. cua-driver uses private SkyLight SPIs (SLEventPostToPid, SLPSPostEventRecordTo, _AXObserverAddNotificationAndCheckRemote) that can break on any macOS update. Pin with HERMES_CUA_DRIVER_VERSION. - Requires Accessibility + Screen Recording permissions — the post-setup prints the Settings path. Supersedes PR #4562 (pyautogui/Quartz foreground backend, Anthropic- native schema). Credit @0xbyt4 for the original #3816 groundwork whose context/eviction/token design is preserved here in generic form. |
||
|
|
e63364b8df |
revert: computer-use cua-driver (PR #16919) (#16927)
Reverts PR #16919 (commits |
||
|
|
dad10a78d0 |
feat(computer-use): cua-driver backend, universal any-model schema
Background macOS desktop control via cua-driver MCP — does NOT steal the user's cursor or keyboard focus, works with any tool-capable model. Replaces the Anthropic-native `computer_20251124` approach from the abandoned #4562 with a generic OpenAI function-calling schema plus SOM (set-of-mark) captures so Claude, GPT, Gemini, and open models can all drive the desktop via numbered element indices. - `tools/computer_use/` package — swappable ComputerUseBackend ABC + CuaDriverBackend (stdio MCP client to trycua/cua's cua-driver binary). - Universal `computer_use` tool with one schema for all providers. Actions: capture (som/vision/ax), click, double_click, right_click, middle_click, drag, scroll, type, key, wait, list_apps, focus_app. - Multimodal tool-result envelope (`_multimodal=True`, OpenAI-style `content: [text, image_url]` parts) that flows through handle_function_call into the tool message. Anthropic adapter converts into native `tool_result` image blocks; OpenAI-compatible providers get the parts list directly. - Image eviction in convert_messages_to_anthropic: only the 3 most recent screenshots carry real image data; older ones become text placeholders to cap per-turn token cost. - Context compressor image pruning: old multimodal tool results have their image parts stripped instead of being skipped. - Image-aware token estimation: each image counts as a flat 1500 tokens instead of its base64 char length (~1MB would have registered as ~250K tokens before). - COMPUTER_USE_GUIDANCE system-prompt block — injected when the toolset is active. - Session DB persistence strips base64 from multimodal tool messages. - Trajectory saver normalises multimodal messages to text-only. - `hermes tools` post-setup installs cua-driver via the upstream script and prints permission-grant instructions. - CLI approval callback wired so destructive computer_use actions go through the same prompt_toolkit approval dialog as terminal commands. - Hard safety guards at the tool level: blocked type patterns (curl|bash, sudo rm -rf, fork bomb), blocked key combos (empty trash, force delete, lock screen, log out). - Skill `apple/macos-computer-use/SKILL.md` — universal (model-agnostic) workflow guide. - Docs: `user-guide/features/computer-use.md` plus reference catalog entries. 44 new tests in tests/tools/test_computer_use.py covering schema shape (universal, not Anthropic-native), dispatch routing, safety guards, multimodal envelope, Anthropic adapter conversion, screenshot eviction, context compressor pruning, image-aware token estimation, run_agent helpers, and universality guarantees. 469/469 pass across tests/tools/test_computer_use.py + the affected agent/ test suites. - `model_tools.py` provider-gating: the tool is available to every provider. Providers without multi-part tool message support will see text-only tool results (graceful degradation via `text_summary`). - Anthropic server-side `clear_tool_uses_20250919` — deferred; client-side eviction + compressor pruning cover the same cost ceiling without a beta header. - macOS only. cua-driver uses private SkyLight SPIs (SLEventPostToPid, SLPSPostEventRecordTo, _AXObserverAddNotificationAndCheckRemote) that can break on any macOS update. Pin with HERMES_CUA_DRIVER_VERSION. - Requires Accessibility + Screen Recording permissions — the post-setup prints the Settings path. Supersedes PR #4562 (pyautogui/Quartz foreground backend, Anthropic- native schema). Credit @0xbyt4 for the original #3816 groundwork whose context/eviction/token design is preserved here in generic form. |