Files
hermes-agent/tests/tools/test_tool_search_multiquery.py
Siddharth Balyan b4d04eb8fd Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding

* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg

The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.

On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.

* docs(tool-search): connectors section — remote tools through the bridge

The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.

* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots

dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.

The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.

The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).

Live, 311 local tools + gateway, before -> after:
  "send gmail email":           5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
  "read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
  "linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.

* refactor(tool-search): connector leg into tools/connector_search.py

tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.

No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.

* fix(tool-search): at most 7 queries per call, the gateway's search limit

One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.

The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.

* fix(tool-search): the model is told that connectors__ names are manage_connections accounts

tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.

The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.

Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.

Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.

* fix(connectors): /stop halts a connector batch before the next remote call

dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.

The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.

Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.

* test(connections): schema assertions become dispatch contracts

test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.

Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.

Test count in the file goes from 26 to 25.

* docs(tool-search): connector batches are one gateway request per entry

The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.

Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.

Docs only, no test.

* fix(tools): the between-turns refresh never rewrites the bridge tools

The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.

The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.

Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.

* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation

The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.

* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote

bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.

The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.

Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.

* fix(connectors): search keeps the twin a colliding name reaches, and says so

format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.

Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
2026-09-10 02:21:16 +05:30

565 lines
23 KiB
Python

"""Multi-query ``tool_search``, batched ``tool_describe``, and stemming.
Covers the upgrade that replaced the single ``query`` string with
``queries: [str, ...]`` (grouped, split-shape response), the single
``name`` with ``names: [str, ...]`` (map response with ``not_found``),
and added Snowball stemming to the shared tokenizer.
"""
import json
from concurrent.futures import ThreadPoolExecutor
import pytest
def _td(name, desc, props=None, required=None):
return {
"type": "function",
"function": {
"name": name,
"description": desc,
"parameters": {
"type": "object",
"properties": props or {},
"required": required or [],
},
},
}
def _register(name, toolset, desc="Deferred capability.", props=None, required=None):
from tools.registry import registry
registry.register(
name=name,
handler=lambda args, **kw: json.dumps({"ok": True}),
schema=_td(name, desc, props, required),
toolset=toolset,
)
return _td(name, desc, props, required)
@pytest.fixture
def issue_defs():
"""A small deferred catalog registered under an MCP toolset."""
return [
_register("mq_linear_create_issue", "mcp-mq-linear",
"Create a new issue in a team.",
{"title": {"type": "string"}, "team": {"type": "string"}},
["title", "team"]),
_register("mq_linear_list_issues", "mcp-mq-linear",
"List issues in the workspace.",
{"query": {"type": "string"}}),
_register("mq_slack_post_message", "mcp-mq-slack",
"Post a message to a channel.",
{"channel": {"type": "string"}, "text": {"type": "string"}},
["channel", "text"]),
]
# ---------------------------------------------------------------------------
# Stemming
# ---------------------------------------------------------------------------
class TestStemming:
def test_tokenize_stems_index_and_query_identically(self):
from tools.tool_search_catalog import _tokenize
# Same stem on both sides is the whole contract.
assert _tokenize("issues") == _tokenize("issue")
assert _tokenize("creating messages") == _tokenize("create message")
def test_plural_query_finds_singular_tool_name(self, issue_defs):
"""The measured miss on the old tokenizer: 'issues' skipped create_issue."""
from tools.tool_search import build_catalog, search_catalog
catalog = build_catalog(issue_defs)
names = [h.name for h in search_catalog(catalog, "issues", limit=5)]
assert "mq_linear_create_issue" in names
assert "mq_linear_list_issues" in names
def test_rarest_query_token_gates_admission(self, issue_defs):
"""A document that lacks the query's rarest token is not a result, however many
common tokens it shares. In this catalog 'issue' and 'linear' are each in two tools
and 'slack' in one, so 'slack' gates: the two linear tools share two of the three
query tokens and still do not come back."""
from tools.tool_search import build_catalog, search_catalog
catalog = build_catalog(issue_defs)
names = [h.name for h in search_catalog(catalog, "linear issue slack", limit=5)]
assert names == ["mq_slack_post_message"]
def test_token_no_document_carries_admits_nothing(self, issue_defs):
"""'send gmail email' against a catalog with no gmail tool returns nothing rather
than five tools that merely share 'message' or 'email'."""
from tools.tool_search import build_catalog, search_catalog
catalog = build_catalog(issue_defs)
assert search_catalog(catalog, "post gmail message", limit=5) == []
def test_single_token_stems_are_cached(self):
from tools.tool_search_catalog import _stem, _tokenize
_stem.cache_clear()
corpus = "issues creating issues creating"
_tokenize(corpus)
hits_before = _stem.cache_info().hits
_tokenize(corpus)
assert _stem.cache_info().hits > hits_before
assert _stem.cache_info().hits > 0
assert _tokenize("issues creating") == ["issu", "creat"]
def test_parallel_tokenize_search_and_dispatch_are_deterministic(self, issue_defs):
from tools.tool_search import (
ToolSearchConfig,
build_catalog,
dispatch_tool_search,
search_catalog,
)
from tools.tool_search_catalog import _stem, _tokenize
corpus = (
"issues",
"issue",
"creating",
"create",
"meetings",
"meeting",
"post slack message",
"messages posted",
)
catalog = build_catalog(issue_defs)
expected = {
text: (
_tokenize(text),
[entry.name for entry in search_catalog(catalog, text, limit=3)],
)
for text in corpus
}
def tokenize_and_search(index):
text = corpus[index % len(corpus)]
return text, _tokenize(text), [
entry.name for entry in search_catalog(catalog, text, limit=3)
]
_stem.cache_clear()
misses_before = _stem.cache_info().misses
with ThreadPoolExecutor(max_workers=8) as pool:
threaded = list(pool.map(tokenize_and_search, range(512)))
assert _stem.cache_info().misses > misses_before
for text, tokens, names in threaded:
assert (tokens, names) == expected[text]
args = {"queries": ["issues", "post slack message", "meetings"]}
config = ToolSearchConfig.from_raw({})
expected_json = dispatch_tool_search(
args,
current_tool_defs=issue_defs,
config=config,
)
def dispatch(_index):
return dispatch_tool_search(
args,
current_tool_defs=issue_defs,
config=config,
)
with ThreadPoolExecutor(max_workers=8) as pool:
dispatched = list(pool.map(dispatch, range(64)))
assert dispatched == [expected_json] * 64
def test_stemmer_is_safe_under_concurrent_cache_misses(self):
"""Hammer the raw stemmer from 8 threads with cache-missing input.
``_stem``'s lru_cache means a small corpus warms after a handful of
misses and later iterations never reach the stemmer, so a shared
(non-thread-local) stemmer instance can survive a threaded test over
repeated tokens. This test bypasses the cache: every call stems a
unique token via ``_stem.__wrapped__``, so thousands of stems execute
concurrently on the underlying per-thread instances. A shared
stemmer's mutable parse state produces wrong stems or raises here.
"""
from tools.tool_search_catalog import _stem
words = ["issues", "creating", "meetings", "categories", "searching"]
def serial_baseline(salt):
return [
_stem.__wrapped__(f"{word}x{salt}n{i}")
for i, word in enumerate(words)
]
expected = {salt: serial_baseline(salt) for salt in range(400)}
def worker(salt):
return salt, [
_stem.__wrapped__(f"{word}x{salt}n{i}")
for i, word in enumerate(words)
]
with ThreadPoolExecutor(max_workers=8) as pool:
for salt, stems in pool.map(worker, range(400)):
assert stems == expected[salt]
# ---------------------------------------------------------------------------
# Exact-name ranking and shared corpus statistics
# ---------------------------------------------------------------------------
class TestCatalogRanking:
def test_exact_name_beats_shorter_siblings(self):
from tools.tool_search import build_catalog, search_catalog
exact = _td(
"github_create_issue",
"Create a new issue with a title, body, assignees, labels, "
"milestone, project metadata, and linked context for a repository.",
)
catalog = build_catalog([
exact,
_td("github_create_issue_comment", "Comment."),
_td("github_create_issue_label", "Label."),
])
assert search_catalog(catalog, "github_create_issue", limit=1) == [catalog[0]]
def test_exact_short_name_beats_prefixed_names(self):
from tools.tool_search import build_catalog, search_catalog
catalog = build_catalog([
_td("list", "List one item."),
_td("list_x", "List x."),
_td("list_all_the_open_items", "List every open item."),
])
assert search_catalog(catalog, "list", limit=1) == [catalog[0]]
def test_precomputed_corpus_stats_preserve_results(self, issue_defs):
from tools.tool_search import build_catalog, search_catalog
from tools.tool_search_catalog import _corpus_stats
catalog = build_catalog(issue_defs)
expected = search_catalog(catalog, "create issues", limit=3)
actual = search_catalog(
catalog,
"create issues",
limit=3,
corpus_stats=_corpus_stats(catalog),
)
assert actual == expected
# ---------------------------------------------------------------------------
# Multi-query dispatch_tool_search
# ---------------------------------------------------------------------------
class TestMultiQuerySearch:
def test_grouped_names_plus_shared_tool_map(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_search
result = json.loads(dispatch_tool_search(
{"queries": ["create linear issue", "post slack message"]},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert result["queries"] == ["create linear issue", "post slack message"]
assert result["total_available"] == 3
# Groups carry NAMES only, in query order.
assert [g["query"] for g in result["results"]] == result["queries"]
for group in result["results"]:
for name in group["matches"]:
assert isinstance(name, str)
assert "mq_linear_create_issue" in result["results"][0]["matches"]
assert "mq_slack_post_message" in result["results"][1]["matches"]
# The shared map holds each matched tool exactly once, and nothing else.
matched = {n for g in result["results"] for n in g["matches"]}
assert set(result["tools"]) == matched
record = result["tools"]["mq_linear_create_issue"]
assert record["source"] == "mcp"
assert record["source_name"] == "mcp-mq-linear"
assert record["description"].startswith("Create a new issue")
assert record["required"] == ["title", "team"]
# All queries matched → no fallback block.
assert "available_sources" not in result
assert "hint" not in result
def test_limit_applies_per_query(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_search
result = json.loads(dispatch_tool_search(
{"queries": ["issues", "message"], "limit": 1},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
for group in result["results"]:
assert len(group["matches"]) <= 1
def test_required_names_are_bounded(self):
from tools.tool_search import ToolSearchConfig, dispatch_tool_search
required = [f"field_{index}_" + ("x" * 5000) for index in range(200)]
name = "mq_bounded_required_fields"
tool_def = _register(name, "mcp-mq-bounds", required=required)
result = json.loads(dispatch_tool_search(
{"queries": [name]},
current_tool_defs=[tool_def],
config=ToolSearchConfig.from_raw({}),
))
record = result["tools"][name]
assert len(record["required"]) <= 32
assert all(len(item) <= 64 for item in record["required"])
@pytest.mark.parametrize("schema", [
{"function": "not an object"},
{"function": {"parameters": ["not", "an", "object"]}},
])
def test_shared_record_handles_non_object_schema_fields(self, schema):
from tools.tool_search import CatalogEntry, _shared_tool_record
entry = CatalogEntry(
name="mq_malformed_schema",
description="Malformed schema fixture.",
schema=schema,
source="mcp",
source_name="mcp-mq-malformed",
)
assert _shared_tool_record(entry)["required"] == []
def test_partial_miss_adds_fallback_to_empty_group(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_search
result = json.loads(dispatch_tool_search(
{"queries": ["issues", "zzzz nonsense qqqq"]},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert result["results"][1]["matches"] == []
assert "available_sources" not in result["results"][0]
assert "hint" not in result["results"][0]
missed = result["results"][1]
assert "This query returned no lexical matches" in missed["hint"]
source_names = {s["name"] for s in missed["available_sources"]}
assert {"mq-linear", "mq-slack"} <= source_names
assert "available_sources" not in result
assert "hint" not in result
def test_bare_string_query_coerced_to_single_query(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_search
result = json.loads(dispatch_tool_search(
{"queries": "post slack message"},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert result["queries"] == ["post slack message"]
def test_max_query_cap_respected(self, issue_defs, monkeypatch):
import tools.tool_search as tool_search
monkeypatch.setattr(tool_search, "_MAX_QUERIES_PER_CALL", 2)
cfg = tool_search.ToolSearchConfig.from_raw({})
ok = json.loads(tool_search.dispatch_tool_search(
{"queries": ["a b", "c d"]}, current_tool_defs=issue_defs, config=cfg))
assert "error" not in ok
over = json.loads(tool_search.dispatch_tool_search(
{"queries": ["a", "b", "c"]}, current_tool_defs=issue_defs, config=cfg))
assert "too many queries" in over["error"]
# ---------------------------------------------------------------------------
# Batched dispatch_tool_describe
# ---------------------------------------------------------------------------
class TestBatchedDescribe:
def test_map_response_with_not_found(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
# Deferrable in the global registry, but NOT in this session's defs —
# the stale/out-of-scope case that lands in not_found.
_register("mq_out_of_scope_op", "mcp-mq-elsewhere")
result = json.loads(dispatch_tool_describe(
{"names": ["mq_linear_create_issue", "mq_slack_post_message",
"mq_out_of_scope_op", "mcp__bogus__missing"]},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert set(result["tools"]) == {"mq_linear_create_issue",
"mq_slack_post_message"}
schema = result["tools"]["mq_linear_create_issue"]
assert schema["description"] == "Create a new issue in a team."
assert schema["parameters"]["required"] == ["title", "team"]
# Deferrable-but-absent and unknown names collect in not_found; found
# ones still resolve.
assert result["not_found"] == ["mq_out_of_scope_op", "mcp__bogus__missing"]
assert "tool_search" in result["hint"]
assert "errors" not in result
def test_real_schemas_and_unknown_name_are_classified_independently(self):
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
tool_defs = [
_register("mcp__linear__get_issue", "mcp-linear"),
_register("mcp__granola__list_meeting_folders", "mcp-granola"),
]
result = json.loads(dispatch_tool_describe(
{
"names": [
"mcp__linear__get_issue",
"mcp__granola__list_meeting_folders",
"mcp__linear__does_not_exist_zzz",
]
},
current_tool_defs=tool_defs,
config=ToolSearchConfig.from_raw({}),
))
assert set(result["tools"]) == {
"mcp__linear__get_issue",
"mcp__granola__list_meeting_folders",
}
assert result["not_found"] == ["mcp__linear__does_not_exist_zzz"]
assert "errors" not in result
def test_unregistered_core_name_is_not_found(self, issue_defs, monkeypatch):
from tools import registry as registry_module
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
# The intent: a name that is NOT registered lands in not_found, even
# when it looks like a core tool. Whether "terminal" is registered in
# this process depends on which test files imported model_tools
# earlier, so force the unregistered condition instead of relying on
# collection order.
real_get_entry = registry_module.registry.get_entry
monkeypatch.setattr(
registry_module.registry,
"get_entry",
lambda name: None if name == "terminal" else real_get_entry(name),
)
result = json.loads(dispatch_tool_describe(
{"names": ["terminal", "mq_linear_create_issue"]},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert "mq_linear_create_issue" in result["tools"]
assert "terminal" in result["not_found"]
assert "errors" not in result
def test_registered_direct_surface_name_keeps_exact_error(self):
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
name = "mq_desktop_direct_action"
tool_def = _register(name, "desktop_ui")
result = json.loads(dispatch_tool_describe(
{"names": [name]},
current_tool_defs=[tool_def],
config=ToolSearchConfig.from_raw({}),
))
assert result["errors"][name] == (
f"'{name}' is not a deferrable tool. If you see it in the tools list "
"already, call it directly; otherwise check the spelling against tool_search."
)
assert name not in result.get("not_found", [])
def test_registry_lookup_failure_is_not_found(self, monkeypatch):
from tools.registry import registry
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
def fail_lookup(name):
raise RuntimeError("registry unavailable")
monkeypatch.setattr(registry, "get_entry", fail_lookup)
result = json.loads(dispatch_tool_describe(
{"names": ["mq_unknown_during_lookup"]},
current_tool_defs=[],
config=ToolSearchConfig.from_raw({}),
))
assert result["not_found"] == ["mq_unknown_during_lookup"]
assert "errors" not in result
def test_duplicates_deduped_silently(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
result = json.loads(dispatch_tool_describe(
{"names": ["mq_linear_create_issue", "mq_linear_create_issue"]},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert list(result["tools"]) == ["mq_linear_create_issue"]
assert "not_found" not in result
def test_empty_and_overcap_names_error(self, issue_defs, monkeypatch):
import tools.tool_search as tool_search
monkeypatch.setattr(tool_search, "_MAX_DESCRIBE_NAMES_PER_CALL", 2)
cfg = tool_search.ToolSearchConfig.from_raw({})
assert "error" in json.loads(tool_search.dispatch_tool_describe(
{}, current_tool_defs=issue_defs, config=cfg))
assert "error" in json.loads(tool_search.dispatch_tool_describe(
{"names": []}, current_tool_defs=issue_defs, config=cfg))
over = ["n%d" % i for i in range(3)]
parsed = json.loads(tool_search.dispatch_tool_describe(
{"names": over}, current_tool_defs=issue_defs, config=cfg))
assert "too many names" in parsed["error"]
def test_bare_string_name_coerced(self, issue_defs):
from tools.tool_search import ToolSearchConfig, dispatch_tool_describe
result = json.loads(dispatch_tool_describe(
{"names": "mq_linear_create_issue"},
current_tool_defs=issue_defs,
config=ToolSearchConfig.from_raw({}),
))
assert "mq_linear_create_issue" in result["tools"]
# ---------------------------------------------------------------------------
# Config + bridge schema
# ---------------------------------------------------------------------------
class TestConfigAndSchema:
def test_limit_default_within_cap(self):
from hermes_cli.config_defaults import DEFAULT_CONFIG
from tools.tool_search import ToolSearchConfig
cfg = ToolSearchConfig.from_raw(DEFAULT_CONFIG["tools"]["tool_search"])
assert cfg.max_search_limit == 25
assert cfg.search_default_limit == 5
assert 1 <= cfg.search_default_limit <= cfg.max_search_limit <= 50
def test_bridge_schema_declares_array_inputs(self):
from tools.tool_search import bridge_tool_schemas
schemas = {s["function"]["name"]: s["function"] for s in bridge_tool_schemas(3)}
search_params = schemas["tool_search"]["parameters"]
assert search_params["required"] == ["queries"]
assert search_params["properties"]["queries"]["type"] == "array"
query_description = search_params["properties"]["queries"]["description"]
assert "single string is accepted" in query_description
assert "one query" in query_description
limit_description = search_params["properties"]["limit"]["description"]
assert "per query" in limit_description
assert "configured maximum (25 by default)" in limit_description
describe_params = schemas["tool_describe"]["parameters"]
assert describe_params["required"] == ["names"]
assert describe_params["properties"]["names"]["type"] == "array"
name_description = describe_params["properties"]["names"]["description"]
assert "single string is accepted" in name_description
assert "one name" in name_description