11 Commits

Author SHA1 Message Date
Teknium c63de5a231 feat(tool_search): long hunts for nonexistent tools now return no results instead of incidental matches
Port from nearai/ironclaw#7965: BM25 admits any document scoring above
zero, i.e. sharing ONE term with the query. A long descriptive search
for a capability that does not exist therefore returned a plausible-
looking ranked list, and the model read 'results exist' as 'it is in
here somewhere' and rephrased instead of stopping (IronClaw production
trace: 652 tool calls, 216 of them tool_search, hunting a 'data' tool
that did not exist).

A document must now match at least half the query's ANSWERABLE terms
(terms present anywhere in the index) before it is offered. Coverage
only engages from four answerable terms up, preserving recall on short
queries; exact tool-name matches remain authoritative; the substring
fallback is unchanged.

Docs: relevance-floor bullet added to tool-search.md implementation
details.
2026-09-13 21:04:08 -07:00
Siddharth Balyan cf4b78e91f fix(tool-search): a query no tool answers returns nothing, not five tools sharing one word (#106676)
search_catalog admitted every document with BM25 score > 0 and then padded
to `limit`. BM25 sums over the tokens a document shares with the query, so
on a 300-tool catalog "send gmail email" returned five incident tools that
shared only "email", and the discriminating word ("gmail", in no document)
had no say. The model read those as the answer.

Admission is now the query's rarest token: a document is a result only if it
contains the query token with the highest IDF, the one that names the intent.
Common verbs ("send", "read", "create") sit in dozens of documents and never
gate; vendor and object words ("gmail", "github", "incident") do. A token no
document carries admits nothing, and the existing empty-group hint tells the
model to retry without it. The name-substring fallback is deleted: it admitted
tools that matched no query token at all.

Result descriptions are clipped at 500 characters instead of 400. Over 353
vendor tool descriptions, 500 keeps 91% whole and every first sentence
(first-sentence max 329); 400 kept 82%.

Measured on the live 311-tool catalog with 25 hand-labelled queries:
precision@5 0.18 -> 0.43, wrong names returned 102 -> 66, false positives on
absent intents 17 -> 13. Live before/after: "send gmail email" went from five
betterstack tools to an empty group with the retry hint; "linear create issue"
and "betterstack incident" are unchanged.
2026-09-10 02:21:15 +05:30
Teknium e0df9656bb refactor(tools): group H — BM25 loop fold, catalog/selection/limits micro-collapses, sync-manager walrus 2026-09-03 01:13:51 -07:00
Teknium 136ac31065 refactor(tools): group H — _registry_toolset helper unifies deferral/source lookups, partial() config loaders, Counter/dict-comprehension folds, legacy-config path collapse 2026-09-03 01:09:12 -07:00
Teknium 0994c01d1d refactor(tools): group H — pack multi-line signatures, config-default one-liners, exact-match score fold 2026-09-03 01:03:20 -07:00
Teknium 5645bc3cd0 refactor(tools): fold tool_search_names into tool_search_catalog (single importer pair), dict-dispatch modal backend selection, compact registry.register calls and shebangs 2026-09-03 00:59:55 -07:00
Teknium dda3aa81d5 refactor(tools): tool_search_catalog — fold guard chains, drop single-use locals, compact BM25/listing docstrings 2026-09-03 00:41:42 -07:00
Teknium cffc1df516 refactor(tools): group H layout compaction (AST-neutral bracket hugging) 2026-09-02 22:42:44 -07:00
Teknium c7bae88fce refactor(tools): tool_search — drop dead build_catalog_listing/_describe_classification/_safe_int, fold describe classification into is_deferrable_tool_name, compact docstrings 2026-09-02 22:18:08 -07:00
Teknium ae67178be1 refactor(tools/infra): split cronjob god-methods into per-action helpers; extract tool_search catalog/validation; compact registry, lazy_deps, tool_backend_helpers, desktop_ui 2026-09-02 14:47:02 -07:00
Teknium d4cec15b47 refactor(tools): first-wave simplification of tools/ (file ops split, lazy_deps, code_exec, approval, browser, delegate, mcp, skills, terminal, voice, media)
Behavior-neutral structural pass over tools/*: god-file extractions into
sibling modules (file_operations_common/lint/search, file_tools_paths/
read_tracking/write, code_execution_env/rpc, tool_search_catalog/names/
validation, tts_command_provider, ...), duplicate helper unification,
if/elif -> dispatch tables, dead-code removal, docstring compaction.
Tool schemas (get_tool_definitions) verified byte-identical to base.
2026-09-02 14:43:45 -07:00