diff options
Diffstat (limited to 'README.md')
| -rw-r--r-- | README.md | 144 |
1 files changed, 97 insertions, 47 deletions
@@ -21,10 +21,56 @@ the model picks. | `query` | string, required | As you would type it into a search box | | `numResults` | integer, optional | 1–20, default 10 | | `recency` | `day` \| `week` \| `month` \| `year`, optional | Omit for no time limit | +| `engine` | `google` \| `duckduckgo` \| `bing` \| `brave`, optional | Overrides the configured default for one call | It returns titles, URLs and snippets — **search results only**. It does not fetch the linked pages; follow up with a read/fetch tool for full content. +## Engines + +Four, with independent indexes and independent rate limits. That last part is +the point: when one starts serving captchas, the others generally still work, +and the tool tells the agent so inside the error. + +| Engine | Recency windows | Notes | +|---|---|---| +| `google` | day, week, month, year | Best results; blocks aggressively under automation | +| `duckduckgo` | day, week, month, year | No-JS endpoint — the most stable markup here, least likely to challenge | +| `bing` | day, week, month | No year window exists | +| `brave` | none | Independent index, unwrapped links; challenges quickly | + +Aliases are accepted: `ddg`, `duck`, `g`, `b`. + +**An unsupported recency window is refused, not ignored.** Asking Bing for the +past year raises `RecencyUnsupportedError` naming the engines that can do it. +Silently returning unfiltered results would be indistinguishable from success. + +### Switching engines + +Exactly like switching the CDP host, plus a per-call override: + +``` +/search-engine show the current engine and where it came from +/search-engine ddg persist a choice (machine-wide, survives restarts) +/search-engine default back to Google +CCS_SEARCH_ENGINE=bing pi … pin for one process +``` + +Precedence: **`engine` parameter > `CCS_SEARCH_ENGINE` > `/search-engine` > +Google.** The stored choice lives in +`~/.pi/agent/castle-cdp-search-engine.json`, written atomically. + +Note the tool parameter sits *above* the environment variable, which is the +opposite of how the CDP target treats an explicit endpoint. That is deliberate: +for the browser, an env var must win because quietly driving a different machine +means automating the operator's own signed-in Chrome — a safety property. +Choosing a different search engine carries no such hazard, and letting the agent +fall back when one engine is blocked is the single most useful thing it can do +with this tool. + +An unrecognised engine name is reported rather than skipped, so a typo in +`CCS_SEARCH_ENGINE` never silently searches Google instead. + ## Which browser it drives Exactly the same configuration as [`pi-browser-harness`](../pi-browser-harness), @@ -61,8 +107,8 @@ transparently on the next search. and aborts. `background: true` so it does not steal focus from whoever is looking at that screen. - **Two concurrent searches**, queued beyond that. -- **20-second deadline** per search, covering connect, navigate and extract - together. Esc aborts an in-flight navigation. +- **20-second deadline** per call, covering the queue wait, connect, navigate and + extract together. Esc aborts an in-flight navigation. - **Every failure throws.** pi only sets `isError: true` when `execute()` throws; a returned `{ error }` object would read to the model as a search that simply found nothing. @@ -72,70 +118,73 @@ transparently on the next search. The browser is shared and belongs to a person. Worth being aware of: - Queries the agent runs land in that browser profile's history and cookies, and - in the Google account's search history if that profile is signed in. + in the search engine's account history if that profile is signed in. - Tabs open and close on someone's screen. `background: true` keeps them from stealing focus, but they are visible. -- Searching hard trips Google's rate limiter for the whole host — a handful of - queries in a few seconds is enough to earn a `/sorry/` page that affects the - human using that browser too. The cap of two concurrent searches limits this; - it does not eliminate it. Measured once: roughly 30 queries in a few minutes - cost about 90 minutes of blocking. -- **A decaying block does not look like a block.** Once the `/sorry/` page stops - being served, Google returns an *empty results page* for a while instead — - which is indistinguishable from a query that genuinely has no hits. It - surfaces as `NoResultsError`, not `SearchChallengeError`. If several unrelated - queries all come back empty, that is rate limiting; wait rather than retrying. +- Searching hard trips rate limiters for the whole host — a handful of queries in + a few seconds is enough to earn a challenge that affects the human using that + browser too. Measured on Google: roughly 30 queries in a few minutes cost about + 90 minutes of blocking. Brave challenges considerably sooner than that. +- **A decaying block does not look like a block.** Once the challenge page stops + being served, engines return an *empty results page* for a while instead — + indistinguishable from a query with no hits. It surfaces as `NoResultsError`, + not `SearchChallengeError`. If one engine comes back empty and another answers + the same query, that is rate limiting. ## When a captcha appears -Google will eventually serve a `/sorry/` interstitial, a consent wall, or a -recaptcha — especially if searches come in fast. The extension detects this and -throws a `SearchChallengeError` naming the challenge, the browser, and the URL a -human has to visit. +Every engine eventually serves an interstitial: Google's `/sorry/` page, Brave's +"Verifying you're not a bot", a consent wall, a Cloudflare challenge. The +extension detects these and throws `SearchChallengeError` naming the challenge, +the engine, the browser, and the URL a human would have to visit. -The agent cannot solve it. The guidelines tell it to stop searching and hand off -to the user, who opens that URL in the browser at the endpoint, clears the -challenge, and lets the agent retry. The cookie is profile-wide, so one pass -unblocks later searches. +The error tells the agent to **try another engine first**, because a challenge on +one engine says nothing about the others and that is a fix it can apply itself. +Only when engines run out should it ask the user to clear the challenge in the +browser. The cookie is profile-wide, so one pass unblocks later searches. The tab is closed rather than left open on the challenge page: leaving it would let the operator solve it in place, but would also litter a shared browser with abandoned tabs on every failure. The URL in the error is enough. -## Search parameters, and two surprises +## Search parameters, and three traps -Plain searches use `udm=14` — Google's "Web" tab: no AI overview, no carousels, -just ranked links, which is both cheaper to parse and closer to what was asked -for. Verified against the real browser: +Everything below was observed against castle's real Chrome, not inferred. All +three fail *silently* — the results look entirely plausible, just wrong. -- `tbs=qdr:*`, the parameter Google's own Tools menu writes, renders an **empty - page** for this profile — `#search` present, no `#rso`, no results. The older - `as_qdr=*` works and genuinely filters. -- Any date restriction combined with `udm=14` also renders that empty page. +1. **Google: `tbs=qdr:*` renders an empty page.** That is the parameter Google's + own Tools menu writes. The older `as_qdr=*` works. +2. **Google: any date filter combined with `udm=14` renders that same empty + page.** So a time-limited search drops `udm` and uses the classic layout. +3. **Bing: `count=` silently cancels `filters=`.** With `ex1:"ez1"` alone every + result is hours old; add `count` in either order and months-old results come + back, unfiltered and unremarkable-looking. So Bing drops `count` whenever a + date filter is present and the caller slices the list instead. -So a time-limited search drops `udm` and uses `as_qdr`. The extractor handles -both layouts. +Each has a regression test asserting the parameter combination, since none of +them would announce itself if it regressed. ## Result extraction -`src/extract.ts` avoids Google's generated class names (`MjjYud`, `kb0PBd`, …) -entirely. It prefers the `data-snhf` / `data-sncf` hooks and otherwise falls back -to a structural walk: from each `<h3>`, take the enclosing link, climb until the -text grows past the header's, and abandon the climb if a second `<h3>` comes into -scope. Checked against both the classic SERP and the `udm=14` layout. +Two modes, because the engines genuinely differ: -Google's markup will change anyway. When it does, a search returns -`NoResultsError` whose message distinguishes "results container rendered but -unparseable" (extractor needs updating) from "no results area at all" (query or -parameters). +- **items** — DuckDuckGo (`.result`), Bing (`li.b_algo`) and Brave + (`.snippet[data-type=web]`) each have a clean per-result container. +- **headings** — Google has no stable container; its class names (`MjjYud`, + `kb0PBd`, `yuRUbf`) are generated. Results are found from each `<h3>` outward, + climbing to the enclosing block but abandoning the climb if a second `<h3>` + comes into scope. -## Install +DuckDuckGo and Bing both route outbound links through redirectors +(`duckduckgo.com/l/?uddg=…`, `bing.com/ck/a?…&u=a1<base64url>`); both are +unwrapped in the page so the agent gets real URLs. Google and Brave link +straight out. -``` -pi install /Volumes/Sense/src/soarez/pi-castle-cdp-search -``` +Markup will change. When it does, `NoResultsError` distinguishes "container +rendered but unparseable" (that engine's extractor needs updating) from "no +container at all" (query, or rate limiting). -Or, once it is on castle alongside the harness: +## Install ``` pi install ssh://sz@10.88.0.25/Users/sz/repos/pi-castle-cdp-search.git @@ -159,8 +208,9 @@ npm run typecheck npm test # unit tests, no browser needed ``` -The tests cover endpoint resolution, URL building, formatting, the challenge -error, and the concurrency semaphore. The semaphore tests point at an +The tests cover engine selection and precedence, URL building for all four +engines (including the three traps above), recency refusal, formatting, the +challenge error, and the concurrency semaphore. The semaphore tests point at an unreachable endpoint on purpose — the questions there are about slot bookkeeping, and hitting a real browser to test a counter would be slow, flaky, and rude to whoever is using it. |
