summaryrefslogtreecommitdiff
path: root/README.md
diff options
context:
space:
mode:
Diffstat (limited to 'README.md')
-rw-r--r--README.md144
1 files changed, 97 insertions, 47 deletions
diff --git a/README.md b/README.md
index a97d105..684bd96 100644
--- a/README.md
+++ b/README.md
@@ -21,10 +21,56 @@ the model picks.
| `query` | string, required | As you would type it into a search box |
| `numResults` | integer, optional | 1–20, default 10 |
| `recency` | `day` \| `week` \| `month` \| `year`, optional | Omit for no time limit |
+| `engine` | `google` \| `duckduckgo` \| `bing` \| `brave`, optional | Overrides the configured default for one call |
It returns titles, URLs and snippets — **search results only**. It does not fetch
the linked pages; follow up with a read/fetch tool for full content.
+## Engines
+
+Four, with independent indexes and independent rate limits. That last part is
+the point: when one starts serving captchas, the others generally still work,
+and the tool tells the agent so inside the error.
+
+| Engine | Recency windows | Notes |
+|---|---|---|
+| `google` | day, week, month, year | Best results; blocks aggressively under automation |
+| `duckduckgo` | day, week, month, year | No-JS endpoint — the most stable markup here, least likely to challenge |
+| `bing` | day, week, month | No year window exists |
+| `brave` | none | Independent index, unwrapped links; challenges quickly |
+
+Aliases are accepted: `ddg`, `duck`, `g`, `b`.
+
+**An unsupported recency window is refused, not ignored.** Asking Bing for the
+past year raises `RecencyUnsupportedError` naming the engines that can do it.
+Silently returning unfiltered results would be indistinguishable from success.
+
+### Switching engines
+
+Exactly like switching the CDP host, plus a per-call override:
+
+```
+/search-engine show the current engine and where it came from
+/search-engine ddg persist a choice (machine-wide, survives restarts)
+/search-engine default back to Google
+CCS_SEARCH_ENGINE=bing pi … pin for one process
+```
+
+Precedence: **`engine` parameter > `CCS_SEARCH_ENGINE` > `/search-engine` >
+Google.** The stored choice lives in
+`~/.pi/agent/castle-cdp-search-engine.json`, written atomically.
+
+Note the tool parameter sits *above* the environment variable, which is the
+opposite of how the CDP target treats an explicit endpoint. That is deliberate:
+for the browser, an env var must win because quietly driving a different machine
+means automating the operator's own signed-in Chrome — a safety property.
+Choosing a different search engine carries no such hazard, and letting the agent
+fall back when one engine is blocked is the single most useful thing it can do
+with this tool.
+
+An unrecognised engine name is reported rather than skipped, so a typo in
+`CCS_SEARCH_ENGINE` never silently searches Google instead.
+
## Which browser it drives
Exactly the same configuration as [`pi-browser-harness`](../pi-browser-harness),
@@ -61,8 +107,8 @@ transparently on the next search.
and aborts. `background: true` so it does not steal focus from whoever is
looking at that screen.
- **Two concurrent searches**, queued beyond that.
-- **20-second deadline** per search, covering connect, navigate and extract
- together. Esc aborts an in-flight navigation.
+- **20-second deadline** per call, covering the queue wait, connect, navigate and
+ extract together. Esc aborts an in-flight navigation.
- **Every failure throws.** pi only sets `isError: true` when `execute()` throws;
a returned `{ error }` object would read to the model as a search that simply
found nothing.
@@ -72,70 +118,73 @@ transparently on the next search.
The browser is shared and belongs to a person. Worth being aware of:
- Queries the agent runs land in that browser profile's history and cookies, and
- in the Google account's search history if that profile is signed in.
+ in the search engine's account history if that profile is signed in.
- Tabs open and close on someone's screen. `background: true` keeps them from
stealing focus, but they are visible.
-- Searching hard trips Google's rate limiter for the whole host — a handful of
- queries in a few seconds is enough to earn a `/sorry/` page that affects the
- human using that browser too. The cap of two concurrent searches limits this;
- it does not eliminate it. Measured once: roughly 30 queries in a few minutes
- cost about 90 minutes of blocking.
-- **A decaying block does not look like a block.** Once the `/sorry/` page stops
- being served, Google returns an *empty results page* for a while instead —
- which is indistinguishable from a query that genuinely has no hits. It
- surfaces as `NoResultsError`, not `SearchChallengeError`. If several unrelated
- queries all come back empty, that is rate limiting; wait rather than retrying.
+- Searching hard trips rate limiters for the whole host — a handful of queries in
+ a few seconds is enough to earn a challenge that affects the human using that
+ browser too. Measured on Google: roughly 30 queries in a few minutes cost about
+ 90 minutes of blocking. Brave challenges considerably sooner than that.
+- **A decaying block does not look like a block.** Once the challenge page stops
+ being served, engines return an *empty results page* for a while instead —
+ indistinguishable from a query with no hits. It surfaces as `NoResultsError`,
+ not `SearchChallengeError`. If one engine comes back empty and another answers
+ the same query, that is rate limiting.
## When a captcha appears
-Google will eventually serve a `/sorry/` interstitial, a consent wall, or a
-recaptcha — especially if searches come in fast. The extension detects this and
-throws a `SearchChallengeError` naming the challenge, the browser, and the URL a
-human has to visit.
+Every engine eventually serves an interstitial: Google's `/sorry/` page, Brave's
+"Verifying you're not a bot", a consent wall, a Cloudflare challenge. The
+extension detects these and throws `SearchChallengeError` naming the challenge,
+the engine, the browser, and the URL a human would have to visit.
-The agent cannot solve it. The guidelines tell it to stop searching and hand off
-to the user, who opens that URL in the browser at the endpoint, clears the
-challenge, and lets the agent retry. The cookie is profile-wide, so one pass
-unblocks later searches.
+The error tells the agent to **try another engine first**, because a challenge on
+one engine says nothing about the others and that is a fix it can apply itself.
+Only when engines run out should it ask the user to clear the challenge in the
+browser. The cookie is profile-wide, so one pass unblocks later searches.
The tab is closed rather than left open on the challenge page: leaving it would
let the operator solve it in place, but would also litter a shared browser with
abandoned tabs on every failure. The URL in the error is enough.
-## Search parameters, and two surprises
+## Search parameters, and three traps
-Plain searches use `udm=14` — Google's "Web" tab: no AI overview, no carousels,
-just ranked links, which is both cheaper to parse and closer to what was asked
-for. Verified against the real browser:
+Everything below was observed against castle's real Chrome, not inferred. All
+three fail *silently* — the results look entirely plausible, just wrong.
-- `tbs=qdr:*`, the parameter Google's own Tools menu writes, renders an **empty
- page** for this profile — `#search` present, no `#rso`, no results. The older
- `as_qdr=*` works and genuinely filters.
-- Any date restriction combined with `udm=14` also renders that empty page.
+1. **Google: `tbs=qdr:*` renders an empty page.** That is the parameter Google's
+ own Tools menu writes. The older `as_qdr=*` works.
+2. **Google: any date filter combined with `udm=14` renders that same empty
+ page.** So a time-limited search drops `udm` and uses the classic layout.
+3. **Bing: `count=` silently cancels `filters=`.** With `ex1:"ez1"` alone every
+ result is hours old; add `count` in either order and months-old results come
+ back, unfiltered and unremarkable-looking. So Bing drops `count` whenever a
+ date filter is present and the caller slices the list instead.
-So a time-limited search drops `udm` and uses `as_qdr`. The extractor handles
-both layouts.
+Each has a regression test asserting the parameter combination, since none of
+them would announce itself if it regressed.
## Result extraction
-`src/extract.ts` avoids Google's generated class names (`MjjYud`, `kb0PBd`, …)
-entirely. It prefers the `data-snhf` / `data-sncf` hooks and otherwise falls back
-to a structural walk: from each `<h3>`, take the enclosing link, climb until the
-text grows past the header's, and abandon the climb if a second `<h3>` comes into
-scope. Checked against both the classic SERP and the `udm=14` layout.
+Two modes, because the engines genuinely differ:
-Google's markup will change anyway. When it does, a search returns
-`NoResultsError` whose message distinguishes "results container rendered but
-unparseable" (extractor needs updating) from "no results area at all" (query or
-parameters).
+- **items** — DuckDuckGo (`.result`), Bing (`li.b_algo`) and Brave
+ (`.snippet[data-type=web]`) each have a clean per-result container.
+- **headings** — Google has no stable container; its class names (`MjjYud`,
+ `kb0PBd`, `yuRUbf`) are generated. Results are found from each `<h3>` outward,
+ climbing to the enclosing block but abandoning the climb if a second `<h3>`
+ comes into scope.
-## Install
+DuckDuckGo and Bing both route outbound links through redirectors
+(`duckduckgo.com/l/?uddg=…`, `bing.com/ck/a?…&u=a1<base64url>`); both are
+unwrapped in the page so the agent gets real URLs. Google and Brave link
+straight out.
-```
-pi install /Volumes/Sense/src/soarez/pi-castle-cdp-search
-```
+Markup will change. When it does, `NoResultsError` distinguishes "container
+rendered but unparseable" (that engine's extractor needs updating) from "no
+container at all" (query, or rate limiting).
-Or, once it is on castle alongside the harness:
+## Install
```
pi install ssh://sz@10.88.0.25/Users/sz/repos/pi-castle-cdp-search.git
@@ -159,8 +208,9 @@ npm run typecheck
npm test # unit tests, no browser needed
```
-The tests cover endpoint resolution, URL building, formatting, the challenge
-error, and the concurrency semaphore. The semaphore tests point at an
+The tests cover engine selection and precedence, URL building for all four
+engines (including the three traps above), recency refusal, formatting, the
+challenge error, and the concurrency semaphore. The semaphore tests point at an
unreachable endpoint on purpose — the questions there are about slot
bookkeeping, and hitting a real browser to test a counter would be slow, flaky,
and rude to whoever is using it.