<feed xmlns='http://www.w3.org/2005/Atom'>
<title>pi-castle-cdp-search.git/src/engines.ts, branch master</title>
<subtitle>Search castle's Chrome tabs and history over CDP</subtitle>
<id>https://git.soarez.net/sz/pi-castle-cdp-search.git/atom/src/engines.ts?h=master</id>
<link rel='self' href='https://git.soarez.net/sz/pi-castle-cdp-search.git/atom/src/engines.ts?h=master'/>
<link rel='alternate' type='text/html' href='https://git.soarez.net/sz/pi-castle-cdp-search.git/'/>
<updated>2026-08-03T21:17:46Z</updated>
<entry>
<title>Record measured rate-limit behaviour per engine</title>
<updated>2026-08-03T21:17:46Z</updated>
<author>
<name>Igor Soarez</name>
<email>igor@soarez.org</email>
</author>
<published>2026-08-03T21:17:46Z</published>
<link rel='alternate' type='text/html' href='https://git.soarez.net/sz/pi-castle-cdp-search.git/commit/?id=92bc4b3cba1beb7aa443c4dcbb1e741ece3a5deb'/>
<id>urn:sha1:92bc4b3cba1beb7aa443c4dcbb1e741ece3a5deb</id>
<content type='text'>
Concrete numbers from development rather than a vague warning: Google
tolerates a lot then blocks for ~90 minutes, Brave challenges after one or
two queries and clears in ~15-20, and DuckDuckGo and Bing never challenged
at all. That last fact is the useful one — it says which engine to reach
for when several searches are needed in a row.
</content>
</entry>
<entry>
<title>Take Brave snippets by position rather than by class name</title>
<updated>2026-08-03T21:09:48Z</updated>
<author>
<name>Igor Soarez</name>
<email>igor@soarez.org</email>
</author>
<published>2026-08-03T21:09:48Z</published>
<link rel='alternate' type='text/html' href='https://git.soarez.net/sz/pi-castle-cdp-search.git/commit/?id=7a8f261ec50b9c15bc37ba8ae1f0022c392111d3'/>
<id>urn:sha1:7a8f261ec50b9c15bc37ba8ae1f0022c392111d3</id>
<content type='text'>
Brave's source name was leaking into every snippet ("Medium March 27,
2025 - ..."), because the .sitename/.netloc selectors guessed for it match
nothing. Rather than guess again, use the ordering: an engine rendering
"source / breadcrumb / title / description" puts all its metadata before
the title, so everything after the title line is the description. Class
names churn; that ordering does not. The named selectors stay as the
fallback for when the title is not on a line of its own.

Verified against Brave, and Google/DuckDuckGo/Bing re-checked for
regressions.
</content>
</entry>
<entry>
<title>Support DuckDuckGo, Bing and Brave, switchable like the CDP host</title>
<updated>2026-08-03T20:43:56Z</updated>
<author>
<name>Igor Soarez</name>
<email>igor@soarez.org</email>
</author>
<published>2026-08-03T20:43:56Z</published>
<link rel='alternate' type='text/html' href='https://git.soarez.net/sz/pi-castle-cdp-search.git/commit/?id=db69207c06e8d5233bff4996e3ef5b43332d533f'/>
<id>urn:sha1:db69207c06e8d5233bff4996e3ef5b43332d533f</id>
<content type='text'>
Engine resolution mirrors the browser target: CCS_SEARCH_ENGINE, then a
/search-engine choice persisted machine-wide, then Google. A per-call
`engine` parameter sits above both so the agent can fall back when one
engine starts serving captchas — the one case where the model, not the
operator, has to make the call. Unlike the CDP target there is no safety
argument for the environment winning: driving the wrong browser means
automating someone's signed-in Chrome, choosing a different index does not.

An unsupported recency window is refused, naming the engines that support
it, rather than dropped. Silently returning unfiltered results is
indistinguishable from success, which is the failure this whole design is
trying to avoid.

Extraction grows a second mode. DuckDuckGo, Bing and Brave have clean
per-result containers; Google does not, so its heading-walk stays as its
own path rather than being bent into the item shape. DuckDuckGo and Bing
route links through redirectors, unwrapped in the page.

Third silent-failure trap found, alongside Google's two: on Bing, `count`
cancels `filters`. With ex1:"ez1" alone every result is hours old; add
count in either order and months-old results return, looking perfectly
ordinary. Bing now drops count whenever a date filter is present.

Challenge detection widened to Brave's "Verifying you're not a bot" and
"Quick check before you continue searching", which the previous Google-
shaped matcher missed entirely — found by tripping it.
</content>
</entry>
</feed>
