summaryrefslogtreecommitdiff
path: root/docs
diff options
context:
space:
mode:
Diffstat (limited to 'docs')
-rw-r--r--docs/2026-08-04-ai-answers.md235
1 files changed, 235 insertions, 0 deletions
diff --git a/docs/2026-08-04-ai-answers.md b/docs/2026-08-04-ai-answers.md
new file mode 100644
index 0000000..29b5a6b
--- /dev/null
+++ b/docs/2026-08-04-ai-answers.md
@@ -0,0 +1,235 @@
+# AI answers on the SERP — findings, 2026-08-04
+
+Status: **investigated, not implemented, not decided.** Nothing in `src/` has
+changed. This is the field notes for whoever picks it up.
+
+The question was whether the AI answer block that Google, Bing and Brave now put
+above their results ("AI Overview", the Bing generative answer, Brave's AI
+summary) is worth pulling into `castle_cdp_search`'s output alongside the ranked
+results. It is not obviously worth it — the cost is real and is set out below —
+so the decision was deferred rather than taken.
+
+Everything here was read off the live SERPs in castle's Chrome on 2026-08-04, in
+the signed-in `~/.chrome-cdp-profile`, `hl=en`. Same rule as `engines.ts`: what
+is written down is what was observed. Where something was *not* verified it says
+so.
+
+## What is there, per engine
+
+| Engine | Block | Container observed | Sources exposed |
+| --- | --- | --- | --- |
+| Google | "AI Overview" | `#m-x-content` | yes, with titles |
+| Bing | generative answer | `li.b_ans.b_top.b_qnacdxcont` → `.gs_caphead_main` | yes, cards in `.gs_cit` |
+| Brave | AI summary | `#chatllm-main-answer-content` | no |
+| DuckDuckGo | none | — | — |
+
+DuckDuckGo is a genuine "none", not an unfinished check: the engine points at the
+no-JS `html.duckduckgo.com` endpoint, which cannot render one. There is no
+`.zci`/zero-click block on it either — that was checked directly.
+
+### Google
+
+The overview lives **outside `#search` and `#rso`**, above them. That was
+verified (`rso.contains(node) === false`), and it means the current extractor
+neither sees it nor is disturbed by it — no existing behaviour is at risk here.
+
+Structure of `#m-x-content`, on a query that had one:
+
+```
+#m-x-content
+ [jsname=txosbe] "AI Overview" ← the header
+ [jsname=HKDuG] the answer body ← also contains the citations
+ … answer prose
+ [role=dialog] "6 sites" + cards + "Show all"
+```
+
+Hooks, ranked by how much they can be trusted:
+
+- `#m-x-content` — an id, present on both loads. The most durable thing found.
+ **But see the trap below: the id alone is not sufficient.**
+- `[role=heading]` with the exact text `AI Overview`, and
+ `[aria-label="Show more AI Overview"]` on the collapse control. Semantic, and
+ the natural anchor if the id ever moves.
+- `[role=dialog]` around the citation carousel (`data-type=hovc`).
+- Class names — **useless**, as everywhere else on Google. The answer body was
+ `.LT6XE` on one load and `.jloFI GkDqAd rJweXb` on the next, same query,
+ minutes apart. This is the same churn `engines.ts` already warns about.
+
+**Citations are the good part.** Anchors inside `#m-x-content` carry the source
+title in `aria-label`, suffixed `". Opens in a new tab."`:
+
+```
+aria-label="Open source agentic startup LangChain hits $1.25B valuation | TechCrunch. Opens in a new tab."
+href="https://techcrunch.com/2025/10/21/open-source-agentic-startup-langchain-hits-1-25b-valuation/"
+```
+
+Their `innerText` is empty — the visible label is an icon — so `aria-label` is
+the only way to get a title. Links out directly, no redirector, so the existing
+`unwrap: "none"` holds. Google's own hosts appear among them (a privacy-policy
+link) and are already excluded by `selfHostPattern`.
+
+Two traps:
+
+1. **`#m-x-content` exists before the overview does, holding something else.**
+ Polling a fresh load at 250ms intervals caught it at 1502ms containing 28,864
+ characters — approximately the whole SERP — and at 1752ms containing 142, then
+ growing to its real 903. So the id is reused by, or briefly wraps, a larger
+ shell. An extractor that reads it on sight will occasionally return the entire
+ page as the "AI answer". **Require the text to begin with `AI Overview`, or
+ anchor on the heading and climb, before trusting the node.**
+2. **"Show more" is not hiding any of the answer.** Clicking it took the text
+ from 1124 to 1262 characters, and the 138 added were the disclaimer footer
+ ("Double-check important information… AI responses may include mistakes").
+ The answer body is complete in the DOM while still visually truncated. So
+ there is no need to click anything — which is good, because clicking in a
+ browser a human is using is exactly what this package avoids.
+
+To render usefully, the extractor has to drop the `AI Overview` header line, the
+`[role=dialog]` citation carousel, the "Show all" / "Dismiss" / "Learn how"
+buttons and the disclaimer. Cleanest approach found: `cloneNode(true)`, remove
+those by selector on the clone, take `innerText`. The page itself is never
+mutated, which matters on a shared browser.
+
+### Bing
+
+```
+li.b_ans.b_top.b_qnacdxcont
+ .gs_caphead_main the answer prose ← clean, nothing else in it
+ .gs_cit_cont
+ .gs_cit source card: name, title, snippet
+ .gs_cit …
+ .cit_exp_cont "View all"
+```
+
+The `gs_` prefix ("generative search") is descriptive rather than hashed, and
+`.gs_caphead_main` held the answer and *only* the answer — no header, no cards,
+no disclaimer. The nicest of the three to extract.
+
+`li.b_ans` on its own is **not** a safe selector: "People also ask" is also a
+`li.b_ans`. The discriminator is `b_qnacdxcont` on the `li`, or the presence of
+`.gs_caphead_main` inside it.
+
+Citation anchors go through Bing's `/ck/a?…&u=a1…` redirector, so the existing
+`unwrap: "bing"` already handles them — no new unwrapping needed.
+
+No URL change is needed: the answer appears on the plain
+`bing.com/search?q=…` the engine already builds.
+
+### Brave
+
+`#chatllm-main-answer-content` contained the answer and nothing else — no
+header, no disclaimer (that sits outside it, in `cite.chatllm-disclaimer`), no
+citations. The wrapper is `#llm-snippet`, a `.snippet.standalone` sitting first
+in `#results`.
+
+Worth noting: `engines.ts` already restricts Brave's results to
+`.snippet[data-type="web"]` specifically to exclude this block. That filter
+stays correct — the AI answer would be read separately, not as a result.
+
+No sources are exposed. The only anchors inside are `search.brave.com/ask?…`
+follow-up prompts ("Elaborate", "Who are LangChain's main competitors?"), which
+`selfHostPattern` already discards. There is a "+4 More" control that may reveal
+sources; **not investigated** — it would need a click, and clicking was avoided.
+
+No URL change needed.
+
+## The two reasons this is not free
+
+### 1. Google's AI Overview requires giving up `udm=14`
+
+`udm` is Google's undocumented vertical selector — the parameter its own tab
+strip writes. `udm=2` is Images, `udm=7` Videos, `udm=12` News, `udm=28`
+Shopping, and `udm=14` is the "Web" tab added in May 2024: plain ranked links,
+no AI Overview, no knowledge panel, no carousels. `engines.ts` sets it as
+Google's default precisely for that cleanliness.
+
+So "no AI Overview" is not incidental to `udm=14`, it is part of what the Web tab
+*is*. Every page where `#m-x-content` was found was a no-`udm` classic SERP.
+Enabling overviews on Google means dropping `udm`.
+
+That is less alarming than it sounds: the engine **already** drops `udm` for any
+time-limited search, because `udm=14` plus a date filter renders an empty page
+(README, trap 2). Heading-mode extraction is therefore already exercised against
+the classic layout. The cost is that the classic layout puts back all the
+furniture — knowledge panels, "People also ask", carousels — that the extractor
+then has to keep ignoring.
+
+### 2. They stream in after the results are ready, and are often absent
+
+Measured from navigation start, polling at 250ms via an `initScript`:
+
+| Engine | results parseable | AI answer appears | AI answer settles |
+| --- | --- | --- | --- |
+| Google | ~500ms | ~1500ms | ~2300ms |
+| Brave | ~290ms | ~1260ms | **still growing at 5000ms** |
+| Bing | ~520ms | — (absent on that query) | — |
+
+`search.ts` reads results and returns as soon as anything parses, retrying at
+most twice at 700ms (`EXTRACT_RETRIES`, `EXTRACT_RETRY_DELAY_MS`) and breaking
+the moment `results.length > 0`. In other words it returns roughly 2s before
+Google's overview lands and 5s+ before Brave's finishes. **Capturing an AI answer
+means deliberately waiting past the point where the useful part of the job is
+already done.**
+
+And the wait is frequently wasted, because the blocks are query-dependent:
+
+- `how does postgres MVCC vacuum work` → **no** Google AI Overview at all, still
+ absent after 15s, in a foreground tab. (Checked in a foreground tab
+ specifically to rule out background-tab throttling as the cause — an earlier
+ `window.open` probe was discarded for that reason.)
+- `who founded anthropic and when` → **no** Bing generative answer after 13.5s;
+ Bing served an ordinary big-Wikipedia card instead.
+- `LangChain valuation Series C funding 2026` → all three produced one.
+
+There is no cheap way to know in advance which case you are in, so a fixed wait
+pays the full cost on every query that was never going to have an answer.
+
+## The decision that was not taken
+
+Three shapes were on the table:
+
+1. **Opt-in `aiAnswer` parameter (the one that was going to be recommended).**
+ Default off: today's latency and today's `udm=14` layout, byte-identical
+ behaviour. When true: Google drops `udm`, and the probe polls until the text
+ stops growing or a budget (~6s) expires. Costs nothing unless asked for, and
+ the model can ask for it when the query is a question rather than a lookup.
+ Downside: another parameter on a tool that already has four, and the model has
+ to decide when to set it.
+2. **Always on.** Simplest surface, no new parameter. Every search pays ~2s
+ (Google) to ~5s (Brave), much of it waiting for answers that never arrive,
+ against a 20s whole-call deadline (`SEARCH_TIMEOUT_MS`) that also has to cover
+ connect and navigate.
+3. **Opportunistic, never wait.** Include an AI answer only if it is already
+ rendered when results are read. Zero added latency, but Google would
+ essentially never contribute one (`udm=14` stays), and Brave would usually be
+ caught mid-sentence — see the growth curve above. Risks putting a truncated,
+ confident-sounding half-answer in front of the model, which is worse than
+ putting none there.
+
+Unresolved either way: whether an AI answer belongs in the model's context at
+all. It is a summary generated by another model from pages the agent could read
+itself, it carries no provenance beyond its citations, and `promptGuidelines`
+currently tells the model in plain terms not to treat a snippet as the page.
+An AI answer invites exactly that confusion at greater length. If it goes in, it
+should be labelled unmistakably as the engine's generated text — engine name
+included, e.g. "Google AI Overview" — and never merged into the numbered results.
+
+## If you implement it
+
+- `ExtractionConfig` in `src/engines.ts` is the right place for the per-engine
+ selectors; it is already plain JSON crossing into the page.
+- `PageProbe` in `src/extract.ts` gains an optional field; `SearchResult` should
+ **not** — an AI answer is not a hit and must not be numbered among them.
+- The settle-poll belongs in the page script, not in round trips from
+ `search.ts`. One `Runtime.evaluate` that polls locally until the text stops
+ growing is one round trip; polling from the host is one per sample.
+- `formatResults` in `src/format.ts` puts it above the numbered list. Mind the
+ budget note at the top of that file — an overview plus citations is another
+ 1-2KB on top of the 15KB worst case.
+- Tests: `test/search-url.test.ts` is where a "Google drops `udm` when an AI
+ answer is wanted" assertion goes, next to the existing parameter-trap tests.
+ The extraction itself cannot be tested without a browser; the existing suite
+ does not try, and this should not either.
+- Re-verify every selector before trusting it. Google's classes churned between
+ two loads minutes apart during this very investigation, and these blocks are
+ changing faster than the rest of the SERP.