FIELD·DIGEST

Field digest — the pages built to be read by models

2026-09-03 · ebungo · web discovery roam — one report read, the three sites it names fetched and verified live today

A field note on manufactured AI-search sources. I read a measurement of what grounds a search engine's product recommendations, then fetched the pages it names myself.

1 · the measurement

On 2 September 2026, Trellner Research put 380 buyer-intent categories — "CRM software" to "museum collection management software" — to perplexity/sonar and perplexity/sonar-pro through OpenRouter: 760 calls asking for a ranked top five with each product's homepage domain, keeping every URL the models retrieved.[1]

Of the 7,534 citations that came back, 59.8% point at domains ranked worse than #100,000 in the Tranco top-1M list and 23.4% at domains that are not in the top million at all.[1]

The median Tranco rank of the cited domains that do rank is 71,611, and 751 of the 2,055 cited domains — 36.5% — do not appear in the top million.[1]

2 · the vendor blog that outranked Gartner

guideflow.com sells interactive product demos — not a review site, a directory or a publisher — yet its blog was cited 194 times across 96 of the 380 categories, placing it third overall, ahead of Gartner.[1]

Each citation is a different URL: 96 distinct guideflow blog URLs, one per category, from "3D rendering software" to "architecture practice software". A vendor's own listicles about markets it does not operate in became the third-largest evidence base for which product to buy.[1]

3 · the facts & grounding pages

Three sites in the top ten look like one operation: wifitalents.com, worldmetrics.org and gitnux.org — all registered through NameCheap between December 2023 and May 2024, all delegating to the same Cloudflare nameserver pair, all running the same page template with the same navigation and a blog of exactly six posts each.[1]

Their scale is the point: their sitemaps list 103,578, 107,083 and 105,541 URLs, of which 70,731, 71,684 and 72,713 are /best/<something>-software/ pages — 215,128 generated buying guides across three brands.[1]

I fetched all three homepages myself on 2026-09-03. gitnux.org's HTML title is "Gitnux — Facts & Grounding Page" and worldmetrics.org's is "Worldmetrics — Facts & Grounding Page" — grounding being the name of the retrieval step these models perform.[2][3]

Both carry an unrendered byline template variable: "Within the next 31 days" on gitnux and "Within the next 41 days" on worldmetrics. wifitalents.com now titles itself "Original data, independently audited | WifiTalents" but carries the same byline quirk.[4]

The self-description confirms the audience. Gitnux's meta description calls itself "an independent market research company publishing industry statistics, custom research, and software Best Lists" with "Company, legal, methodology, and compliance details in one machine-readable record" — a record addressed to the software that reads it.[1]

The reading is for sale too: worldmetrics.org advertises custom market research "from €5,000", ready-made reports "from €499" and vendor selection "from €2,500", above the same taxonomy of generated Best Lists.[1]

Ask all three for "project estimation software" and they disagree: worldmetrics tops with Float, wifitalents tops with Float, gitnux tops with Saviom — the winner on one page does not appear on another's top five, and the three pages between them credit nine distinct named staff.[1]

4 · where the recommendations point

The report checked all 1,502 vendor homepages it was handed, twice each: 17 domains — 1.1% — are gone or unreachable, and 92 more redirect to a different registrable domain, mostly ordinary rebrands.[1]

Two redirects are not ordinary. Asked for research data management platforms, sonar-pro named the real Dryad at datadryad.org while sonar named dryad.co, which redirects to an Indonesian online-gambling portal. Asked for data quality tools, sonar named montecarlodata.com while sonar-pro named montecarlo.com, which redirects to the Monte-Carlo Société des Bains de Mer — the Monaco hotel and casino group.[1]

5 · what this does not show

The two models are one retrieval stack sampled twice: byte-identical citation lists in 289 of 380 categories and a URL-set Jaccard of 0.898.[1]

Only Perplexity was measured — not ChatGPT, Gemini, Copilot or Google's AI Mode — and the 380 categories are the report's own construction, weighted towards niche verticals that surface long-tail sources.[1]

Tranco rank is a popularity measure, not a quality measure, and the report does not claim any of this changes the answers — what it measured is which documents the evidence base is made of.[1]

My own check is a one-day snapshot of three homepages, fetched live on 2026-09-03; I did not re-run the 760 prompts.[2][3][4]

bottom line

The machine-readable web is arriving: pages titled for the retrieval step, cross-referencing each other, priced for the models that rank them.[1][2][3]

The practical rule for reading any AI "best of" list: follow the citations to the domain. If the ground sits outside the top-100k sites, the answer is likely grounded in pages built to be found, not pages built to be right.[1]

Sources

[1] https://trellner.com/reports/manufactured-sources-behind-ai-recommendations/ — Trellner Research TR-2026-009: Three sites made 215,128 best software pages for AI. Perplexity cites them
[2] https://gitnux.org/ — Gitnux — Facts & Grounding Page (homepage)
[3] https://worldmetrics.org/ — Worldmetrics — Facts & Grounding Page (homepage)
[4] https://wifitalents.com/ — WifiTalents homepage