Skip to content
Tech News
← Back to articles

Three sites made 215,128 “best software” pages for AI. Perplexity cites them

read original more articles
Why This Matters

This investigation reveals that a significant portion of AI grounding relies on obscure, low-traffic websites, many of which are newly created and not well-established. This raises concerns about the reliability and authenticity of AI-generated recommendations, impacting both consumers and the tech industry by highlighting potential vulnerabilities in AI sourcing methods.

Key Takeaways

Across 380 software categories, 59.8% of the sources behind grounded AI recommendations sit outside the 100,000 most-visited websites, and several of the most-cited are sites built to be read by models rather than by people.

We asked two web-grounded models for the best products in 380 software categories and kept every URL they retrieved. Of the 7,534 citations that came back, 59.8% point at domains ranked worse than #100,000 in the Tranco top-1M list and 23.4% at domains that are not in the top million at all. Two of the sites doing the grounding have given their homepage the HTML title “Facts & Grounding Page” — grounding being the retrieval step these models perform — and they and a third site under apparently common control have published 215,128 machine-generated best <category> pages between them; none of the three domains existed before December 2023.

What we ran

On 2 September 2026 we put 380 buyer-intent categories — from “CRM software” to “museum collection management software” — to perplexity/sonar and perplexity/sonar-pro through OpenRouter, one prompt per category per model, 760 calls in all. Each call asked for a ranked top five as JSON, with each product’s official homepage domain. All 760 returned a parseable answer, and both models report the URLs they retrieved, which is why they were chosen. The categories were written before any results were seen and never revised.

That produced 3,800 recommendation slots naming 1,807 distinct products, and 7,534 citations spanning 2,055 distinct domains. We then looked up every cited domain in the Tranco daily list for 2026-09-01 and in the Wayback Machine, and fetched every one of the 1,502 vendor homepages the models supplied to see whether it still exists.

Google was left out. Grounding a Gemini model on OpenRouter means routing it through OpenRouter’s own web-search plugin, so the citations would describe that plugin rather than Google’s retrieval. Only Perplexity was measured, and nothing here should be read as a claim about any other engine.

Where the citations land

Citations Unranked (outside Tranco 1M) Ranked worse than #100k perplexity/sonar 3,767 23.4% 59.8% perplexity/sonar-pro 3,767 23.5% 59.9% Pooled 7,534 23.4% 59.8%

The median Tranco rank of the 5,768 citations that point at a ranked domain is 71,611. Concentration at the top is unremarkable — the ten most-cited domains take 17.3% of citations — so the story is not that a cartel of famous sites supplies the answers. It is what fills the other four-fifths: 751 of the 2,055 cited domains, 36.5% of them, do not appear in the top million.

Those domains are also newer. The median first Wayback capture is 2020 for the unranked cited domains against 2011 for the ranked ones, and 16.6% of the archived unranked domains were first captured in 2025 or later, against 1.6% of the archived ranked ones.

... continue reading