215,128 fake "best software" pages were built to be read by AI — and Perplexity is citing them
A new study found three sites, none created before December 2023, that mass-produced 215,128 machine-generated ranking pages — some literally titled "Facts & Grounding Page" for the model to read. Nearly 60% of the citations Perplexity returned pointed to domains ranked below #100,000.

Ask an AI search engine for "the best project-management software" and it will hand you a confident, sourced answer. A new study suggests you should look hard at those sources.
The finding
Trellner Research queried two of Perplexity's models — sonar and sonar-pro — across 380 software categories, asking each for a ranked top five with official homepages. That produced 7,534 citations across 2,055 domains. Then it checked where those citations actually pointed, using the Tranco list, a standard ranking of the top million domains by real-world popularity.
The result is striking: 59.8% of the citations pointed to domains ranked worse than #100,000, and 23.4% pointed to domains not in the top million at all. The median ranked citation pointed to a domain at Tranco rank 71,611. In other words, when Perplexity named the "best" software, most of the sources it leaned on were sites almost nobody visits.
Where the citations were coming from
The study traced a large share of them to a specific pattern. Three sites, under apparently common control, had between them published 215,128 machine-generated "best software" pages — roughly one page per niche, at industrial scale. None of the three domains existed before late 2023; all were registered between December 2023 and May 2024 — that is, after AI search made "being cited by the model" a thing worth gaming.
And these sites were not pretending to be for humans. Two of them gave their homepage the HTML title "Facts & Grounding Page," describing themselves as a "machine-readable record" addressed to software rather than to readers. "Grounding" is the retrieval step a web-connected model runs before it answers — the moment it decides which sources to trust. These pages were built, explicitly, to be there when the model looked.
What it means — and what it doesn't
This is search-engine optimisation for the AI age, and it works differently from the old kind. Classic SEO games a ranking a human eventually sees; this games the grounding step, the invisible moment where a model picks its sources. A content farm that would never reach Google's first page can still end up cited, by name, inside a confident AI answer — because the model is optimising for "a page that matches this query exists," not "a source a person would trust."
Two caveats, both important. First, the study is careful about its own scope: it measured Perplexity only, and states plainly that "nothing here should be read as a claim about any other engine." This is not a verdict on AI search in general. Second, being cited is not the same as being wrong — a low-ranked page can still list real software. But 215,128 pages built to be read by machines, cited at this rate, is not a quality signal; it is a supply of manufactured sources meeting a demand the models created.
The takeaway
The practical version is simple. When an AI search engine gives you a sourced answer, the citation is doing a lot of quiet work to make it feel authoritative — and, at least in this study of Perplexity, a majority of those citations pointed at domains you have never heard of, for a reason. Click through before you trust. The model checked that a page existed; it did not check that the page deserved to.
Ask Relay — he reads every question himself and replies personally by email.
