Research record · June 11, 2026

What we checked, what we counted, and where the data stops.

This is the methodology and limitation record for one bounded scan of 50 affiliate publisher sites. It documents observable link behaviour—not market-wide prevalence, customer outcomes, or revenue impact.

One observed sampleChecked-link denominator onlyNo revenue estimates

6,550

outbound links checked

11,676 found; 5,126 outside the time budget

5.8%

visibly broken

380 non-blocked 4xx, 5xx, or timeout observations

9.1%

attribution failures

597 stripped-parameter or homepage-redirect observations

34/50

sites with substantive coverage

The other 16 returned fewer than three public pages and limit site-level conclusions

All rates above use the 6,550 links actually checked within budget. The 5,126 skipped links are not silently treated as healthy.

Collection method

A repeatable crawl, with the shortfalls left visible

The process was deliberately bounded. Partial coverage is part of the result, not a reason to convert missing observations into zeros.

  1. 1. Build a fixed research panel

    The panel contained 50 established affiliate publisher sites across several niches. It was curated, not randomly sampled, so it does not represent every publisher or the whole market.

  2. 2. Discover a bounded set of pages

    The crawler checked sitemap entries first, then used a same-site crawl when needed. It respected robots.txt and capped each site at 20 public pages.

  3. 3. Enforce the same time limits

    Each site received a 100-second soft budget and a 300-second safety stop. 23 sites returned partial results; skipped links never entered a rate denominator.

  4. 4. Follow each checked link

    The checker started with HEAD, retried suspicious responses with GET, allowed up to 5 redirect hops, and used a 10-second request timeout.

  5. 5. Keep blocks separate from breaks

    308 responses still returned 403, 405, or 429 after the fallback. They were labelled unverifiable by the automated check and excluded from the visibly-broken count.

Counting rules

Six labels, three reporting groups

A status-code failure and a tracking failure answer different questions. They are shown separately and never converted into a financial claim.

ObservationCountChecked linksReporting group
Non-blocked 4xx1712.6%Visible break
5xx server response821.3%Visible break
Timeout1271.9%Visible break
Tracking parameter stripped5698.7%Attribution failure
Deep link redirected to homepage280.4%Attribution failure
Bot-blocked or rate-limited3084.7%Reported separately

The central observed pattern

569 of the 597 attribution failures were tracking parameters that disappeared by the final URL. Those links can still return a successful page response, which is why status-only checking misses the category.

Read before citing

What this study cannot establish

The honest use of a bounded study is narrow. These findings describe the links observed on this panel, on this date, under these crawl limits.

Not a representative sample

The 50 sites were curated. The percentages are not a universal rate for publishers, niches, or the web.

Not complete site inventories

23 sites exhausted the soft budget, and only 34 returned at least three pages. Results are bounded observations, not full audits.

Not proof of revenue impact

The crawler did not have commission rates, traffic, conversion data, account access, or customer outcomes. No dollar inference is supported.

Not permanently current

Links, redirects, servers, and bot rules change. The record describes June 11, 2026 and should not be presented as a live measurement.

Apply the same evidence standard

Check a site without pretending the crawler knows its revenue.

The free leak snapshot reviews a limited set of public pages and reports observable link behaviour. Fit for deeper work is reviewed separately.

Get a free leak snapshot