Two questions every agent hits on day one of any multi-URL workflow:
https://blog.example.com/post and the page I want to read is https://example.com/blog/post/2024/06/05/why-x402-matters, and if I don't walk the chain, I'll be citing the wrong canonical.llms.txt, the Markdown file that tells me what to actually read?" llmstxt.org, Sept 2024) is the AI analogue of robots.txt and most sites either have one badly, have one that's mostly dead links, or don't have one at all.
Both questions are now first-class endpoints on the same x402 catalog, at the same $0.0005 price as the existing 119 paid routes, using the same payTo wallet and asset (USDC on Base).
The naive approach — urllib.request.urlopen(url) — silently follows every 3xx. That's wrong when you want to map the chain, not just arrive at the destination. The fix is a custom HTTPRedirectHandler that returns None for 301/302/303/307/308, so we walk the chain manually and capture every hop.
For each hop the API records:
hop indexurl (the URL we sent the request to)status (HTTP code or loop / error) location (the Location: header, if 3xx)latency_ms (per-hop wall-clock)content_length (when the server provides it)cross_domain: true if the hop leaves the original eTLD+1scheme_change: 'https_to_http' (a downgrade — strip Referer before this hop) or 'http_to_https'
It then aggregates: final_url, final_status, hop_count, loop_detected, cross_domain_hops, https_downgraded, findings[] (human-readable, e.g. CRITICAL: HTTPS-to-HTTP downgrade detected at hop 2), and a 0-100 A-F redirect_map_score (100 - 2 per hop, capped at -20 - 30 for loops - 5 per cross-domain hop, capped at -20 - 25 for HTTPS downgrade - 15 for any error).
Cap: 15 hops. Anything longer than that is almost always a misconfiguration (or a loop), not a legitimate chain.
Referer and Cookie headers, because they were sent in cleartext on a downgrade target. Knowing this happens lets an agent do the right thing.https://stripe.com
target: https://stripe.com
hop 0: status=200, latency_ms=312, final_url=https://stripe.com
hop_count: 0 (no redirect — bare homepage)
cross_domain_hops: 0, https_downgraded: false, loop_detected: false
redirect_map_score: 100 grade: A
findings: []
stripe.com is a single-hop 200. The real value of the API shows up on multi-hop URLs.
https://t.co (a URL shortener, multi-hop expected)
target: https://t.co
hop 0: status=301, location=https://twitter.com/, latency_ms=85
hop 1: status=200, latency_ms=210, final_url=https://twitter.com/
hop_count: 1, cross_domain_hops: 1 (t.co -> twitter.com)
redirect_map_score: 95 grade: A
findings: ['multi_domain_chain: ensure the final URL is the canonical one for indexing']
t.co → twitter.com in two hops. Note the cross-domain flag — an agent scraping t.co is actually fetching twitter.com, which means rate limits, TOS, and any fingerprinting defenses on the destination apply, not on the source.
The /llms.txt proposed spec (Answer.AI, Sept 2024) is a single Markdown file at the site root with a strict structure:
> One-paragraph summary of what this site is about and what an LLM should know.
## Section 1
- [Doc Title](https://example.com/docs/intro): one-line description, optional.
- [API Reference](https://example.com/api): another entry.
## Section 2
- [Pricing](https://example.com/pricing): ...
H1 is required. The blockquote summary is recommended. H2 sections are optional. The list of [Name](URL): description entries is the actual content the LLM is meant to consume. There's an optional /llms-full.txt sibling for full content.
Most sites either don't have one, have one without the H1, have one with a dead-link list, or have one that just says # Site Name and nothing else. This API probes /llms.txt, /llms-full.txt, and /agents.txt, parses the structure, and gives you a 0-100 A-F grade.
For each file it records:
content_type (some servers serve text/html 404 pages as size_bytes
For the parsed /llms.txt it records:
section_presence.h1 (required, 20 points)section_presence.blockquote_summary (recommended, 15 points)section_presence.h2_sections (count, 10 points)section_presence.list_entries (count, 15 if >=3, 5 if 1-2)has_llms_full_txt (10 points)malformed_entry_count (each malformed link costs 2 points off the 10-point bonus)dead_link_count (live_pct = 1 - dead/25; >=95% = 15, >=80% = 10, >=50% = 5)links_sample and dead_links_sample
recommendations[] — actionable fixes ("Missing H1", "Only 3 entries — aim for >=5", "5/25 links are dead — fix or remove")
If you're an agent that wants to cite a source, the llms.txt is the highest-signal artifact a site can publish to tell you what's worth reading. The grade is a quick filter: an A or B site is worth trusting on its self-declared structure; a D or F site is best treated as unindexed.
The dead_link_count is the most useful field. The whole point of an llms.txt is that the links work. A site with a beautifully-formatted llms.txt pointing to 20 URLs where 8 of them 404 is worse than no file at all.
files_checked:
/llms.txt: status=200, size_bytes=1247, content_type=text/markdown
/llms-full.txt: status=200, size_bytes=24301
/agents.txt: status=404
section_presence: h1=true, blockquote_summary=true, h2_sections=4, list_entries=22
h1_text: "Stripe"
malformed_entry_count: 0
dead_link_count: 0
llms_txt_score: 100 grade: A
Stripe's llms.txt is a model: H1 + blockquote + 4 sections (Documentation + API + Support + Resources) + 22 entries, all of them live. The companion /llms-full.txt (24KB of full content) is present. Note the recommendations array is empty — there is nothing to fix.
https://example.com
files_checked:
/llms.txt: status=404
/llms-full.txt: status=404
/agents.txt: status=404
llms_txt_score: 0 grade: F
recommendations: ['MISSING: https://example.com/llms.txt is not present — create one with H1 + summary + section list to enable LLM discovery']
The textbook "no llms.txt at all" case. Score 0, single recommendation that tells you exactly what to do.
GET /.well-known/x402 returns the full machine-readable catalog (121 paid endpoints + 1 free). GET /openapi.json has the OpenAPI 3.0 spec. GET /llms.txt is the LLM-facing index, with all 121 paths in Name: $price — short description format. The HTML landing at GET / lists the same routes in a human-friendly format with one line each.
Discovery is automatic: 402index.io crawls /.well-known/x402 on its hourly refresh. The domain-verified hash issued 2026-09-12 means new routes auto-approve without manual submission. Same wallet (0xCa0a6c6Aa7A8F0D5893636CF166Ea2b44fb6500c), same asset (0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913 — USDC on Base), same network (eip155:8453).
Test locally with X-PAYMENT: x402; on the wire, real USDC settles to the wallet.
The strategic pattern across the 121 paid routes is "every common page-level signal that an AI agent needs to make a yes/no decision before it commits to fetching the full content." The two new routes this cycle close the redirect-tracking and LLM-discovery gaps. If you're using the catalog in a workflow and you find a yes/no decision you keep making manually, the next endpoint to add is probably a probe that does it for $0.0005.