{"slug": "clean-article-text-a-tl-dr-from-any-news-url-in-one-api-call", "title": "Clean article text + a TL;DR from any news URL in one API call", "summary": "A developer built an API that extracts clean article text from news URLs and optionally generates AI summaries in a single call, returning structured metadata such as authors, publish dates, word counts, token estimates, and Markdown body text. The service, demonstrated on an October 2026 Guardian article about Oura's postponed IPO, returned extraction results in 0.48 seconds and summaries in 1.74 to 3.92 seconds, and it rejects non-article pages like section fronts before invoking any model.", "body_md": "If you have ever tried to feed news articles to an LLM, you know the boring part is not the summary. It is everything before it: cookie banners, \"related stories\", share buttons, newsletter boxes, author bios, and pages that are not articles at all (a home page, a section front, a login wall).\n\nI built a small API that does the boring part and then, only if you ask, the summary. Below are real calls I made on 1 October 2026, with the actual responses.\n\n```\ncurl --request POST \\\n  --url 'https://article-extractor-ai-summarizer-api.p.rapidapi.com/api/v1/article' \\\n  --header 'X-RapidAPI-Key: YOUR_RAPIDAPI_KEY' \\\n  --header 'X-RapidAPI-Host: article-extractor-ai-summarizer-api.p.rapidapi.com' \\\n  --header 'Content-Type: application/json' \\\n  --data '{\"url\": \"https://www.theguardian.com/technology/2026/sep/29/oura-ring-public-offering\", \"maxChars\": 400}'\n```\n\nThe response came back in **0.48 s** (plain HTTP, no browser needed). Trimmed to the fields that matter:\n\n```\n{\n  \"success\": true,\n  \"isArticle\": true,\n  \"pageType\": \"article\",\n  \"articleType\": \"NewsArticle\",\n  \"title\": \"Smart ring maker Oura puts off initial public offering due to market ‘uncertainty’\",\n  \"authors\": [\"Guardian staff reporter\", \"Associated Press\"],\n  \"publishedAt\": \"2026-09-29T17:57:01.000Z\",\n  \"modifiedAt\": \"2026-09-29T18:20:35.000Z\",\n  \"publisher\": \"The Guardian\",\n  \"language\": \"en\",\n  \"tags\": [\"Technology\", \"IPOs\", \"AI (artificial intelligence)\", \"Business\", \"US news\"],\n  \"paywalled\": false,\n  \"wordCount\": 296,\n  \"readingTimeMinutes\": 1,\n  \"tokensEstimate\": 445,\n  \"format\": \"markdown\",\n  \"markdown\": \"…(body as Markdown, cut at 400 characters because of maxChars)…\",\n  \"truncated\": true,\n  \"summaryGenerated\": false\n}\n```\n\nA few details I cared about:\n\n`authors` is an `wordCount` and `tokensEstimate` describe the `maxChars` cuts the body. Handy for deciding whether a page fits your context window before you download all of it.`format` can be `markdown`, `text` or `both`.\nThe summary lives on a separate endpoint, `/api/v1/article/summary`, so extraction-only calls never touch the model (and never use your summary quota).\n\n```\ncurl --request POST \\\n  --url 'https://article-extractor-ai-summarizer-api.p.rapidapi.com/api/v1/article/summary' \\\n  --header 'X-RapidAPI-Key: YOUR_RAPIDAPI_KEY' \\\n  --header 'X-RapidAPI-Host: article-extractor-ai-summarizer-api.p.rapidapi.com' \\\n  --header 'Content-Type: application/json' \\\n  --data '{\"url\": \"https://www.theguardian.com/technology/2026/sep/29/oura-ring-public-offering\", \"summary\": \"tldr\"}'\n```\n\n**1.74 s**, and the summary block looked like this:\n\n```\n\"summary\": {\n  \"style\": \"tldr\",\n  \"language\": \"same as source\",\n  \"maxWords\": 30,\n  \"text\": \"Oura postpones its initial public offering due to market uncertainty despite strong demand and revenue growth.\",\n  \"inputChars\": 1866,\n  \"inputTruncated\": false,\n  \"tokens\": { \"input\": 575, \"output\": 20 }\n}\n```\n\nYou get the token counts back, so you can see what the model actually read.\n\nSame article, `\"summary\": \"bullets\", \"maxWords\": 80, \"language\": \"Spanish\"` — **3.92 s**, 7 bullets. The first three:\n\n```\n- Oura Inc está retrasando su oferta pública inicial debido a la incertidumbre del mercado.\n- La empresa había planeado vender 50m acciones en la oferta pública inicial a un precio entre $40 y $44.\n- Casi tres cuartas partes de las acciones serían vendidas por accionistas actuales.\n```\n\nThere is also `\"summary\": \"paragraph\"` and a `focus` option (for example `\"focus\": \"numbers and dates\"`). If you already have the text, send `text` instead of `url`.\n\nThis is the part I wanted most. Summarising a home page gives you a confident paragraph about nothing. So the API checks the page type first and refuses before any AI runs:\n\n```\n--data '{\"url\": \"https://www.bbc.com/news\"}'\n{\n  \"success\": false,\n  \"error\": {\n    \"code\": \"not_an_article\",\n    \"message\": \"This page is not an article: this page lists articles (a section, category or tag page) rather than being one.\",\n    \"details\": { \"pageType\": \"listing\", \"finalUrl\": \"https://www.bbc.com/news\" }\n  }\n}\n```\n\nThat came back in 0.67 s with HTTP 422. If you do want the page anyway, pass `\"allowNonArticle\": true` and you will get it with `isArticle: false`.\n\nSection fronts were the hard case. My first test run let the Guardian's technology section through as an \"article\" (no author, no date). The check now also looks at the page's structured data (a `CollectionPage` is not an article, even when `og:type` says so), at how much of the page is headline cards with short teasers, and at URL shapes like `/section/`, `/category/` and `/tag/`. In a re-run, 15 section, category and tag pages (Guardian, BBC, Reuters, Spiegel, TechCrunch, The Verge, Medium, Cloudflare's blog and others) all came back `422` with `pageType: listing`, and 13 real articles (news in English, German and Japanese, Wikipedia in English and Japanese, a Paul Graham essay, Medium, Substack, Cloudflare and personal blog posts) were all still accepted. It is still a heuristic, so if you feed it mixed URLs, `pageType` and `publishedAt` are worth a glance.\n\nAnd some sites simply refuse automated requests. One news site I tried answered the first extraction and then returned HTTP 405 to the next calls; the API reported that as `422 target_blocked` instead of making something up.\n\n``` python\nimport requests\n\nHOST = \"article-extractor-ai-summarizer-api.p.rapidapi.com\"\nHEADERS = {\"X-RapidAPI-Key\": \"YOUR_RAPIDAPI_KEY\", \"X-RapidAPI-Host\": HOST}\n\ndef tldr(url: str) -> str | None:\n    r = requests.post(f\"https://{HOST}/api/v1/article/summary\",\n                      json={\"url\": url, \"summary\": \"tldr\"}, headers=HEADERS, timeout=60)\n    body = r.json()\n    if not body.get(\"success\"):\n        print(url, \"->\", body[\"error\"][\"code\"])   # not_an_article, target_blocked, timeout...\n        return None\n    return body[\"summary\"][\"text\"]\n\nfor link in [\"https://www.theguardian.com/technology/2026/sep/29/oura-ring-public-offering\",\n             \"https://www.bbc.com/news\"]:\n    print(link, \"=>\", tldr(link))\n```\n\n`inputTruncated` tells you.\nOn RapidAPI there is a free plan (200 requests and 30 summaries a month, hard limit; RapidAPI may ask for a card even for free plans). The plans as of 1 October 2026:\n\n| Plan | Per month | Extraction requests | Summaries | \n|---|---|---|---|\n| Basic | $0 | 200 | 30 | \n| Pro | $12.99 | 10,000 | 1,500 | \n| Ultra | $39.99 | 50,000 | 6,000 | \n| Mega | $99.99 | 150,000 | 15,000 | \n\nEvery call counts as a request; only `/article/summary` calls count as summaries. RapidAPI also has its own bandwidth fee above 10 GB a month, which is theirs, not mine.\n\n**Try it:** [https://rapidapi.com/tidytools/api/article-extractor-ai-summarizer-api](https://rapidapi.com/tidytools/api/article-extractor-ai-summarizer-api)\n\nIf you would rather run it in bulk without code, the same summariser is an Apify Actor, [Article Summarizer & Web Page Summarizer - AI Text Summarizer](https://apify.com/tidytools/web-page-summarizer) ($6 per 1,000 pages).\n\n*Disclosure: I built this API and the Apify Actor, and I earn money when people use them. All outputs above are real responses from 1 October 2026; the article quoted is Associated Press/Guardian reporting, used here only as a test input.*", "url": "https://wpnews.pro/news/clean-article-text-a-tl-dr-from-any-news-url-in-one-api-call", "canonical_source": "https://dev.to/tidytools/clean-article-text-a-tldr-from-any-news-url-in-one-api-call-1074", "published_at": "2026-10-08 02:07:11+00:00", "updated_at": "2026-10-08 02:17:29.968227+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "developer-tools"], "entities": ["RapidAPI", "The Guardian", "Oura"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/clean-article-text-a-tl-dr-from-any-news-url-in-one-api-call", "markdown": "https://wpnews.pro/news/clean-article-text-a-tl-dr-from-any-news-url-in-one-api-call.md", "text": "https://wpnews.pro/news/clean-article-text-a-tl-dr-from-any-news-url-in-one-api-call.txt", "jsonld": "https://wpnews.pro/news/clean-article-text-a-tl-dr-from-any-news-url-in-one-api-call.jsonld"}}