# Turn Your Blog Archive Into a Knowledge Base AI Can Use

> Source: <https://www.digitalapplied.com/blog/blog-archive-ai-knowledge-base-method>
> Published: 2026-10-01 00:00:00+00:00

Most companies that have blogged for a few years own something they never use: a record of what they advised, when, and on what evidence. AI writing tools do not see it. They see the brief, a search result or two, and whatever the model remembers. The fix is not to feed the model every post. It is to extract what the posts claim into records that carry a source and a date, then put those records where a tool can retrieve them. This is the method.

**Editorial note:** A general method, written October 3 as an October 1, 2026 dispatch. It uses no corpus, count or result from any client engagement. The record schema is illustrative.

1. 01Inventory from the outsideThe CMS list is what you meant to publish. The sitemap plus a crawl is what readers and crawlers can reach. Start from the second.
2. 02Records, not postsA knowledge base is a set of claims with a URL and a date each, not a folder of articles. Extraction is the work.
3. 03Staleness is a fieldEvery record carries the date the claim was true and a confidence tag. Retrieval without those two fields repeats old advice with new confidence.
4. 04Small enough to read wholeAnthropic’s own retrieval guidance says a base under about 200,000 tokens can go into the prompt directly. Many archives, once extracted, fit.

## 01 — Step oneInventory from the sitemap and a crawl

The CMS export is the wrong starting list. It contains drafts, posts that redirect, pages that were unpublished but still indexed, and nothing about the URLs a migration left behind. Pull the XML sitemap, then run a crawl from the home page, and keep the union. Every URL that appears in one list and not the other is a finding before the work has begun: an orphan the crawl could not reach, or a live page the sitemap forgot.

For each URL, record the publish and modified dates from the page itself rather than from the database. The schema.org Article vocabulary defines datePublished as the date of first publication and dateModified as the date the work was most recently changed, and most sites emit both in structured data. They are the two fields that make staleness computable later. If a site’s dates are missing or wrong, that is the first repair, because nothing downstream can be trusted without them.

## 02 — Step twoExtract claims, advice and examples

Read each post for three kinds of sentence and ignore the rest. A claim is a statement about the world that could be checked: a figure, a rule a platform enforces, a date. Advice is a recommendation the post makes, with its condition if it states one. An example is a worked case, a template or a before-and-after that the advice points at. Introductions, transitions and conclusions are not extracted. They are the part of a post that the model can write again; the three kinds above are the part it cannot.

This is a reading job that a model can do at scale and a person must sample. Give the model one post and the schema below, ask for records only, and have an editor check one post in ten against the original. The two failure modes to look for are invented precision, where a vague sentence becomes a number, and dropped conditions, where “if the site is under a thousand pages” falls off the front of the advice. Our note on [preserving facts through AI rewrites](https://www.digitalapplied.com/blog/ai-content-update-fact-preservation) describes the same two failures from the other direction.

## 03 — The artefactOne record schema

Keep the schema small enough that a model fills it reliably and a spreadsheet can hold it. These are the fields we would start with for any archive; a specialist site adds its own.

| An illustrative record schema. Field names are generic; the example values are invented to show the shape. |  |  | 
|---|---|---|
| Field | What goes in it | Example | 
|---|---|---|
| type | claim, advice or example | advice | 
| text | The statement in one sentence, in the post’s own terms, with its condition kept. | Add a visible last-reviewed date to any page that cites a platform rule. | 
| source_url | The post it came from, and the section anchor where one exists. | /blog/example-post#review-dates | 
| evidence | What the post cited for it: a primary, a vendor figure, an observation, or nothing. | vendor documentation, linked | 
| as_of | The date the claim was true, from the post’s published or modified date. | 2025-03-14 | 
| confidence | high, medium or low, set by the evidence field and the editor’s sample. | medium | 
| staleness | stable, dated or expired, judged on whether the subject changes yearly, quarterly or never. | dated | 
| topics | Two or three terms from a taxonomy you control, not free tags. | content-review, platform-policy | 

Two fields do most of the work. The evidence field is what lets a later reader tell a measured figure from a remembered one. The as_of field is what lets a tool say “this was true in March 2025” rather than asserting it as current. A base without them is a quote bank; with them it is a record.

## 04 — Step threeTag, group and deduplicate

An archive repeats itself. The same advice appears in a 2023 guide, a 2024 checklist and a 2025 refresh, worded three ways and sometimes contradicting itself in a detail. Group records by topic, then within a topic sort by as_of and read the duplicates together. Keep the newest version that has evidence; mark the older ones as superseded rather than deleting them, because the change itself is information. A claim that moved from “always” to “usually” between two posts is worth a sentence in the next one.

The taxonomy should be small, with a few dozen terms at most, and written before tagging starts. Free tags drift into synonyms within a week. If the archive already has a category and tag structure, start from it and prune; our [keep, rewrite or retire](https://www.digitalapplied.com/blog/old-posts-keep-rewrite-retire-decision-data) method is the companion pass for the posts themselves, and the two share a spreadsheet comfortably. If the immediate job is a bulk rewrite rather than a standing base, our note on [indexing the facts before an agent rewrites an archive](https://www.digitalapplied.com/blog/source-index-before-agent-rewrites-content) is the narrower, faster version of the same extraction.

Anthropic’s engineering note on contextual retrieval makes a practical point before it gets to retrieval: if a knowledge base is smaller than about 200,000 tokens, roughly 500 pages, it can simply be included in the prompt. An archive of a few hundred posts often extracts to fewer tokens than that, because the records are a fraction of the prose. Measure before building a retrieval layer. Many teams do not need one.

## 05 — PayoffFour things the base is for

The base earns its cost in four jobs that every content team already does by memory and search.

##### Briefs

Pull every record on the topic. The brief now says what the company has already claimed, with dates, so the new post extends the position rather than restating or contradicting it.

##### Refreshes

Filter records by staleness and as_of. The list of expired claims is the edit list, and the evidence field says which ones need a new source rather than a new sentence.

##### Internal links

A record that supports a sentence in the draft is a link candidate to its source_url. The anchor text is the record’s text; the link lands on the section, not the home page.

##### Voice and position

The advice records, read together, are the company’s stated positions. Hand them to the model with the brand voice guide and the draft stops inventing opinions the company never held.

The fourth use pairs with a [brand voice guide](https://www.digitalapplied.com/blog/extract-brand-voice-guide-ai-content-2026): the guide says how the company sounds, the base says what it thinks. A model given only the first writes fluent pages with no spine. Given both, it writes pages the editor recognises.

## 06 — Serving itHow AI tools read it

Once the records exist, there are three ways to put them in front of a tool, and the right one depends on size and on who is asking.

The third option deserves a caveat the proposal itself makes. The llms.txt format is for inference, when an agent needs information about a topic while helping a user. It is not a search-ranking input, and the proposal does not claim to be one. Our [review of llms.txt in practice](https://www.digitalapplied.com/blog/llms-txt-in-practice-adoption-evidence-2026) has the adoption evidence. Publish one because it is a cheap, readable index of what the company stands behind, not because it will move a ranking.

Where a team wants this built as a repeatable system rather than a one-off spreadsheet, our [content engine](https://www.digitalapplied.com/services/content-engine) work sets up the extraction, the editor sample and the refresh filter so the base stays current as the archive grows.

### Extract ten posts by hand before you automate anything

Take the ten most-linked posts, fill the schema for each, and read the records together. That afternoon tells you whether the taxonomy holds, how often the evidence field is empty, and whether the archive is worth the rest of the work. Usually it is.
