cd /news/ai-tools/openstax-llm-tools-for-openstax-llm-… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-144118] src=github.com β†— pub= topic=ai-tools verified=true sentiment=↑ positive

OpenStax-LLM: tools for OpenStax LLM access

OpenStax-LLM, a new open-source toolkit built on the openstax-md project, converts OpenStax college textbooks into citation-aware, formula-safe datasets for vector search and LLM fine-tuning. The repository ships a Python library and CLI, a Model Context Protocol server (openstax-llm-mcp), and an agent skill, and can compile and chunk any of the 90-plus OpenStax catalog volumes into JSONL with globally unique chunk_ids compatible with Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face datasets. The project targets generic chunkers that split math formulas, separate worked examples from solutions, and lose chapter hierarchy, enforcing a max_words ceiling and never placing a boundary inside unclosed $...$ or $$...$$ blocks or fenced code blocks.

read4 min views3 publishedOct 2, 2026
OpenStax-LLM: tools for OpenStax LLM access
Image: Michielbdejong (auto-discovered)

Pedagogical semantic chunking, RAG dataset preparation, and LLM fine-tuning pipelines from OpenStax textbooks.

Built on top of openstax-md, openstax-llm transforms OpenStax college textbooks into structured, citation-aware, formula-safe datasets for vector search (RAG) and model fine-tuning.

This repository ships three things that share one core:

Artifact What it is Entry point
openstax-llm Python library and CLI openstax-llm
openstax-llm-mcp Model Context Protocol server openstax-llm-mcp
skills/openstax-llm Agent skill for coding assistants /skill:openstax-llm

Generic chunkers (simple character or recursive token splitters) break down on technical academic textbooks:

  • They cut mathematical formulas in half ($x^2 + \dots$ split from\dots + y^2$ ).
  • They separate worked examples from their solutions.
  • They lose the chapter and section hierarchy that citations depend on.

openstax-llm provides:

  1. Pedagogical boundary awareness β€” worked examples (Example 1.1 ), problem sets, definitions, and summaries are kept whole.
  2. Formula integrity β€” a chunk boundary is never placed inside an unclosed$...$ or$$...$$ block, or inside a fenced code block.
  3. Enforced size ceiling β€”max_words is actually honoured; oversized paragraphs are split at sentence boundaries that lie outside math.
  4. On-demand compilation β€” any textbook in the OpenStax catalog (90 volumes and counting) is pulled and chunked without manual data management.
  5. Vector-store-ready output β€” JSONL with globally uniquechunk_id s, compatible with Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Facedatasets .
uv add openstax-llm

uv tool install openstax-llm

uvx openstax-llm-mcp

Or run the container:

docker build -t openstax-llm-mcp .

To track an unreleased commit instead, install from git:

uv add "git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm"
uvx --from "git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm-mcp" openstax-llm-mcp
openstax-llm search physics

openstax-llm info astronomy-2e

openstax-llm prepare astronomy-2e -o datasets/astronomy-2e.jsonl

openstax-llm validate datasets/astronomy-2e.jsonl

validate checks formula integrity, chunk_id uniqueness, and provenance, and exits non-zero on structural errors. Add --strict to also fail on warnings such as front matter that has no section number. Always run it before a dataset into an index: a split formula that reaches an embedding store is very hard to detect afterwards.

from openstax_llm import DocumentChunker, TextBookDataset

chunker = DocumentChunker(target_words=400, max_words=600, overlap_words=50)
dataset = TextBookDataset.from_textbook("calculus-volume-1", chunker=chunker)
dataset.to_jsonl("calculus.jsonl")

print(dataset.summary())

DocumentChunker also works on arbitrary markdown:

from openstax_llm import DocumentChunker

chunks = DocumentChunker().chunk_markdown(
    "## 1.2 Functions\n\nAn inline formula $f(x)=x^2$ stays intact.\n",
    book_slug="my-notes",
    section="1.2",
    section_title="Functions",
)

Exposes the same three operations to any MCP client over stdio or streamable HTTP.

pi mcp add openstax-llm -- uvx openstax-llm-mcp

docker run --rm -p 8765:8765 openstax-llm-mcp
Tool Purpose
search_catalog(query, limit) Resolve a subject to a canonical slug (offline)
inspect_textbook(target) Chunk totals plus a per-section index
prepare_textbook(target, out) Export JSONL, sandboxed to the server's output directory

Resources: textbook://<slug> for an overview and textbook://<slug>/<section> for every chunk in one section. See packages/openstax-llm-mcp/README.md for options, tool schemas, and failure modes.

npx skills add michaelnavazhylau/openstax-llm

Also on skills.sh β€” that directory is populated from anonymous install telemetry, so the listing appears only after the first npx skills add (the command above is always the canonical path).

The skill teaches an agent when and how to reach for these tools: resolving slugs instead of guessing titles, verifying exports before indexing, tuning chunk sizes, and the result into Chroma, Qdrant, Pinecone, or Hugging Face. It deliberately contains no chunking logic of its own.

Field Type Description
chunk_id string Unique within a book: <section>-c<index> , e.g.1.2-c003
text string Markdown with intact LaTeX math
book_slug string Canonical OpenStax slug
book_title string Human-readable title
chapter string Chapter number from the section hierarchy
section string Section number, e.g. 1.2
section_title string Module title
chunk_type string prose ,example ,exercise ,definition ,summary
word_count integer Whitespace-delimited word count
token_est integer Heuristic estimate, words * 1.3
metadata object Carries module_id ; free for downstream use

Front matter β€” prefaces, formula tables, chapter introductions β€” has no section number. Its chunk_id is prefixed with the module id and its provenance lives in metadata.module_id. The machine-readable schema is at skills/openstax-llm/assets/chunk.schema.json.

uv sync --all-groups --all-packages

uv run pytest -v
uv run ruff check . && uv run ruff format --check .
uv run mypy

OPENSTAX_LLM_NETWORK_TESTS=1 uv run pytest tests/test_integration.py -v

The repository is a uv workspace: the root pyproject.toml declares members and owns the shared tool configuration, while each package under packages/ is independently distributable. See AGENTS.md for the architecture and engineering invariants, and docs/PUBLISHING.md for the release process.

MIT

── more in #ai-tools 4 stories Β· sorted by recency
── more on @openstax-llm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/openstax-llm-tools-f…] indexed:0 read:4min 2026-10-02 Β· β€”