OpenStax-LLM: tools for OpenStax LLM access OpenStax-LLM, a new open-source toolkit built on the openstax-md project, converts OpenStax college textbooks into citation-aware, formula-safe datasets for vector search and LLM fine-tuning. The repository ships a Python library and CLI, a Model Context Protocol server (openstax-llm-mcp), and an agent skill, and can compile and chunk any of the 90-plus OpenStax catalog volumes into JSONL with globally unique chunk_ids compatible with Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face datasets. The project targets generic chunkers that split math formulas, separate worked examples from solutions, and lose chapter hierarchy, enforcing a max_words ceiling and never placing a boundary inside unclosed $...$ or $$...$$ blocks or fenced code blocks. Pedagogical semantic chunking, RAG dataset preparation, and LLM fine-tuning pipelines from OpenStax textbooks. Built on top of openstax-md https://github.com/michaelnavazhylau/openstax-md , openstax-llm transforms OpenStax college textbooks into structured, citation-aware, formula-safe datasets for vector search RAG and model fine-tuning. This repository ships three things that share one core: | Artifact | What it is | Entry point | |---|---|---| | openstax-llm https://github.com/michaelnavazhylau/openstax-llm/blob/main/packages/openstax-llm | Python library and CLI | openstax-llm | | openstax-llm-mcp https://github.com/michaelnavazhylau/openstax-llm/blob/main/packages/openstax-llm-mcp | Model Context Protocol server | openstax-llm-mcp | | skills/openstax-llm https://github.com/michaelnavazhylau/openstax-llm/blob/main/skills/openstax-llm | Agent skill for coding assistants | /skill:openstax-llm | Generic chunkers simple character or recursive token splitters break down on technical academic textbooks: - They cut mathematical formulas in half $x^2 + \dots$ split from \dots + y^2$ . - They separate worked examples from their solutions. - They lose the chapter and section hierarchy that citations depend on. openstax-llm provides: 1. Pedagogical boundary awareness — worked examples Example 1.1 , problem sets, definitions, and summaries are kept whole. 2. Formula integrity — a chunk boundary is never placed inside an unclosed $...$ or $$...$$ block, or inside a fenced code block. 3. Enforced size ceiling — max words is actually honoured; oversized paragraphs are split at sentence boundaries that lie outside math. 4. On-demand compilation — any textbook in the OpenStax catalog 90 volumes and counting is pulled and chunked without manual data management. 5. Vector-store-ready output — JSONL with globally unique chunk id s, compatible with Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face datasets . Library + CLI uv add openstax-llm CLI as a standalone tool uv tool install openstax-llm MCP server, one-shot via uvx uvx openstax-llm-mcp Or run the container: docker build -t openstax-llm-mcp . To track an unreleased commit instead, install from git: uv add "git+https://github.com/michaelnavazhylau/openstax-llm.git subdirectory=packages/openstax-llm" uvx --from "git+https://github.com/michaelnavazhylau/openstax-llm.git subdirectory=packages/openstax-llm-mcp" openstax-llm-mcp Search the catalog for available textbooks openstax-llm search physics Inspect a textbook's chunk statistics and section index openstax-llm info astronomy-2e Compile and chunk a textbook into a JSONL dataset openstax-llm prepare astronomy-2e -o datasets/astronomy-2e.jsonl Verify an export before using it downstream openstax-llm validate datasets/astronomy-2e.jsonl validate checks formula integrity, chunk id uniqueness, and provenance, and exits non-zero on structural errors. Add --strict to also fail on warnings such as front matter that has no section number. Always run it before loading a dataset into an index: a split formula that reaches an embedding store is very hard to detect afterwards. python from openstax llm import DocumentChunker, TextBookDataset chunker = DocumentChunker target words=400, max words=600, overlap words=50 dataset = TextBookDataset.from textbook "calculus-volume-1", chunker=chunker dataset.to jsonl "calculus.jsonl" print dataset.summary {'book slug': 'calculus-volume-1', 'total chunks': 1445, 'total words': 262184, ...} DocumentChunker also works on arbitrary markdown: python from openstax llm import DocumentChunker chunks = DocumentChunker .chunk markdown " 1.2 Functions\n\nAn inline formula $f x =x^2$ stays intact.\n", book slug="my-notes", section="1.2", section title="Functions", Exposes the same three operations to any MCP client over stdio or streamable HTTP. Register with Pi pi mcp add openstax-llm -- uvx openstax-llm-mcp Or serve over HTTP docker run --rm -p 8765:8765 openstax-llm-mcp | Tool | Purpose | |---|---| | search catalog query, limit | Resolve a subject to a canonical slug offline | | inspect textbook target | Chunk totals plus a per-section index | | prepare textbook target, out | Export JSONL, sandboxed to the server's output directory | Resources: textbook://