{"slug": "openstax-llm-tools-for-openstax-llm-access", "title": "OpenStax-LLM: tools for OpenStax LLM access", "summary": "OpenStax-LLM, a new open-source toolkit built on the openstax-md project, converts OpenStax college textbooks into citation-aware, formula-safe datasets for vector search and LLM fine-tuning. The repository ships a Python library and CLI, a Model Context Protocol server (openstax-llm-mcp), and an agent skill, and can compile and chunk any of the 90-plus OpenStax catalog volumes into JSONL with globally unique chunk_ids compatible with Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face datasets. The project targets generic chunkers that split math formulas, separate worked examples from solutions, and lose chapter hierarchy, enforcing a max_words ceiling and never placing a boundary inside unclosed $...$ or $$...$$ blocks or fenced code blocks.", "body_md": "**Pedagogical semantic chunking, RAG dataset preparation, and LLM fine-tuning pipelines from OpenStax textbooks.**\n\nBuilt on top of [`openstax-md`](https://github.com/michaelnavazhylau/openstax-md), `openstax-llm`\ntransforms OpenStax college textbooks into structured, citation-aware, formula-safe datasets\nfor vector search (RAG) and model fine-tuning.\n\nThis repository ships three things that share one core:\n\n| Artifact | What it is | Entry point | \n|---|---|---|\n| [`openstax-llm`](https://github.com/michaelnavazhylau/openstax-llm/blob/main/packages/openstax-llm) | Python library and CLI | `openstax-llm` | \n| [`openstax-llm-mcp`](https://github.com/michaelnavazhylau/openstax-llm/blob/main/packages/openstax-llm-mcp) | Model Context Protocol server | `openstax-llm-mcp` | \n| [`skills/openstax-llm`](https://github.com/michaelnavazhylau/openstax-llm/blob/main/skills/openstax-llm) | Agent skill for coding assistants | `/skill:openstax-llm` | \n\nGeneric chunkers (simple character or recursive token splitters) break down on technical academic textbooks:\n\n- They cut mathematical formulas in half (`$x^2 + \\dots$` split from`\\dots + y^2$` ).\n- They separate worked examples from their solutions.\n- They lose the chapter and section hierarchy that citations depend on.\n\n`openstax-llm` provides:\n\n1. **Pedagogical boundary awareness** — worked examples (`Example 1.1` ), problem sets,\ndefinitions, and summaries are kept whole.\n2. **Formula integrity** — a chunk boundary is never placed inside an unclosed`$...$` or`$$...$$` block, or inside a fenced code block.\n3. **Enforced size ceiling** —`max_words` is actually honoured; oversized paragraphs are\nsplit at sentence boundaries that lie outside math.\n4. **On-demand compilation** — any textbook in the OpenStax catalog (90 volumes and\ncounting) is pulled and chunked without manual data management.\n5. **Vector-store-ready output** — JSONL with globally unique`chunk_id` s, compatible with\nChroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face`datasets` .\n\n```\n# Library + CLI\nuv add openstax-llm\n\n# CLI as a standalone tool\nuv tool install openstax-llm\n\n# MCP server, one-shot via uvx\nuvx openstax-llm-mcp\n```\n\nOr run the container:\n\n```\ndocker build -t openstax-llm-mcp .\n```\n\nTo track an unreleased commit instead, install from git:\n\n```\nuv add \"git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm\"\nuvx --from \"git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm-mcp\" openstax-llm-mcp\n# Search the catalog for available textbooks\nopenstax-llm search physics\n\n# Inspect a textbook's chunk statistics and section index\nopenstax-llm info astronomy-2e\n\n# Compile and chunk a textbook into a JSONL dataset\nopenstax-llm prepare astronomy-2e -o datasets/astronomy-2e.jsonl\n\n# Verify an export before using it downstream\nopenstax-llm validate datasets/astronomy-2e.jsonl\n```\n\n`validate` checks formula integrity, `chunk_id` uniqueness, and provenance, and exits\nnon-zero on structural errors. Add `--strict` to also fail on warnings such as front matter\nthat has no section number. Always run it before loading a dataset into an index: a split\nformula that reaches an embedding store is very hard to detect afterwards.\n\n``` python\nfrom openstax_llm import DocumentChunker, TextBookDataset\n\nchunker = DocumentChunker(target_words=400, max_words=600, overlap_words=50)\ndataset = TextBookDataset.from_textbook(\"calculus-volume-1\", chunker=chunker)\ndataset.to_jsonl(\"calculus.jsonl\")\n\nprint(dataset.summary())\n# {'book_slug': 'calculus-volume-1', 'total_chunks': 1445, 'total_words': 262184, ...}\n```\n\n`DocumentChunker` also works on arbitrary markdown:\n\n``` python\nfrom openstax_llm import DocumentChunker\n\nchunks = DocumentChunker().chunk_markdown(\n    \"## 1.2 Functions\\n\\nAn inline formula $f(x)=x^2$ stays intact.\\n\",\n    book_slug=\"my-notes\",\n    section=\"1.2\",\n    section_title=\"Functions\",\n)\n```\n\nExposes the same three operations to any MCP client over stdio or streamable HTTP.\n\n```\n# Register with Pi\npi mcp add openstax-llm -- uvx openstax-llm-mcp\n\n# Or serve over HTTP\ndocker run --rm -p 8765:8765 openstax-llm-mcp\n```\n\n| Tool | Purpose | \n|---|---|\n| `search_catalog(query, limit)` | Resolve a subject to a canonical slug (offline) | \n| `inspect_textbook(target)` | Chunk totals plus a per-section index | \n| `prepare_textbook(target, out)` | Export JSONL, sandboxed to the server's output directory | \n\nResources: `textbook://<slug>` for an overview and `textbook://<slug>/<section>` for every\nchunk in one section. See [`packages/openstax-llm-mcp/README.md`](https://github.com/michaelnavazhylau/openstax-llm/blob/main/packages/openstax-llm-mcp/README.md)\nfor options, tool schemas, and failure modes.\n\n```\nnpx skills add michaelnavazhylau/openstax-llm\n```\n\nAlso on [skills.sh](https://skills.sh/michaelnavazhylau/openstax-llm) — that directory is populated\nfrom anonymous install telemetry, so the listing appears only after the first\n`npx skills add` (the command above is always the canonical path).\n\nThe skill teaches an agent *when* and *how* to reach for these tools: resolving slugs\ninstead of guessing titles, verifying exports before indexing, tuning chunk sizes, and\nloading the result into Chroma, Qdrant, Pinecone, or Hugging Face. It deliberately contains\nno chunking logic of its own.\n\n| Field | Type | Description | \n|---|---|---|\n| `chunk_id` | `string` | Unique within a book: `<section>-c<index>` , e.g.`1.2-c003` | \n| `text` | `string` | Markdown with intact LaTeX math | \n| `book_slug` | `string` | Canonical OpenStax slug | \n| `book_title` | `string` | Human-readable title | \n| `chapter` | `string` | Chapter number from the section hierarchy | \n| `section` | `string` | Section number, e.g. `1.2` | \n| `section_title` | `string` | Module title | \n| `chunk_type` | `string` | `prose` ,`example` ,`exercise` ,`definition` ,`summary` | \n| `word_count` | `integer` | Whitespace-delimited word count | \n| `token_est` | `integer` | Heuristic estimate, `words * 1.3` | \n| `metadata` | `object` | Carries `module_id` ; free for downstream use | \n\nFront matter — prefaces, formula tables, chapter introductions — has no section number. Its\n`chunk_id` is prefixed with the module id and its provenance lives in `metadata.module_id`.\nThe machine-readable schema is at\n[`skills/openstax-llm/assets/chunk.schema.json`](https://github.com/michaelnavazhylau/openstax-llm/blob/main/skills/openstax-llm/assets/chunk.schema.json).\n\n```\nuv sync --all-groups --all-packages\n\nuv run pytest -v\nuv run ruff check . && uv run ruff format --check .\nuv run mypy\n\n# Real-textbook integration tests (clones a book, needs network)\nOPENSTAX_LLM_NETWORK_TESTS=1 uv run pytest tests/test_integration.py -v\n```\n\nThe repository is a uv workspace: the root `pyproject.toml` declares members and owns the\nshared tool configuration, while each package under `packages/` is independently\ndistributable. See [AGENTS.md](https://github.com/michaelnavazhylau/openstax-llm/blob/main/AGENTS.md) for the architecture and engineering invariants,\nand [docs/PUBLISHING.md](https://github.com/michaelnavazhylau/openstax-llm/blob/main/docs/PUBLISHING.md) for the release process.\n\nMIT", "url": "https://wpnews.pro/news/openstax-llm-tools-for-openstax-llm-access", "canonical_source": "https://github.com/michaelnavazhylau/openstax-llm", "published_at": "2026-10-02 20:03:08+00:00", "updated_at": "2026-10-02 20:36:25.885919+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "agent-protocols", "developer-tools", "ai-research"], "entities": ["OpenStax-LLM", "openstax-md", "Michael Navazhylau", "Model Context Protocol", "OpenStax", "Chroma", "Qdrant", "Pinecone"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openstax-llm-tools-for-openstax-llm-access", "markdown": "https://wpnews.pro/news/openstax-llm-tools-for-openstax-llm-access.md", "text": "https://wpnews.pro/news/openstax-llm-tools-for-openstax-llm-access.txt", "jsonld": "https://wpnews.pro/news/openstax-llm-tools-for-openstax-llm-access.jsonld"}}