{"slug": "kompyla-a-self-hosted-research-monitor-that-builds-a-cited-wiki", "title": "Kompyla – a self-hosted research monitor that builds a cited wiki", "summary": "Kompyla is a self-hosted research monitor that treats an LLM like a compiler, turning raw documents into structured, cited wiki pages. The tool extends Andrej Karpathy's LLM Knowledge Base pattern with an active retrieval layer that automatically fetches sources from the web, arXiv, GitHub, RSS and YouTube, plus a self-evolving feedback loop that detects gaps and flags stale knowledge. It adds a presentation pipeline exporting HTML, DOCX, PPTX, Marp, PDF, Markdown and charts, and supports scheduled full cycles via daemon or cron.", "body_md": "**Autonomous research agent that builds and evolves a structured knowledge base.**\n\nKompyla treats an LLM like a compiler: raw documents go in, structured wiki pages come out. It extends Andrej Karpathy's LLM Knowledge Base pattern with an active retrieval layer that fetches new sources automatically, a self-evolving feedback loop that detects gaps and flags stale knowledge, and a full presentation pipeline (HTML, DOCX, PPTX, slide decks, charts).\n\n```\nResearch query\n     │\n     ▼\n┌──────────────┐    web · arxiv · GitHub · RSS · YouTube\n│ Search Agent │──────────────────────────────────────────\n└──────┬───────┘   dedup + relevance filter\n       │ clean .md files\n       ▼\n  raw/ directory\n       │\n       ▼\n┌──────────────┐\n│  Compiler    │──── LLM: raw → structured wiki page\n└──────┬───────┘\n       │\n       ▼\n  wiki/ (the KB)  ◄──── Health check · Gap detection · Feedback\n       │\n       ▼\n┌──────────────┐\n│   Present    │──── Q&A · slides · charts · HTML · DOCX · PPTX\n└──────────────┘\nflowchart TD\n    %% ============ 1. INITIALISATION ============\n    subgraph INIT[\"1 — Initialisation\"]\n        U1[/User: research domain/]\n        I1[kompyla init]\n        I2[LLM: generate schema]\n        I3[(Domain Schema)]\n        I4[(KB directory structure)]\n        U1 --> I1 --> I2 --> I3\n        I2 --> I4\n    end\n\n    %% ============ 2. RETRIEVAL ============\n    subgraph RETR[\"2 — Retrieval\"]\n        R0[Retrieval Orchestrator]\n        R1[Source connectors:<br/>Web · arXiv · GitHub<br/>RSS · YouTube]\n        R2[Dedup: SHA-256<br/>+ MinHash]\n        R3[Relevance filter<br/>LLM scoring]\n        R4[(raw/ documents)]\n        R0 --> R1 --> R2 --> R3 --> R4\n    end\n\n    %% ============ 3. COMPILATION ============\n    subgraph COMP[\"3 — Compilation\"]\n        C1[Pending raw docs]\n        C2[LLM Compiler]\n        C3{Wiki page<br/>exists?}\n        C4[Write new page]\n        C5[LLM merge pass]\n        C6[Master Index update]\n        C1 --> C2 --> C3\n        C3 -->|No| C4 --> C6\n        C3 -->|Yes| C5 --> C6\n    end\n\n    %% ============ 4. EVOLUTION ============\n    subgraph EVOL[\"4 — Evolution\"]\n        E0[(Wiki KB)]\n        E1[Lint: links, stale,<br/>orphans, confidence]\n        E2[Gap Detection:<br/>broken-link + LLM gaps]\n        E3[/User Feedback/]\n        E4[Confidence delta]\n        E0 --> E1\n        E0 --> E2\n        E0 --> E3 --> E4 --> E0\n    end\n\n    %% ============ 5. QUERY & PRESENTATION ============\n    subgraph PRES[\"5 — Query & Presentation\"]\n        Q1[/User question/]\n        Q2[Q&A Engine]\n        Q3[LLM answer<br/>with citations]\n        Q4{Save as<br/>synthesis?}\n        Q5[(Synthesis page)]\n        P0[Export: HTML · DOCX<br/>PPTX · Marp · PDF<br/>Markdown · Charts]\n        Q1 --> Q2 --> Q3 --> Q4\n        Q4 -->|Yes| Q5\n    end\n\n    %% ============ 6. AUTOMATION ============\n    subgraph AUTO[\"6 — Automation\"]\n        A1[Scheduler<br/>daemon / cron]\n        A2[Full cycle: retrieve<br/>compile · lint · gaps]\n        A3[Cross-Reference Bridge]\n        A4[(Other domain KB)]\n        A5[(Comparison report)]\n        A1 --> A2\n        A3 --> A4\n        A3 --> A5\n    end\n\n    %% ============ CROSS-LAYER FLOWS (minimal) ============\n    I3 -->|seed queries| R0\n    R4 --> C1\n    C6 --> E0\n    E2 -->|auto-fill| R0\n    E0 --> Q2\n    E0 --> P0\n    Q5 -.-> E0\n    A2 -->|triggers| R0\n\n    %% ============ STYLING ============\n    classDef initLayer fill:#fff4d6,stroke:#b8860b,stroke-width:2px,color:#000\n    classDef retrLayer fill:#d6e8ff,stroke:#1f4e8c,stroke-width:2px,color:#000\n    classDef compLayer fill:#d6f5d6,stroke:#2e7d32,stroke-width:2px,color:#000\n    classDef evolLayer fill:#ffe2c4,stroke:#cc5500,stroke-width:2px,color:#000\n    classDef presLayer fill:#e8d6ff,stroke:#5b2c8c,stroke-width:2px,color:#000\n    classDef autoLayer fill:#e0e0e0,stroke:#555555,stroke-width:2px,color:#000\n\n    class INIT initLayer\n    class RETR retrLayer\n    class COMP compLayer\n    class EVOL evolLayer\n    class PRES presLayer\n    class AUTO autoLayer\n\n    ...\n```\n\n| Capability | Description | \n|---|---|\n| **KB scaffolding** | LLM generates a domain schema (page types, entity categories, seed queries) from a plain-English topic | \n| **Agentic retrieval** | Searches the web (Serper / Brave / Exa / SerpAPI; DuckDuckGo fallback), arXiv, GitHub, RSS feeds, and YouTube transcripts | \n| **Deduplication** | SHA-256 exact matching + MinHash LSH for near-duplicate detection (Jaccard ≥ 0.85) | \n| **Incremental compilation** | Raw `.md` files are compiled into structured wiki pages; when a page already exists, a second LLM pass merges new information in | \n| **Health checks** | Finds broken internal links, stale pages (>180 days), low-confidence pages, and orphans | \n| **Gap detection** | Deterministic broken-link gaps + LLM-suggested missing topics | \n| **Q&A** | Natural-language questions answered from the wiki with citations; answers can be saved back as synthesis pages | \n| **Presentation** | HTML, Markdown bundle, DOCX, PPTX, Marp slide decks, and PDF (optional) exports | \n| **Charts** | Confidence histogram, pages-by-type, and raw-docs-by-source PNGs | \n| **Web UI** | Streamlit app with Browse, Search, Ask, and Stats tabs; guided setup screen on first run | \n| **Scheduler** | Periodic research cycle (fetch → compile → lint → gaps) with configurable interval | \n| **Cross-referencing** | Find shared topics between two separate domain KBs | \n| **Feedback** | Flag pages as wrong, outdated, excellent, or unclear; apply to confidence scores | \n| **Synthetic data** | Generate Q&A training pairs from high-confidence pages for model fine-tuning | \n\n- Python 3.11 or later\n- One LLM provider — see [LLM providers](#llm-providers) below\n\nOptional extras:\n\n```\npip install \"kompyla[pdf]\"          # WeasyPrint for PDF export\npip install \"kompyla[search]\"       # Exa and SerpAPI backends\npip install \"kompyla[pdf,search]\"   # both\n```\n\nKompyla supports four providers. The `provider` key in `config.yaml` selects which one is used.\n\n| Provider | Config value | Required env var | Notes | \n|---|---|---|---|\n| **Ollama** | `ollama` | — | Fully offline. Install from [ollama.com](https://ollama.com) , then`ollama pull llama3.2` | \n| **Anthropic** | `anthropic` | `ANTHROPIC_API_KEY` | Claude models (e.g. `claude-sonnet-4-6` ) | \n| **OpenAI** | `openai` | `OPENAI_API_KEY` | GPT-4o and other OpenAI models | \n| **Gemini** | `gemini` | `GEMINI_API_KEY` | Gemini 2.0 Flash and other Google models (uses `google-genai` SDK) | \n\n```\ngit clone https://github.com/damien220/kompyla.git\ncd kompyla\npython -m venv .venv\nsource .venv/bin/activate        # Windows: .venv\\Scripts\\activate\npip install -e \".[dev]\"\npip install kompyla\n```\n\nKompyla ships with a `Dockerfile` and `docker-compose.yml` so you can run the full stack — Streamlit UI, CLI, scheduler, and optionally an offline Ollama instance — without touching your local Python environment.\n\n```\ndocker build -t kompyla:latest .\n```\n\nThe image installs `kompyla[pdf,search]` so all LLM providers and all web search backends are available out of the box.\n\nOn first start the KB directory (`/kb`) is empty. The Streamlit UI detects this automatically and shows a **setup screen** where you type your research domain and click **Initialise KB**. The LLM generates the domain schema and the app reloads into the normal Browse / Search / Ask / Stats view.\n\nIf you prefer the CLI:\n\n```\ndocker compose --profile cli run --rm cli init \"electric vehicles\" --kb /kb\n# With Anthropic\nANTHROPIC_API_KEY=sk-... KOMPYLA_KB_PATH=./my_kb docker compose up kompyla-ui\n\n# With a locally running Ollama on the host\nOLLAMA_BASE_URL=http://host.docker.internal:11434 \\\nKOMPYLA_KB_PATH=./my_kb docker compose up kompyla-ui\n```\n\nOpen `http://localhost:8501` in your browser.\n\n```\n# Pull the model once (cached in a named Docker volume)\ndocker compose --profile ollama run --rm ollama ollama pull llama3.2\n\n# Start UI + Ollama together\nKOMPYLA_KB_PATH=./my_kb docker compose --profile ollama up\n```\n\nThe `cli` service lets you run any `kompyla` command against the mounted KB:\n\n```\n# Compile documents\ndocker compose --profile cli run --rm cli compile --kb /kb\n\n# Ask a question\ndocker compose --profile cli run --rm cli query \"What is solid-state battery?\" --kb /kb\nKOMPYLA_KB_PATH=./my_kb docker compose --profile scheduler up -d scheduler\n```\n\nCopy `.env.example` to `.env` in the project root (never commit `.env`):\n\n```\n# .env\nKOMPYLA_KB_PATH=./my_kb      # host path mounted as /kb inside the container\nKOMPYLA_PORT=8501             # host port for the UI (default 8501)\n\n# LLM — uncomment the provider you use\n# ANTHROPIC_API_KEY=sk-ant-...\n# OPENAI_API_KEY=sk-...\n# GEMINI_API_KEY=AIza...\n# OLLAMA_BASE_URL=http://host.docker.internal:11434\n\n# Web search — first key present wins; DuckDuckGo used if none are set\n# SERPER_API_KEY=...\n# BRAVE_API_KEY=...\n# EXA_API_KEY=...\n# SERPAPI_API_KEY=...\n\n# Optional\n# GITHUB_TOKEN=ghp_...\n```\n\nThen simply:\n\n```\ndocker compose up                          # UI only\ndocker compose --profile ollama up         # UI + bundled Ollama\ndocker compose --profile scheduler up -d  # add background scheduler\n```\n\nCopy the example config and edit it:\n\n```\nmkdir -p ~/.kompyla\ncp config.yaml.example ~/.kompyla/config.yaml\n```\n\nDefault uses Ollama + `llama3.2`. To use Claude instead:\n\n```\n# ~/.kompyla/config.yaml\nllm:\n  provider: anthropic\n  model: claude-sonnet-4-6\n```\n\nSet `ANTHROPIC_API_KEY` in your shell (or add `anthropic_api_key:` to the config file).\n\n```\nkompyla init \"electric vehicles\" --path ./ev_kb\ncd ev_kb\n```\n\nThis calls the LLM to generate a domain schema with page types, entity categories, and seed search queries.\n\n```\nkompyla search                       # uses schema seed queries, all enabled sources\nkompyla search \"solid-state batteries\" --sources web,arxiv\nkompyla fetch https://example.com/article\nkompyla add-youtube https://www.youtube.com/watch?v=...\n```\n\nFetched documents land in `raw/` as markdown files.\n\n```\nkompyla compile\n```\n\nEach raw document is transformed by the LLM into a structured wiki page in `wiki/`. If a page for that topic already exists, the new content is merged in.\n\n```\nkompyla status                       # overview metrics\nkompyla query \"What is the range of the Tesla Model 3?\"\nkompyla serve                        # open http://localhost:8501\nkompyla init <domain>                  Scaffold a new KB and generate its domain schema\nkompyla compile                        Compile raw/ documents into wiki/ pages\nkompyla status                         Show KB metrics (pages, raw docs, confidence)\n\nkompyla search [query]                 Retrieve from web/arxiv/GitHub/RSS/YouTube\nkompyla fetch <url>                    Fetch and save a single URL to raw/\nkompyla add-youtube <url>              Fetch a YouTube transcript into raw/\n\nkompyla query <question>               Answer a question from the wiki\nkompyla lint                           Run health checks (broken links, stale, orphans)\nkompyla gaps [--auto-fill]             Detect knowledge gaps; optionally fill them\n\nkompyla export <title> -f html|md|docx|pptx|pdf|marp\nkompyla export --all -f md|html        Whole-KB bundle\nkompyla slides <title> [--html]        Generate a Marp slide deck\nkompyla chart                          Generate stats PNGs\n\nkompyla serve [--port 8501]            Launch Streamlit UI\n\nkompyla schedule --enable --interval 24   Enable periodic research cycle\nkompyla schedule --run-now               Run one cycle immediately\nkompyla schedule --daemon                Loop forever\nkompyla schedule --status                Show current schedule\n\nkompyla crossref --kb-target path/to/other/kb   Find topic overlaps across KBs\nkompyla feedback <title> --signal wrong|outdated|excellent|unclear\nkompyla feedback --apply               Apply feedback deltas to confidence scores\nkompyla synth [--out training.jsonl]   Generate Q&A training data\n```\n\nFastest things you can do today without code changes:\n\n1. Switch to Groq — it's ~5× faster than other providers\n2. Set use_relevance_filter: false in config — removes the per-doc LLM scoring pass\n3. Reduce max_per_source — fewer docs = fewer compile calls\n4. kompyla search --no-filter / kompyla gaps --no-llm for quick runs\n\n`~/.kompyla/config.yaml` (env vars override file values):\n\n```\nllm:\n  provider: ollama # \"ollama\", \"anthropic\", \"openai\", \"gemini\", or \"groq\"\n  model: llama3.2 # any Ollama model; or \"claude-sonnet-4-6\", \"gpt-4o\", \"gemini-2.0-flash\", \"llama-3.3-70b-versatile\"\n  ollama_base_url: http://localhost:11434\n  # anthropic_api_key: ...  # or ANTHROPIC_API_KEY env var\n  # openai_api_key: ...     # or OPENAI_API_KEY env var\n  # gemini_api_key: ...     # or GEMINI_API_KEY env var\n  # groq_api_key: ...       # or GROQ_API_KEY env var  (free tier at console.groq.com)\n\nretrieval:\n  enabled_sources: [web, arxiv, github, rss] # add \"youtube\" if needed\n  max_per_source: 5\n  min_relevance: 0.5\n  use_relevance_filter: true\n  # Web search — first key present wins; DuckDuckGo used as free fallback if none set\n  # serper_api_key: ...     # or SERPER_API_KEY env var\n  # brave_api_key: ...      # or BRAVE_API_KEY env var\n  # exa_api_key: ...        # or EXA_API_KEY env var   (requires kompyla[search])\n  # serpapi_api_key: ...    # or SERPAPI_API_KEY env var (requires kompyla[search])\n  # github_token: ...       # or GITHUB_TOKEN env var\n  rss_feeds:\n    - https://hnrss.org/frontpage\n  youtube_languages: [en]\n```\n\n| Variable | Purpose | \n|---|---|\n| `ANTHROPIC_API_KEY` | Anthropic (Claude) API key | \n| `OPENAI_API_KEY` | OpenAI API key | \n| `GEMINI_API_KEY` | Google Gemini API key | \n| `OLLAMA_BASE_URL` | Override Ollama server URL (default `http://localhost:11434` ) | \n| `SERPER_API_KEY` | Serper web search — Google results, $1/1K queries | \n| `BRAVE_API_KEY` | Brave Search API — free tier: 2 K queries/month | \n| `EXA_API_KEY` | Exa.ai semantic search — free tier available ( `kompyla[search]` required) | \n| `SERPAPI_API_KEY` | SerpAPI multi-engine search ( `kompyla[search]` required) | \n| `GITHUB_TOKEN` | GitHub API token (raises rate limits for the GitHub connector) | \n| `KOMPYLA_KB` | Default KB path (used by `kompyla serve` and the Docker image) | \n\n**Web search fallback** — if no search API key is set, Kompyla falls back to DuckDuckGo automatically (no key required, no extra package). Snippet results are enriched with full page text via trafilatura.\n\n```\nmy_kb/\n├── kompyla.yaml          Domain config + schedule state\n├── raw/                  Fetched source documents (auto-populated)\n│   ├── web/\n│   ├── arxiv/\n│   ├── github/\n│   └── youtube/\n├── wiki/                 Compiled structured wiki pages\n├── index/\n│   ├── schema.yaml       Domain schema (page types, entities, relationships)\n│   ├── index.md          Master index grouped by page type\n│   ├── meta.db           SQLite metadata index\n│   └── feedback.db       User feedback store\n└── outputs/              Exports (HTML, DOCX, PPTX, charts, training data)\n    ├── charts/\n    └── training_data.jsonl\n```\n\n`skills/kompyla-research/` packages Kompyla's engine as an [Agent Skill](https://agentskills.io), so an agent host can drive an existing KB — ingest, compile, query, lint, gaps — from inside its own session. The bundled script does only the deterministic work (retrieval + dedup, the SQLite index, page rendering, link analysis, keyword ranking) by importing `kompyla` in-process; the host agent's model does the compiling, merging and answering, so **the skill needs no LLM API key of its own**. Pages it writes are identical to those from `kompyla compile`, so a KB can be worked on from both sides.\n\n```\ncp -r skills/kompyla-research ~/.claude/skills/    # then: pip install kompyla\n```\n\nIt operates on an existing KB only — create one first with `kompyla init`. See [` skills/kompyla-research/README.md`](https://github.com/damien220/kompyla/blob/master/skills/kompyla-research/README.md).\n\nKompyla works entirely offline with Ollama — no API key or internet connection required for the LLM step.\n\n```\n# Install Ollama: https://ollama.com\nollama serve\nollama pull llama3.2          # or llama3.1, mistral, qwen2.5, etc.\n\n# Set provider in config\n# llm:\n#   provider: ollama\n#   model: llama3.2\n```\n\nRetrieval connectors (web, GitHub, YouTube) still require internet access, but the compilation, Q&A, and gap-detection steps are all local.\n\n```\npip install -e \".[dev]\"\npytest                   # 70 tests across all phases\npytest tests/test_phase5.py -v   # Phase 5 only\n```\n\nThe test suite covers: deduplication, YouTube transcript parsing, KB health checks, Q&A page selection, all presenter/export modules, scheduler logic, feedback store, cross-KB referencing, and synth data parsing.\n\n```\nkompyla/\n├── schema/         Domain schema generation and Pydantic models\n├── storage/        KBLayout (filesystem), MetaIndex (SQLite)\n├── llm/            LLMProvider ABC + OllamaProvider, AnthropicProvider, OpenAIProvider, GeminiProvider\n├── retriever/      SourceConnector ABC + Web (multi-backend), arXiv, GitHub, RSS, YouTube\n├── filter/         RelevanceScorer (LLM), Deduplicator (SHA-256 + MinHash)\n├── compiler/       raw/ → wiki/ pipeline, incremental merge, linker\n├── evolver/        lint, gap detection, confidence helpers\n├── query/          Q&A with citation, synthesis page filing\n├── presenter/      HTML, Markdown, DOCX, PPTX, Marp, PDF, charts\n├── ui/             Streamlit app (Setup / Browse / Search / Ask / Stats)\n├── scheduler/      Periodic cycle runner and schedule state\n├── crossref/       Multi-KB topic bridge\n├── feedback/       Feedback store and confidence delta application\n├── synth/          Synthetic Q&A training data generator\n└── cli.py          Typer CLI — all 17 commands\n```\n\n- **LLM as compiler, not chatbot** — the model transforms raw sources into structured knowledge; it does not answer from its own weights.\n- **Incremental over one-shot** — every operation touches only the relevant slice of the KB.\n- **Relevance before ingest** — the filter layer rejects noise at the edge; a small clean KB beats a large noisy one.\n- **Markdown + SQLite as substrate** — plain files give portability and git-diffable history; SQLite adds queryable metadata without a server.\n- **Confidence and provenance are first-class** — every wiki page carries a confidence score and a list of source documents.\n- **Search with graceful degradation** — API-backed search (Serper → Brave → Exa → SerpAPI) is preferred when a key is configured; DuckDuckGo is the always-available zero-config fallback.\n\nContributions are welcome. Please follow these steps:\n\n1. \n**Fork** the repository and create a feature branch:\n\n```\ngit checkout -b feature/my-improvement\n```\n\n2. \n**Install dev dependencies:**\n\n```\npip install -e \".[dev]\"\n```\n\n3. \n**Write tests** for your change. The project targets 100% test coverage for deterministic modules (no LLM, no network).\n4. \n**Run the full test suite** before opening a PR:\n\n```\npytest\n```\n\n5. \n**Follow existing code style:**\n  - No type comments — use type annotations throughout.\n  - No docstrings for obvious methods — a clear name beats a paragraph.\n  - New source connectors must implement `SourceConnector` from`kompyla/retriever/base.py` .\n  - New CLI commands go in `kompyla/cli.py` using the Typer pattern already established.\n6. \n**Open a pull request** with a clear description of what changed and why.\n\n1. Create `kompyla/llm/<name>_provider.py` implementing`LLMProvider` (single`chat(messages, system) -> str` method).\n2. Import it in `kompyla/llm/__init__.py` and add a branch to`get_provider()` .\n3. Add `<name>_api_key: str | None = None` to`LLMConfig` in`kompyla/config.py` and the`os.getenv(...)` override in`KompylaConfig.load()` .\n4. Expose the env var in `.env.example` and in`docker-compose.yml` (the`x-env` anchor covers all services automatically).\n\n``` python\n# kompyla/retriever/my_source.py\nfrom kompyla.retriever.base import FetchedDoc, SourceConnector\n\nclass MySourceConnector(SourceConnector):\n    @property\n    def name(self) -> str:\n        return \"mysource\"\n\n    def search(self, query: str, max_results: int = 5) -> list[FetchedDoc]:\n        ...\n\n    def fetch_url(self, url: str) -> FetchedDoc | None:\n        ...\n\n    def is_available(self) -> bool:\n        ...\n```\n\nRegister it in `kompyla/retriever/__init__.py` and add it to `_build_connectors()` in `cli.py`.\n\n- KB scaffolding — domain schema generation from a plain-English topic\n- Agentic retrieval — web (Serper / Brave / Exa / SerpAPI / DuckDuckGo fallback), arXiv, GitHub, RSS, and YouTube connectors\n- Incremental compilation with LLM merge pass\n- Health checks — broken links, stale pages, orphans, low-confidence\n- Gap detection — deterministic + LLM-suggested topics\n- Natural-language Q&A with citation and synthesis page filing\n- Presentation exports — HTML, Markdown bundle, DOCX, PPTX, Marp slides, charts, PDF (optional)\n- Streamlit web UI — Setup (first-run), Browse, Search, Ask, Stats\n-  Scheduled research cycle (`kompyla schedule --daemon` )\n-  Multi-KB cross-referencing (`kompyla crossref` )\n-  User feedback integration (`kompyla feedback` )\n-  Synthetic Q&A training data generator (`kompyla synth` )\n-  Docker image + `docker-compose.yml` — all LLM providers, all search backends, first-run UI\n-  Five LLM providers: Ollama, Anthropic, OpenAI, Gemini (`google-genai` SDK), Groq (free tier — Llama 3.3 70B)\n- Comprehensive README with architecture overview, contributing guide, and license\n- Architecture flowchart in README (Mermaid)\n-  Published to PyPI (`pip install kompyla` )\n\n- GitHub Actions CI for automated test runs on every push\n- Embedding-based semantic search (sentence-transformers) as an alternative to keyword overlap\n- Graph view of the wiki (entity relationships, cross-links) in the Streamlit UI\n- Multi-user collaboration mode with shared feedback\n- Example pre-built knowledge bases (electric vehicles, WebGPU frameworks)\n\n**MIT License**\n\nCopyright (c) 2026 Kompyla Contributors\n\nPermission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the \"Software\"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:\n\nThe above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.\n\nTHE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.\n\nKompyla is inspired by [Andrej Karpathy's `llm-wiki` gist](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f) — a three-layer pattern of raw source documents, an LLM compiler, and a linted, queryable wiki output. Kompyla extends this with an active multi-source retrieval agent, a self-evolving feedback loop, cross-KB referencing, and a full presentation/export pipeline.", "url": "https://wpnews.pro/news/kompyla-a-self-hosted-research-monitor-that-builds-a-cited-wiki", "canonical_source": "https://github.com/damien220/kompyla", "published_at": "2026-09-29 17:08:27+00:00", "updated_at": "2026-09-29 17:18:30.155315+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "large-language-models", "ai-research"], "entities": ["Kompyla", "Andrej Karpathy", "LLM Knowledge Base", "arXiv", "GitHub", "YouTube"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/kompyla-a-self-hosted-research-monitor-that-builds-a-cited-wiki", "markdown": "https://wpnews.pro/news/kompyla-a-self-hosted-research-monitor-that-builds-a-cited-wiki.md", "text": "https://wpnews.pro/news/kompyla-a-self-hosted-research-monitor-that-builds-a-cited-wiki.txt", "jsonld": "https://wpnews.pro/news/kompyla-a-self-hosted-research-monitor-that-builds-a-cited-wiki.jsonld"}}