cd /news/ai-agents/kompyla-a-self-hosted-research-monit… · home › topics › ai-agents › article
[ARTICLE · art-141875] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Kompyla – a self-hosted research monitor that builds a cited wiki

Kompyla is a self-hosted research monitor that treats an LLM like a compiler, turning raw documents into structured, cited wiki pages. The tool extends Andrej Karpathy's LLM Knowledge Base pattern with an active retrieval layer that automatically fetches sources from the web, arXiv, GitHub, RSS and YouTube, plus a self-evolving feedback loop that detects gaps and flags stale knowledge. It adds a presentation pipeline exporting HTML, DOCX, PPTX, Marp, PDF, Markdown and charts, and supports scheduled full cycles via daemon or cron.

read16 min views4 publishedSep 29, 2026
Kompyla – a self-hosted research monitor that builds a cited wiki
Image: Michielbdejong (auto-discovered)

Autonomous research agent that builds and evolves a structured knowledge base.

Kompyla treats an LLM like a compiler: raw documents go in, structured wiki pages come out. It extends Andrej Karpathy's LLM Knowledge Base pattern with an active retrieval layer that fetches new sources automatically, a self-evolving feedback loop that detects gaps and flags stale knowledge, and a full presentation pipeline (HTML, DOCX, PPTX, slide decks, charts).

Research query
     │
     ▼
┌──────────────┐    web · arxiv · GitHub · RSS · YouTube
│ Search Agent │──────────────────────────────────────────
└──────┬───────┘   dedup + relevance filter
       │ clean .md files
       ▼
  raw/ directory
       │
       ▼
┌──────────────┐
│  Compiler    │──── LLM: raw → structured wiki page
└──────┬───────┘
       │
       ▼
  wiki/ (the KB)  ◄──── Health check · Gap detection · Feedback
       │
       ▼
┌──────────────┐
│   Present    │──── Q&A · slides · charts · HTML · DOCX · PPTX
└──────────────┘
flowchart TD
    %% ============ 1. INITIALISATION ============
    subgraph INIT["1 — Initialisation"]
        U1[/User: research domain/]
        I1[kompyla init]
        I2[LLM: generate schema]
        I3[(Domain Schema)]
        I4[(KB directory structure)]
        U1 --> I1 --> I2 --> I3
        I2 --> I4
    end

    %% ============ 2. RETRIEVAL ============
    subgraph RETR["2 — Retrieval"]
        R0[Retrieval Orchestrator]
        R1[Source connectors:<br/>Web · arXiv · GitHub<br/>RSS · YouTube]
        R2[Dedup: SHA-256<br/>+ MinHash]
        R3[Relevance filter<br/>LLM scoring]
        R4[(raw/ documents)]
        R0 --> R1 --> R2 --> R3 --> R4
    end

    %% ============ 3. COMPILATION ============
    subgraph COMP["3 — Compilation"]
        C1[Pending raw docs]
        C2[LLM Compiler]
        C3{Wiki page<br/>exists?}
        C4[Write new page]
        C5[LLM merge pass]
        C6[Master Index update]
        C1 --> C2 --> C3
        C3 -->|No| C4 --> C6
        C3 -->|Yes| C5 --> C6
    end

    %% ============ 4. EVOLUTION ============
    subgraph EVOL["4 — Evolution"]
        E0[(Wiki KB)]
        E1[Lint: links, stale,<br/>orphans, confidence]
        E2[Gap Detection:<br/>broken-link + LLM gaps]
        E3[/User Feedback/]
        E4[Confidence delta]
        E0 --> E1
        E0 --> E2
        E0 --> E3 --> E4 --> E0
    end

    %% ============ 5. QUERY & PRESENTATION ============
    subgraph PRES["5 — Query & Presentation"]
        Q1[/User question/]
        Q2[Q&A Engine]
        Q3[LLM answer<br/>with citations]
        Q4{Save as<br/>synthesis?}
        Q5[(Synthesis page)]
        P0[Export: HTML · DOCX<br/>PPTX · Marp · PDF<br/>Markdown · Charts]
        Q1 --> Q2 --> Q3 --> Q4
        Q4 -->|Yes| Q5
    end

    %% ============ 6. AUTOMATION ============
    subgraph AUTO["6 — Automation"]
        A1[Scheduler<br/>daemon / cron]
        A2[Full cycle: retrieve<br/>compile · lint · gaps]
        A3[Cross-Reference Bridge]
        A4[(Other domain KB)]
        A5[(Comparison report)]
        A1 --> A2
        A3 --> A4
        A3 --> A5
    end

    %% ============ CROSS-LAYER FLOWS (minimal) ============
    I3 -->|seed queries| R0
    R4 --> C1
    C6 --> E0
    E2 -->|auto-fill| R0
    E0 --> Q2
    E0 --> P0
    Q5 -.-> E0
    A2 -->|triggers| R0

    %% ============ STYLING ============
    classDef initLayer fill:#fff4d6,stroke:#b8860b,stroke-width:2px,color:#000
    classDef retrLayer fill:#d6e8ff,stroke:#1f4e8c,stroke-width:2px,color:#000
    classDef compLayer fill:#d6f5d6,stroke:#2e7d32,stroke-width:2px,color:#000
    classDef evolLayer fill:#ffe2c4,stroke:#cc5500,stroke-width:2px,color:#000
    classDef presLayer fill:#e8d6ff,stroke:#5b2c8c,stroke-width:2px,color:#000
    classDef autoLayer fill:#e0e0e0,stroke:#555555,stroke-width:2px,color:#000

    class INIT initLayer
    class RETR retrLayer
    class COMP compLayer
    class EVOL evolLayer
    class PRES presLayer
    class AUTO autoLayer

    ...
Capability Description
KB scaffolding LLM generates a domain schema (page types, entity categories, seed queries) from a plain-English topic
Agentic retrieval Searches the web (Serper / Brave / Exa / SerpAPI; DuckDuckGo fallback), arXiv, GitHub, RSS feeds, and YouTube transcripts
Deduplication SHA-256 exact matching + MinHash LSH for near-duplicate detection (Jaccard ≥ 0.85)
Incremental compilation Raw .md files are compiled into structured wiki pages; when a page already exists, a second LLM pass merges new information in
Health checks Finds broken internal links, stale pages (>180 days), low-confidence pages, and orphans
Gap detection Deterministic broken-link gaps + LLM-suggested missing topics
Q&A Natural-language questions answered from the wiki with citations; answers can be saved back as synthesis pages
Presentation HTML, Markdown bundle, DOCX, PPTX, Marp slide decks, and PDF (optional) exports
Charts Confidence histogram, pages-by-type, and raw-docs-by-source PNGs
Web UI Streamlit app with Browse, Search, Ask, and Stats tabs; guided setup screen on first run
Scheduler Periodic research cycle (fetch → compile → lint → gaps) with configurable interval
Cross-referencing Find shared topics between two separate domain KBs
Feedback Flag pages as wrong, outdated, excellent, or unclear; apply to confidence scores
Synthetic data Generate Q&A training pairs from high-confidence pages for model fine-tuning
  • Python 3.11 or later
  • One LLM provider — see LLM providers below

Optional extras:

pip install "kompyla[pdf]"          # WeasyPrint for PDF export
pip install "kompyla[search]"       # Exa and SerpAPI backends
pip install "kompyla[pdf,search]"   # both

Kompyla supports four providers. The provider key in config.yaml selects which one is used.

Provider Config value Required env var Notes
Ollama ollama — Fully offline. Install from ollama.com , thenollama pull llama3.2
Anthropic anthropic ANTHROPIC_API_KEY Claude models (e.g. claude-sonnet-4-6 )
OpenAI openai OPENAI_API_KEY GPT-4o and other OpenAI models
Gemini gemini GEMINI_API_KEY Gemini 2.0 Flash and other Google models (uses google-genai SDK)
git clone https://github.com/damien220/kompyla.git
cd kompyla
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
pip install kompyla

Kompyla ships with a Dockerfile and docker-compose.yml so you can run the full stack — Streamlit UI, CLI, scheduler, and optionally an offline Ollama instance — without touching your local Python environment.

docker build -t kompyla:latest .

The image installs kompyla[pdf,search] so all LLM providers and all web search backends are available out of the box.

On first start the KB directory (/kb) is empty. The Streamlit UI detects this automatically and shows a setup screen where you type your research domain and click Initialise KB. The LLM generates the domain schema and the app reloads into the normal Browse / Search / Ask / Stats view.

If you prefer the CLI:

docker compose --profile cli run --rm cli init "electric vehicles" --kb /kb
ANTHROPIC_API_KEY=sk-... KOMPYLA_KB_PATH=./my_kb docker compose up kompyla-ui

OLLAMA_BASE_URL=http://host.docker.internal:11434 \
KOMPYLA_KB_PATH=./my_kb docker compose up kompyla-ui

Open http://localhost:8501 in your browser.

docker compose --profile ollama run --rm ollama ollama pull llama3.2

KOMPYLA_KB_PATH=./my_kb docker compose --profile ollama up

The cli service lets you run any kompyla command against the mounted KB:

docker compose --profile cli run --rm cli compile --kb /kb

docker compose --profile cli run --rm cli query "What is solid-state battery?" --kb /kb
KOMPYLA_KB_PATH=./my_kb docker compose --profile scheduler up -d scheduler

Copy .env.example to .env in the project root (never commit .env):

KOMPYLA_KB_PATH=./my_kb      # host path mounted as /kb inside the container
KOMPYLA_PORT=8501             # host port for the UI (default 8501)



Then simply:

docker compose up                          # UI only
docker compose --profile ollama up         # UI + bundled Ollama
docker compose --profile scheduler up -d  # add background scheduler

Copy the example config and edit it:

mkdir -p ~/.kompyla
cp config.yaml.example ~/.kompyla/config.yaml

Default uses Ollama + llama3.2. To use Claude instead:

llm:
  provider: anthropic
  model: claude-sonnet-4-6

Set ANTHROPIC_API_KEY in your shell (or add anthropic_api_key: to the config file).

kompyla init "electric vehicles" --path ./ev_kb
cd ev_kb

This calls the LLM to generate a domain schema with page types, entity categories, and seed search queries.

kompyla search                       # uses schema seed queries, all enabled sources
kompyla search "solid-state batteries" --sources web,arxiv
kompyla fetch https://example.com/article
kompyla add-youtube https://www.youtube.com/watch?v=...

Fetched documents land in raw/ as markdown files.

kompyla compile

Each raw document is transformed by the LLM into a structured wiki page in wiki/. If a page for that topic already exists, the new content is merged in.

kompyla status                       # overview metrics
kompyla query "What is the range of the Tesla Model 3?"
kompyla serve                        # open http://localhost:8501
kompyla init <domain>                  Scaffold a new KB and generate its domain schema
kompyla compile                        Compile raw/ documents into wiki/ pages
kompyla status                         Show KB metrics (pages, raw docs, confidence)

kompyla search [query]                 Retrieve from web/arxiv/GitHub/RSS/YouTube
kompyla fetch <url>                    Fetch and save a single URL to raw/
kompyla add-youtube <url>              Fetch a YouTube transcript into raw/

kompyla query <question>               Answer a question from the wiki
kompyla lint                           Run health checks (broken links, stale, orphans)
kompyla gaps [--auto-fill]             Detect knowledge gaps; optionally fill them

kompyla export <title> -f html|md|docx|pptx|pdf|marp
kompyla export --all -f md|html        Whole-KB bundle
kompyla slides <title> [--html]        Generate a Marp slide deck
kompyla chart                          Generate stats PNGs

kompyla serve [--port 8501]            Launch Streamlit UI

kompyla schedule --enable --interval 24   Enable periodic research cycle
kompyla schedule --run-now               Run one cycle immediately
kompyla schedule --daemon                Loop forever
kompyla schedule --status                Show current schedule

kompyla crossref --kb-target path/to/other/kb   Find topic overlaps across KBs
kompyla feedback <title> --signal wrong|outdated|excellent|unclear
kompyla feedback --apply               Apply feedback deltas to confidence scores
kompyla synth [--out training.jsonl]   Generate Q&A training data

Fastest things you can do today without code changes:

  1. Switch to Groq — it's ~5× faster than other providers
  2. Set use_relevance_filter: false in config — removes the per-doc LLM scoring pass
  3. Reduce max_per_source — fewer docs = fewer compile calls
  4. kompyla search --no-filter / kompyla gaps --no-llm for quick runs

~/.kompyla/config.yaml (env vars override file values):

llm:
  provider: ollama # "ollama", "anthropic", "openai", "gemini", or "groq"
  model: llama3.2 # any Ollama model; or "claude-sonnet-4-6", "gpt-4o", "gemini-2.0-flash", "llama-3.3-70b-versatile"
  ollama_base_url: http://localhost:11434

retrieval:
  enabled_sources: [web, arxiv, github, rss] # add "youtube" if needed
  max_per_source: 5
  min_relevance: 0.5
  use_relevance_filter: true
  rss_feeds:
    - https://hnrss.org/frontpage
  youtube_languages: [en]
Variable Purpose
ANTHROPIC_API_KEY Anthropic (Claude) API key
OPENAI_API_KEY OpenAI API key
GEMINI_API_KEY Google Gemini API key
OLLAMA_BASE_URL Override Ollama server URL (default http://localhost:11434 )
SERPER_API_KEY Serper web search — Google results, $1/1K queries
BRAVE_API_KEY Brave Search API — free tier: 2 K queries/month
EXA_API_KEY Exa.ai semantic search — free tier available ( kompyla[search] required)
SERPAPI_API_KEY SerpAPI multi-engine search ( kompyla[search] required)
GITHUB_TOKEN GitHub API token (raises rate limits for the GitHub connector)
KOMPYLA_KB Default KB path (used by kompyla serve and the Docker image)

Web search fallback — if no search API key is set, Kompyla falls back to DuckDuckGo automatically (no key required, no extra package). Snippet results are enriched with full page text via trafilatura.

my_kb/
├── kompyla.yaml          Domain config + schedule state
├── raw/                  Fetched source documents (auto-populated)
│   ├── web/
│   ├── arxiv/
│   ├── github/
│   └── youtube/
├── wiki/                 Compiled structured wiki pages
├── index/
│   ├── schema.yaml       Domain schema (page types, entities, relationships)
│   ├── index.md          Master index grouped by page type
│   ├── meta.db           SQLite metadata index
│   └── feedback.db       User feedback store
└── outputs/              Exports (HTML, DOCX, PPTX, charts, training data)
    ├── charts/
    └── training_data.jsonl

skills/kompyla-research/ packages Kompyla's engine as an Agent Skill, so an agent host can drive an existing KB — ingest, compile, query, lint, gaps — from inside its own session. The bundled script does only the deterministic work (retrieval + dedup, the SQLite index, page rendering, link analysis, keyword ranking) by importing kompyla in-process; the host agent's model does the compiling, merging and answering, so the skill needs no LLM API key of its own. Pages it writes are identical to those from kompyla compile, so a KB can be worked on from both sides.

cp -r skills/kompyla-research ~/.claude/skills/    # then: pip install kompyla

It operates on an existing KB only — create one first with kompyla init. See skills/kompyla-research/README.md.

Kompyla works entirely offline with Ollama — no API key or internet connection required for the LLM step.

ollama serve
ollama pull llama3.2          # or llama3.1, mistral, qwen2.5, etc.

Retrieval connectors (web, GitHub, YouTube) still require internet access, but the compilation, Q&A, and gap-detection steps are all local.

pip install -e ".[dev]"
pytest                   # 70 tests across all phases
pytest tests/test_phase5.py -v   # Phase 5 only

The test suite covers: deduplication, YouTube transcript parsing, KB health checks, Q&A page selection, all presenter/export modules, scheduler logic, feedback store, cross-KB referencing, and synth data parsing.

kompyla/
├── schema/         Domain schema generation and Pydantic models
├── storage/        KBLayout (filesystem), MetaIndex (SQLite)
├── llm/            LLMProvider ABC + OllamaProvider, AnthropicProvider, OpenAIProvider, GeminiProvider
├── retriever/      SourceConnector ABC + Web (multi-backend), arXiv, GitHub, RSS, YouTube
├── filter/         RelevanceScorer (LLM), Deduplicator (SHA-256 + MinHash)
├── compiler/       raw/ → wiki/ pipeline, incremental merge, linker
├── evolver/        lint, gap detection, confidence helpers
├── query/          Q&A with citation, synthesis page filing
├── presenter/      HTML, Markdown, DOCX, PPTX, Marp, PDF, charts
├── ui/             Streamlit app (Setup / Browse / Search / Ask / Stats)
├── scheduler/      Periodic cycle runner and schedule state
├── crossref/       Multi-KB topic bridge
├── feedback/       Feedback store and confidence delta application
├── synth/          Synthetic Q&A training data generator
└── cli.py          Typer CLI — all 17 commands
  • LLM as compiler, not chatbot — the model transforms raw sources into structured knowledge; it does not answer from its own weights.
  • Incremental over one-shot — every operation touches only the relevant slice of the KB.
  • Relevance before ingest — the filter layer rejects noise at the edge; a small clean KB beats a large noisy one.
  • Markdown + SQLite as substrate — plain files give portability and git-diffable history; SQLite adds queryable metadata without a server.
  • Confidence and provenance are first-class — every wiki page carries a confidence score and a list of source documents.
  • Search with graceful degradation — API-backed search (Serper → Brave → Exa → SerpAPI) is preferred when a key is configured; DuckDuckGo is the always-available zero-config fallback.

Contributions are welcome. Please follow these steps:

Fork the repository and create a feature branch:

git checkout -b feature/my-improvement

Install dev dependencies:

pip install -e ".[dev]"

Write tests for your change. The project targets 100% test coverage for deterministic modules (no LLM, no network). 4. Run the full test suite before opening a PR:

pytest

Follow existing code style:

  • No type comments — use type annotations throughout.
  • No docstrings for obvious methods — a clear name beats a paragraph.
  • New source connectors must implement SourceConnector fromkompyla/retriever/base.py .
  • New CLI commands go in kompyla/cli.py using the Typer pattern already established.

Open a pull request with a clear description of what changed and why.

  1. Create kompyla/llm/<name>_provider.py implementingLLMProvider (singlechat(messages, system) -> str method).
  2. Import it in kompyla/llm/__init__.py and add a branch toget_provider() .
  3. Add <name>_api_key: str | None = None toLLMConfig inkompyla/config.py and theos.getenv(...) override inKompylaConfig.load() .
  4. Expose the env var in .env.example and indocker-compose.yml (thex-env anchor covers all services automatically).
from kompyla.retriever.base import FetchedDoc, SourceConnector

class MySourceConnector(SourceConnector):
    @property
    def name(self) -> str:
        return "mysource"

    def search(self, query: str, max_results: int = 5) -> list[FetchedDoc]:
        ...

    def fetch_url(self, url: str) -> FetchedDoc | None:
        ...

    def is_available(self) -> bool:
        ...

Register it in kompyla/retriever/__init__.py and add it to _build_connectors() in cli.py.

  • KB scaffolding — domain schema generation from a plain-English topic

  • Agentic retrieval — web (Serper / Brave / Exa / SerpAPI / DuckDuckGo fallback), arXiv, GitHub, RSS, and YouTube connectors

  • Incremental compilation with LLM merge pass

  • Health checks — broken links, stale pages, orphans, low-confidence

  • Gap detection — deterministic + LLM-suggested topics

  • Natural-language Q&A with citation and synthesis page filing

  • Presentation exports — HTML, Markdown bundle, DOCX, PPTX, Marp slides, charts, PDF (optional)

  • Streamlit web UI — Setup (first-run), Browse, Search, Ask, Stats

  • Scheduled research cycle (kompyla schedule --daemon )

  • Multi-KB cross-referencing (kompyla crossref )

  • User feedback integration (kompyla feedback )

  • Synthetic Q&A training data generator (kompyla synth )

  • Docker image + docker-compose.yml — all LLM providers, all search backends, first-run UI

  • Five LLM providers: Ollama, Anthropic, OpenAI, Gemini (google-genai SDK), Groq (free tier — Llama 3.3 70B)

  • Comprehensive README with architecture overview, contributing guide, and license

  • Architecture flowchart in README (Mermaid)

  • Published to PyPI (pip install kompyla )

  • GitHub Actions CI for automated test runs on every push

  • Embedding-based semantic search (sentence-transformers) as an alternative to keyword overlap

  • Graph view of the wiki (entity relationships, cross-links) in the Streamlit UI

  • Multi-user collaboration mode with shared feedback

  • Example pre-built knowledge bases (electric vehicles, WebGPU frameworks)

MIT License

Copyright (c) 2026 Kompyla Contributors

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

Kompyla is inspired by Andrej Karpathy's llm-wiki gist — a three-layer pattern of raw source documents, an LLM compiler, and a linted, queryable wiki output. Kompyla extends this with an active multi-source retrieval agent, a self-evolving feedback loop, cross-KB referencing, and a full presentation/export pipeline.

── more in #ai-agents 4 stories · sorted by recency
── more on @kompyla 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/kompyla-a-self-hoste…] indexed:0 read:16min 2026-09-29 · —