cd /news/ai-agents/hephaestus-local-first-open-source-a… · home › topics › ai-agents › article
[ARTICLE · art-149173] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Hephaestus: Local-First, Open-Source AI Agents That Train ML Models While You Go Outside

A developer built Hephaestus, a local-first, open-source multi-agent system that takes a plain-language model description and autonomously handles research, dataset ingestion, PyTorch code generation, sandboxed training, and debugging. The stack runs entirely on the user's own machine using the open-weight qwen2.5-coder:14b model via Ollama, with Docker Compose, Redis, MongoDB, and SearXNG, keeping code, data, and API keys off third-party servers. Successful training scripts are saved as reusable skills so subsequent runs start from prior results.

by read4 min views1 publishedOct 11, 2026

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass

Hephaestus is a local-first, autonomous coordinator for machine learning pipelines. You describe a model in plain language, for example "train a CNN to classify these images", and a team of AI agents takes it from there: they read the research, find a dataset, write the PyTorch code, run the training in a sandbox, fix their own bugs, and remember what worked for next time.

How it gets people off the screen. This isn't a hiking app, and I won't pretend it is. The "grass" here is the hours that ML engineers spend babysitting: hunting for papers, reading shape-mismatch tracebacks, re-running a script that died on a missing import, and watching a GPU in case it runs out of memory. Hephaestus is built so the human's screen time is the shortest part of the job. You write one prompt, the agents run the whole loop unattended (research, data, code, training, debugging), and the dashboard exists to be glanced at, not stared at. When a run finishes, the best scripts are saved as reusable skills, so the next run starts further along.

Who it's for: students and researchers who have a GPU and a deadline, and who would rather go outside than debug tensor dimensions. It's also for anyone who wants an agent system that keeps code, data and API keys on their own machine.

Hephaestus is local-first by design, so it runs on your own machine instead of at a public URL, which keeps your data and keys off any server you don't control. You can start it in four commands (Docker is required, plus Ollama for local models):

cp .env.example .env
docker compose up -d redis mongodb searxng
cd backend && uv sync && uv run python main.py     # API on :8000
cd frontend && pnpm install && pnpm dev            # dashboard on :3000

Then open http://localhost:3000. The dashboard shows a live GPU memory graph, a streaming terminal of the training logs, and a multi-agent chat where each message is tagged with the agent that sent it (Research, MLOps, Orchestrator). The repo README has the full setup for both native and full-Docker modes.

Follow these steps to spin up the entire Hephaestus stack locally using Docker Compose.

First, create a .env file from the example template:

cp .env.example .env        # macOS / Linux
copy .env.example .env       # Windows (cmd / PowerShell)

Open the .env file and set the required variables (like SEARXNG_SECRET).

Run the following command from the root of the project to build the images and start the services in detached mode:

docker compose up --build -d

This will spin up:

Once the containers are…

The open-source AI at its core:

qwen2.5-coder:14b, an open-weight coding model running on the user's own GPU. No cloud account is needed to use Hephaestus. How a request flows:

prompt -> Orchestrator -> Triage Router
   simple script -> General Coder ----------------------+
   ML pipeline  -> Research (arXiv + SearXNG, in parallel)
                -> Data Ingestion (Hugging Face Hub -> Arrow files)
                -> Skill retrieval (past successful scripts)
                -> DataPrep agent -> Architecture agent -> Integration agent
                -> Best-of-N candidates (3) -> AST pre-flight gate
                                                          |
   evict LLM from VRAM -> Docker sandbox training <-------+
        |  failure (up to 3 retries)        | success
        v                                   v
   Analyzer -> Patcher -> re-run     Parse metrics -> save skill if better

The design decisions that matter most

<imports>, <model_class>, <training_loop>), so nothing leaks between parts. nn or Data and injects the missing imports. These checks take microseconds, compared with the seconds a Docker start-up costs, so trivial mistakes never reach the sandbox. For quality, the system generates three candidate scripts and scores them on syntax, import completeness and the presence of the metrics line.HEPHAESTUS_METRICS::{...}). If a run beats the previous baseline by at least 0.02, the script is saved to MongoDB as a learned skill, and later runs retrieve matching skills as reference code.pynvml (falling back to simulated numbers on machines without a GPU) and tells Ollama to unload the language model (keep_alive: 0) before training starts, so the LLM and PyTorch never fight over memory. User API keys, for the optional cloud providers, are encrypted at rest with Fernet (AES-128 plus HMAC) and support key rotation through MultiFernet. The backend has 22 pytest modules covering the guardrails, sandbox, encryption, tools and end-to-end pipeline.

Hephaestus works because the pieces underneath it are open. Here is what that made possible that a closed API wouldn't.

The honest trade-off: cloud models are often stronger, and the 14B default needs a reasonably capable GPU. The open approach asks for hardware and gives back control, privacy and zero marginal cost.

Prize Categories #

Best Use of Render ($200 USD) : The Express backend (RAG ingestion, retrieval, generation, and the hourly commit job) and the managed PostgreSQL database both run on Render.

── more in #ai-agents 4 stories · sorted by recency
── more on @hephaestus 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/hephaestus-local-fir…] indexed:0 read:4min 2026-10-11 · —