cd /news/artificial-intelligence/show-hn-speck-cognitive-architecture… · home › topics › artificial-intelligence › article
[ARTICLE · art-145941] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Show HN: Speck – Cognitive architecture for small local LLMs

Speck, a cognitive runtime built on the Genesis runtime, was released as a Show HN project that moves memory, planning, evidence tracking, confidence estimation and metacognition out of the prompt and into persistent deterministic software so that 1B–7B local language models can offload cognitive work. Speck stores cognitive state independently of the loaded model, letting a worker unload, swap or migrate models without losing task state, and it incorporates mechanisms derived from the Artificial Cognitive Architecture Omega Gen2 without being a fork or rewrite of it.

read9 min views1 publishedOct 6, 2026
Show HN: Speck – Cognitive architecture for small local LLMs
Image: Michielbdejong (auto-discovered)

Speck is a cognitive runtime designed to make small language models substantially more capable by moving cognition out of the prompt and into persistent, deterministic software.

Instead of repeatedly asking an LLM to simulate memory, attention, planning, evidence tracking, confidence, learning and metacognition, Speck implements those mechanisms in the runtime itself.

The model is disposable. The cognitive state is not.

A worker can unload its model, change models, restart, or migrate to another compute device without losing the cognitive state of its task.

Speck is built on the Genesis runtime and incorporates mechanisms derived from experiments in Artificial Cognitive Architecture Omega Gen2, but it is not a fork or rewrite of Omega.

Most LLM agents place an enormous amount of responsibility on the model.

The model is expected to:

  • remember what happened
  • determine what matters
  • maintain task state
  • recognise contradictions
  • track evidence
  • estimate confidence
  • plan
  • reconsider failed approaches
  • decide what to retain
  • use tools correctly
  • learn from outcomes
  • produce the final answer

Much of this work does not inherently require a language model.

Speck starts from a different principle:

Anything that can be remembered, measured, ranked, checked, persisted, scheduled, validated, learned procedurally or calculated deterministically should not consume LLM intelligence twice.

The LLM is therefore treated as a specialised semantic processor inside a larger cognitive system rather than as the entire system.

This is particularly important for small local models.

A large model can often compensate for weak agent architecture with raw intelligence.

A 1B–7B model cannot.

Speck attempts to provide the missing machinery.

At a high level, Speck separates cognitive state from model inference.

                    USER / ENVIRONMENT
                           │
                           ▼
                        INTAKE
                           │
              ┌────────────┴────────────┐
              │                         │
              ▼                         ▼
           MEMORY                    EVIDENCE
         ACTIVATION                  / CLAIMS
              │                         │
              └────────────┬────────────┘
                           ▼
                     COMPETITION
                           │
                           ▼
                    WORKING CONTEXT
                           │
             ┌─────────────┼─────────────┐
             │             │             │
             ▼             ▼             ▼
          PLANNER        WORKER     METACOGNITION
             │             │             │
             └─────────────┼─────────────┘
                           ▼
                       VALIDATION
                           │
                    ┌──────┴──────┐
                    ▼             ▼
                  REPLY          TOOLS
                    │             │
                    └──────┬──────┘
                           ▼
                  OUTCOME / EVIDENCE
                           │
                           ▼
               MEMORY + BELIEF UPDATE
                           │
                           └──────► next cycle

The model participates in cognition.

It does not own cognition.

Speck contains runtime mechanisms for several processes normally delegated to prompts or large models.

Memory is stored independently of the currently loaded model.

The runtime supports mechanisms including:

  • persistent task state
  • memory admission
  • embedding-based retrieval
  • activation
  • associative edges
  • memory graphs
  • grounding
  • domain anchors
  • contextual recall

Relevant memories can compete for inclusion in working context rather than simply being dumped into the model's context window.

Not everything remembered should occupy the model's attention.

Speck can rank and compete candidate information before constructing working context.

This allows limited model context to be spent on information that is more likely to matter to the current task.

Speck distinguishes between information being present and information being established.

The runtime can maintain:

  • claims
  • supporting evidence
  • conflicting evidence
  • confidence
  • provisional hypotheses
  • belief revision

New evidence can therefore change persistent state without requiring the model to reconstruct the entire reasoning history.

Candidate information can be checked against user statements and existing evidence before it becomes persistent knowledge.

Grounding can combine semantic similarity with deterministic signals rather than relying entirely on an LLM to decide what the user previously said.

Uncertain information can remain task-local until sufficiently supported.

Planning is separated from execution.

Different models can therefore be assigned to different cognitive roles.

For example:

Intake       → small fast model
Planner      → stronger reasoning model
Worker       → general-purpose model
Validator    → independent model or deterministic check
Embeddings   → dedicated embedding model

There is no requirement for every role to use the same model.

Speck maintains runtime information about its own performance rather than merely prompting a model to "reflect."

This includes mechanisms for:

  • assessment
  • calibration
  • competence tracking
  • prediction
  • outcome comparison

The objective is not introspection for its own sake.

It is to improve future decisions.

Repeated successful behaviour should not require rediscovery forever.

Speck can represent procedures independently of conversational memory, allowing successful patterns to become reusable runtime knowledge.

The long-term goal is simple:

Do not repeatedly spend inference on problems the system has already learned how to solve.

Speck is deliberately developed against small local models.

This exposes weaknesses that large frontier models can hide.

Small models may:

  • misunderstand complex prompts
  • lose state
  • mishandle negation
  • hallucinate structure
  • perform poor arithmetic
  • choose inappropriate tools
  • accept unsupported conclusions
  • prematurely ask the user for information
  • drift from supplied evidence

Speck does not assume these problems can all be solved with better prompting.

Where practical, responsibility is moved out of the model entirely.

LLM responsibility
        │
        ▼
Can software perform this reliably?
        │
   ┌────┴────┐
  YES        NO
   │          │
Runtime      Model
mechanism   judgement

Model outputs can then be treated as proposals rather than unquestioned state transitions.

One of Speck's fundamental design constraints is:

No model should be the identity of the agent.

Models are replaceable cognitive resources.

A task may begin using one worker model and continue using another.

A model can be:

  • unloaded
  • upgraded
  • downgraded
  • replaced
  • moved to another device
  • assigned a different cognitive role

without discarding the persistent state surrounding the task.

This makes heterogeneous local inference practical.

Speck includes a tool layer rather than allowing arbitrary model output to directly become action.

The runtime contains support for:

  • tool registration
  • tool intents
  • workspace operations
  • transactional workspace changes
  • browser automation
  • Genesis tools
  • validation before execution

Tool selection and tool execution can therefore be independently inspected and constrained.

Speck is built on the Genesis runtime.

Genesis provides the surrounding agent infrastructure, including plugin hosting, profiles, UI integration and runtime services.

Speck provides the cognitive layer.

Conceptually:

┌────────────────────────────────────────────┐
│                  SPECK                     │
│                                            │
│ Memory • Attention • Evidence • Beliefs   │
│ Planning • Metacognition • Procedures     │
│ Validation • Cognitive State              │
└─────────────────────┬──────────────────────┘
                      │
┌─────────────────────▼──────────────────────┐
│                 GENESIS                    │
│                                            │
│ Plugins • Tools • Profiles • Runtime      │
│ Services • Interfaces • Infrastructure    │
└────────────────────────────────────────────┘

Speck is intended to be experimentally testable.

The repository includes benchmarking infrastructure for:

  • role qualification
  • trajectories
  • tool behaviour
  • validation
  • comparison
  • ablation testing
  • cognitive metrics

This is important because adding more cognitive machinery does not automatically make an agent better.

A mechanism should be removable and testable.

If removing a component produces no measurable difference, its value should be questioned.

A conventional agent often resembles:

Prompt
  ↓
LLM
  ↓
Tool
  ↓
LLM
  ↓
Tool
  ↓
LLM
  ↓
Answer

Speck is closer to:

Environment
     ↓
Persistent Cognitive State
     ↓
Evidence + Memory + Attention
     ↓
Working Context
     ↓
Specialised Model Judgement
     ↓
Independent Validation
     ↓
Action
     ↓
Measured Outcome
     ↓
Belief / Memory / Procedure Update
     ↓
Next Cognitive Cycle

The distinction is intentional.

Speck is not primarily attempting to build a better prompt loop.

It is attempting to build the machinery surrounding the model.

Speck is designed to work with locally hosted models and services.

A default installation:

  • binds to loopback
  • stores runtime data locally
  • does not require a cloud model
  • does not enable model, embedding or browser services unless configured
  • keeps credentials and personal runtime data outside the repository

Cloud models can still be used when desired.

They are resources available to the cognitive system rather than a requirement for the architecture.

  • Node.js 18 or newer

Clone the repository:

git clone https://github.com/doctarock/Speck.git
cd Speck

Install dependencies:

npm install

Start Speck:

npm start

Then open:

http://127.0.0.1:4310

Fresh runtime data will be created under:

data/

Model, embedding and browser services remain disabled until explicitly configured.

Build:

npm run build

Type-check:

npm run check

Run the test suite:

npm test

Run the benchmark suite:

npm run benchmark:suite

Speck is an active experimental project.

The architecture is functional, but it should not yet be interpreted as a claim that every cognitive mechanism improves every task or every model.

The project deliberately includes benchmarks, probes and ablation infrastructure so those claims can be tested rather than assumed.

Some mechanisms will work.

Some will need refinement.

Some may eventually be removed.

That is part of the experiment.

Speck incorporates ideas explored in Artificial Cognitive Architecture Omega Gen2, particularly around persistent cognition, memory, attention, competition, evidence and cognitive state.

The objectives are different.

Omega asks:

What happens if we attempt to construct increasingly mind-like persistent artificial cognition?

Speck asks:

How much model intelligence can be replaced or amplified by persistent computational cognitive machinery?

Omega explores artificial cognition.

Speck attempts to make that cognition useful as infrastructure.

The model is disposable. The cognitive state is not. 2. Do not use inference for deterministic work. 3. Models make semantic judgements; software enforces policy. 4. Model output is evidence, not automatically truth. 5. Memory should compete for attention rather than flood context. 6. Successful reasoning should become reusable knowledge where possible. 7. Failures should modify future behaviour. 8. Different cognitive jobs may require different models. 9. Small models are a constraint, not an afterthought. 10. Cognitive mechanisms should be measurable and removable.

Speck ultimately tests a fairly simple hypothesis:

How much of what we currently call LLM intelligence actually needs to live inside the LLM?

Modern agents repeatedly ask enormous neural networks to remember, organise, reconsider, compare, track, schedule and validate information that conventional software can often handle more reliably.

Speck moves those responsibilities outward.

If successful, increasingly capable agents should become possible using smaller models, less inference, persistent knowledge and measurable cognitive machinery.

The goal is not to make a small model pretend to be a large model.

The goal is to give the small model a better brain around it.

Small model. Persistent mind.

Speck is inspired by mechanisms developed in Artificial Cognitive Architecture Omega Gen2, but it is not a rewrite or fork of Omega.

Omega is an experiment in artificial cognition and autonomous mind-like behaviour.

Speck has a different objective:

Use conventional software to provide memory, attention, state, learning, planning support, metacognition, evidence tracking and proceduralisation so that a small LLM only performs operations that genuinely require semantic intelligence.

The central design principle is:

Anything that can be remembered, measured, ranked, checked, persisted, scheduled, validated, learned procedurally or calculated deterministically should not consume LLM intelligence twice.

A second critical principle:

The model is disposable. The cognitive state is not.

A worker must survive un its model, changing models, restarting the process, or migrating to another compute device without losing its task state.

This distribution includes Speck's compiled server, browser interface, and the Genesis runtime modules it uses. It does not include optional plugins, credentials, or user data.

Requirements: Node.js 18 or newer.

npm install
npm start

Open http://127.0.0.1:4310. The server binds to loopback by default and creates fresh runtime data under data/. Model, embedding, and browser services are disabled unless explicitly configured. Keep credentials and personal data in local settings, never in this repository.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @speck 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-speck-cognit…] indexed:0 read:9min 2026-10-06 · —