cd /news/artificial-intelligence/new-results-in-square-packing-proble… · home topics artificial-intelligence article
[ARTICLE · art-124669] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

New results in square packing problem by a hobbyist

A hobbyist researcher using AI agents has improved the lower bound for the square packing problem s(11) to 381/100 = 3.81, surpassing Walter Stromquist's 1984 bound of 3.7888543, marking the first recorded improvement for this open case. The results, published in a GitHub repository, also provide new bounds for n = 12, 20, and 21, and revise values for n = 17 through 21, with seven lower bounds proved by the automated workflow.

read27 min views3 publishedSep 9, 2026
New results in square packing problem by a hobbyist
Image: Michielbdejong (auto-discovered)

This repository contains:

  • New results. The lower bound ons(11) has moved. It improves Stromquist’s3.7888543… bound, stated in1984, Memo III, p. 10 and published in 2003; no intervening improvement was found by the recorded search. With it come the first bounds located in the public record for twelve, twenty and twenty-one squares, and values fromn = 17 throughn = 21 that displace what was in print.
  • A survey of the whole problem. Every casen = 1…100 , the primary literature retained and transcribed, and the bound a sourcereports kept apart from the bound this repository hasverified . Seven of the lower bounds it shows were proved here.
  • An automated research workflow. The results and the survey are produced and checked by AI agents running a recorded process: hypotheses registered before measurement, every claim graded, every defect logged.

The explainer page is the best introduction to the proof: the s(11) bound and its five conditions in one page, with every figure drawn from the certificate it explains.

The retained n = 1…100 atlas, with each packing normalized to its own container and labeled by its best-known side upper bound. For open cases, the strongest lower bound independently verified here appears beneath it. A crimson star marks a lower bound proved here. The image is available in SVG, PDF, and high-resolution PNG.

The register now runs to n = 324, the end of the catalogue’s audited range, and a second, poster-sized composite draws all of it: known-best-1-324, an 18-by-18 grid with the same cards, badges and legend, available as SVG and PDF (44 by 51 inches). The first figure is unchanged; the atlas README describes both.

s(n) is the side of the smallest square that holds n non-overlapping unit squares. The problem is elementary to state and remains open even at small n.

New Results · Survey · Repository Guide · Getting Started · Reports · Autonomous Research Process · Conventions · Layout

The results register collects first-party and load-bearing whole results. Each result has a T-NNN ID and the classifications defined in epistemics.md: V, the highest verification rung supported by its cited evidence, and C, what this repository has recorded or performed itself. The gate checks the structural support for both classifications. apparently-novel means a recorded source search did not find the named contribution; it is not a claim of priority.

Each result also carries S, a significance score from 1 to 5 against the same file’s rubric. The two groups below are split on it rather than on taste: S4 is its anchor for a reusable technique, bound family or resolved disputed value, and S5 for movement on a central open case.

Results first established here, as far as the recorded source searches show:

  • T-018: s(11) ≥ 381/100 = 3.81, improving the lower bound for the smallest open case (S5). s(11) is the case this project exists for, and the recorded public search found no stronger lower bound after Stromquist stated2 + 4/√5 = 3.788854… in 1984 and published it in 2003. Thememo review distinguishes that early statement from the later published presentation. A first-partyweighted fractional unavoidable-set certificate —1121 weighted atoms, total mass434547/40000 , every placement of a shrunken square covering mass at least1 —proves that eleven unit squares do not fit in a container of side381/100 = 3.81 . This narrows the interval from0.088230 to about0.067084 ; the gap is not closed. Two certificate rungs are retained below381/100 :19/5 , the value that first passed Stromquist, and189/50 , the calibration rung below him that was run first on purpose and proves nothing new. ScoredS5 , the rubric’s anchor for movement on a central open case. The shortest complete statement of the proof, with the certificate’s hash and the one command that checks it from the standard library alone, is theproof card . Aself-contained package for third-party checking ships with it, so the19/5 rung can be decided without trusting anything else here. A small refinement, recorded in theT-022 proof packet , givess(11) ≥ 3.810025723614703… as a weak limit bound; it does not decide fit at that endpoint. The certificate at3.81 supplies the proof explained here.
  • T-019: s(17), s(18), s(19) ≥ 459/100, improving the register ( S4). The adopted bound forthese three cases was Massaccesi’s4.5058 , taken from a source rather than proved here. The same generator returns4.59 , on 1184 atoms with total mass423327/25000 = 16.9331 againstn = 17 and least covered mass200009/200000 , so the repository now carries a first-party certificate0.0842 above the number it had adopted, with the229/50 and451/100 rungs it climbed through retained below. A stronger public candidate at9141/2000 = 4.5705 , posted to GitHub on 16 August 2026 and neither peer reviewed nor replayed here, was outside the search corpus when this was registered; against it the movement is0.0195 atn = 17,18 . The DS7 audit now records a stronger reportedn = 19 bound, approximately4.6172815 , with an unresolved source caveat; T-019 does not improve that report. The bounds that came closest, each archived here with its source and replayed where the tools allowed: anabologyco-maker’s4.5705 and4.57 (16 and 13 August), Massaccesi’s4.5058 (21 August), Burns’s4.4811 (6 August), Mira’s4.468292 and4.450837 (11 and 10 August), Fort’s4.456575 (11 August), and Brandwijk’s89/20 = 4.45 (18 July); Mira’s, Fort’s and Brandwijk’s are exact sixteen-point certificates and the rest weighted ones, and every 2026 author but Brandwijk discloses a model-written proof. One certificate covers all three sizes without a monotonicity step: onlyCondition 2 mentionsn , so an atom set certifies its side for every integer above its own mass.T-020 has since carriedn = 19 past it;n = 17 andn = 18 are this result’s alone, being too small for the heavier atom set that moved the other three.
  • T-020: s(19), s(20), s(21) ≥ 24/5, improving the verified register ( S4). The verified fields for twenty and twenty-one squares previously carried Nagamochi’s 2005 general formula,1 + √13 = 4.6055… and1 + √14 = 4.7416… . TheDS7 source audit now records stronger external reports separately, with missing proofs and source caveats explicit. Acertificate at 4.80 —2260 atoms, total mass946131/50000 , least covered mass50007/50000 —movesn = 20 by 0.194449 ,n = 21 by0.058343 , andn = 19 by0.21 , the largest single-case movement in the register. The three sizes again come out ofCondition 2 alone. From this4.80 rung the method had0.1885 of room atn = 20 andn = 21 beforeits own ceiling , and0.0856 atn = 19 before it would contradict the best-known packing.T-021 has since raised then = 20 andn = 21 bounds to97/20 , leaving0.1385 of room there; this24/5 rung remains current forn = 19 .
  • T-017: s(12) ≥ 99/25, from nothing case-specific at all ( S4). n = 12 had only then = 11 bound inherited by monotonicity; the frontier record said in as many words that nothing specific ton = 12 had ever been proved. An eight-rung ladder—19/5 ,77/20 ,97/25 ,39/10 ,393/100 ,197/50 ,79/20 ,99/25 —is retained, all from one generator that applies at everyn , which is why this is scoredS4 as a bound family rather than a case result. At99/25 = 3.96 it also separates the cases:s(12) > s(11) , since Trump’s 1979 packing putss(11) ≤ 3.877084 . That did not follow from anything on record before. The case is now0.04 from its conjectured optimum of4 . On the retained 181-direction net, the proved ceiling for twelve squares is approximately3.990816 ; refining the net can raise that ceiling. Every finite net still has a ceiling strictly below4 , so no single certificate of this shape can close the case. A family of certificates approaching4 is not ruled out; whether one exists is a question about the covering value.
  • T-010: s(11) ≥ 2 + 4/√5, repaired ( S4). The printed 2003 Figure 14 unavoidability claim has a strict counterexample, so the literature’s standings(11) bound rested on a broken step. Thecase report walks through what survived it. A preregistered, source-distinct replacement point set restores the full lower-bound argument and certifies exactly.T-018 has since passed the repaired value, but the repair is what made it a value worth passing.

These checked results have narrower scope: a single case, a catalogue refinement, or an erratum.

T-021: s(20), s(21) ≥ 97/20 (S3). Acertificate at 4.85 has total mass19848723/1000000 = 19.848723 , so the same exact object proves both cases without a monotonicity step. It raises each bound by0.05 aboveT-020 ; the heavier atom set does not apply ton = 19 . #

T-001 / T-002: s(17) ≥ 4.426213 and s(18) ≥ 4.426213. A sixteen-point unavoidable set is certified by exact rational cover verification and an independent interval branch-and-bound over the full pose space. Both are superseded as the verified lower bound: first by the source-backed4.5058 adopted on 2026-09-03, and now byT-019 , which proves more than either. #

T-009: s(29) ≤ 5.93383346267692918974379895098. A Krawczyk interval certificate encloses a unique exact solution around a rational witness. #

T-012 / T-013: exact rigidity determinations. The retainedn = 5 optimum is not infinitesimally rigid but is second-order rigid. The retainedn = 40 packing is infinitesimally flexible, with every recorded first-order flex refused at second order. Both refine catalogue annotations that say only “Rigid.” #

T-014: Goebel’s n = 5 optimum is locally rigid at fixed side. At the exact side2 + √2/2 the labeled pose is an isolated point of the feasible set: no nonconstant continuous feasible path leaves it and no sequence of distinct feasible poses converges to it. Proved exactly overQ(√2) from a complete accounting of all 400 local inequalities, by curve selection and an order-2m coefficient induction, and independently reviewed. The side is fixed throughout; nothing is claimed about an isolation radius, about any othern = 5 optimum, or about global uniqueness, and nothing follows for the side as a variable: with the side free the obstruction fails, which X-007 measured. #

T-005: an erratum in Bentz 2010. Lemma 10’s middle replacement point is transposed in print. An exact escape certificate refutes the printed point, and the corrected reading certifies exactly against the journal page image.

In each case, the theorem belongs to the source; this repository adds an exact machine check.

  • T-004 / T-008: Bentz 2010, Theorem 8, including both halves ofs(46) = 7 .
  • T-011: exact verification of Trump’s 1979n = 11 record witness over its degree-eight field, including the zero-gap contacts that finite precision cannot certify.

The complete statements, scopes, evidence, limitations, classifications, and next actions live in the register. Results that still rest on a source read rather than a machine check are labeled there accordingly.

The survey records the best-known packing and strongest lower bound independently verified here for every n ≤ 324, with provenance and separate reported and verified fields. Its source is one schema-validated case file under packing/frontier/; the generated status table is the reader view, and the atlas above renders every retained known-best packing.

The literature archive retains each primary source, a cleaned Markdown transcription, and the unedited extraction used to check it. The generated evidence inventory shows what each recorded claim rests on, who performed the work, and how far it has been checked.

The survey audits rather than merely transcribes. For example, the earliest published proof of s(7) = 3 carries four recorded defects in its printed route, so the case’s proved status rests on independent later proofs. The n = 7 case states that disposition and links the relevant source audit.

Where What
Tutorial First-principles introduction to the objects, bounds, cells, stationary branches, search, and proof obligations
Synopsis Current technical state, established results, terminology, experiment roll-up, and handoff
Results register Whole-result bounds, audits, structural theorems, and errata graded under epistemics.md
Frontier One record per case for n = 1…324 , with reported and verified bounds kept separate
Atlas Known-best and prospective packings, contact-scaffold enumeration, and deterministic renderings
Literature Retained primary sources, cleaned transcriptions, and raw extractions
Reports Research reports on the mathematics, algorithms, infrastructure, formal proof, and search strategy
Code and development guide Exact verification, search, promotion, and the validation tiers and behavioral lanes that gate every change
Campaign record Hypotheses, preregistered experiments, session records, agendas, and generated ledger
Defect log Generated record of defects, detection methods, fixes, and regressions

Long-lived tests and runs retain detailed timing evidence under OR-14. The validation efficiency and checkpoints plan tracks improvements to everyday feedback and full final checkpoints, with measurements and preserved coverage required before accepting a speedup.

SYNOPSIS.md is the technical root and current-state document. The generated day-to-day views are the frontier status table, results register, campaign ledger, and agenda map. To resume work, use the synopsis’s current handoff, which names the owning work item and next bounded slice.

Read TUTORIAL.md once for the mathematical orientation, then SYNOPSIS.md for current results and open work. Run commands from packing/; the project uses Python 3.14 through uv.

These are the terms a reader encounters most often. The synopsis terminology gives the full definitions.

Term Meaning
configuration A placement of all n squares plus the container side:3n + 1 coordinates
cell A separating axis and order for every pair of squares; with angles fixed, one cell is one linear program
quench Deterministic refinement from a configuration to a local optimum
basin The preimage of one returned pose under a fixed deterministic quench; one connected terminal component may contain several point-basins
polish Refinement within the current basin
exploration Work intended to reach a different basin; the term implies no assurance level
standing best The best published side for that n , hence an upper bound rather than known optimality in open cases
gap best_side − standing_best , always signed
assurance reported ,numerically-checked , orverified ; method, arithmetic, origin, limitations, and novelty are recorded separately

One ID names one durable thing, and IDs are not reused. The prefix identifies the record’s layer; conventions.md is the definitive registry.

ID Names
n-NNN One frontier case, such as n-011
T-NNN One whole result in the results register; the synopsis also has older local T-N shorthand
X-NNN One exploration report from which hypotheses may be derived
H-NNN One falsifiable hypothesis or open question
exp-NNN One durable experiment record; a lower-level run is one command invocation or seed trial
series-NNN One campaign-wide tooling and comparability regime
agenda-NNN One ordered queue of bounded commitments
BC-NNN One bounded commitment in an agenda; other agendas may declare another two-letter prefix
session-NNN One escalated agent-session record containing ordered workflow phases
D-NNN One defect and its detection, consequence, fix, and regression
think-xxxx One git-native tbd bead: durable work and dependency state
W1W10 A workflow entry point, not a durable artifact ID

Other rules needed to read the repository:

  • Structured values live in YAML or frontmatter; prose supplies explanation and judgment. A consumer does not scrape prose for fields.
  • Declared paths are repository-relative. Generated views are regenerated from their source records and are not edited by hand.
  • Evidence assurance, method, origin, precision, limitations, and novelty are separate facts. Whole-result V/C classifications do not replace evidence-level fields.
  • Source-faithful archive material is not cleaned up as project prose. Reconstructed source text is marked and counted.
  • Corrections preserve the original record and add a dated statement of what remains valid. IDs and scientific outcomes are not silently rewritten.

The project keeps numerical exploration, symbolic reconstruction, exact verification, and research records as separate layers.

Layer Tools Role here
Work and issue state tbd Git-native beads, dependencies, specs, guidelines, and handoffs
Structured research records softschema ,PyYAML ,Python jsonschema , and jsonschema-rs Mixed prose-and-data artifacts, JSON Schema contracts, in-process checks, and fast repository-wide validation
Documentation Flowmark andPractical Prose Semantic Markdown formatting and the common documentation guidelines
High-precision numerics mpmath Arbitrary-precision refinement, interval endpoints, and decimal-to-exact promotion
Arrays and optimization NumPy andSciPy Geometry arrays, nonlinear refinement, and fixed-cell linear programs
Symbolic mathematics SymPy Contact-system assembly, elimination probes, minimal-polynomial recovery, and independent symbolic checks
Exact mathematics sqpack.field ,sqpack.verify , and the case-specific certifiers Rational and algebraic sign decisions, unavoidable-set certificates, Krawczyk enclosures, and proof replay
Parallel search The local sqsearch crate,Rayon , and theRust toolchain Multicore f64 screening and annealing; formal promotion remains on the Python side
Python environment and QA uv ,Ruff ,BasedPyright , andpytest Locked environments, linting, formatting, type checking, and behavioral tests
Git hooks lefthook Runs the pinned Markdown formatter and re-stages its changes before commit

The dependency and tool versions are owned by packing/pyproject.toml, packing/uv.lock, packing/sqsearch/Cargo.toml, and the root Makefile and hook configuration. development.md explains how the layers interact.

uv sync --frozen --all-extras --group dev
uv run --frozen packing-witness inspect witnesses/schadt-n029-2025-decimal.yaml
uv run --frozen packing-witness check witnesses/schadt-n029-2025-decimal.yaml \
  --method numerical-multiprecision --precision 300 --tolerance 1e-100
uv run --frozen packing-witness verify witnesses/schadt-n029-2025-rational.yaml
uv run --frozen python -m cases.trump11.verify_exact
uv run --frozen --all-extras --group dev packing-validate --edit

--edit is the smallest of five validation tiers. Which steps each tier runs, what it costs, and which of the three behavioral lanes a test lands in are tabulated in development.md → Validation Loops; the ceilings themselves are data the gate reads, in packing/devtools/gate-budgets.yaml. In short: a contributor runs --edit while editing and --push before pushing, every pull request runs --fast, and the complete gate runs on main and at the end of a research block.

Witness/v2 is the interchange format for supported rational, algebraic, and decimal witnesses. Exact verification covers rational witnesses and algebraic witnesses whose field preconditions the tool can certify. Recovering exact geometry from arbitrary decimal input remains the hard step; development.md and the module docstrings under packing/src/sqpack/ define the supported APIs and limits.

sqpack.render creates deterministic, self-contained SVG figures while preserving the input’s evidence tier in captions and metadata. The rendering guide owns the CLI, gallery, contact annotations, portability contract, and Motion Lab. The Motion Lab is an exploratory instrument, not a citable research result.

These ten research reports are the durable topical syntheses:

Report Scope
Packing 11 Unit Squares in a Square What is proved for s(11) , what remains conjectural, and why the available proof techniques do not close the gap
Algorithms and Tooling for Square Packing Search, numerical-to-exact promotion, verification, and the record landscape
FrankenSim as a Rust Toolkit for Square Packing Assessment of certified-arithmetic and determinism components in a larger Rust framework
Infrastructure for Square-Packing Exploration Build order, latency tiers, language boundaries, and symbolic tooling
Lean for Square-Packing Proofs and Validation Where proof assistants fit and which certificate layers are suitable first targets
A Search Philosophy for Square Packing Basin cartography, structural diversity, relaxation ladders, and search strategy
Public Sources Beyond n = 100 Which catalogues carry geometry above 100, their reuse terms, and why 324 is a source boundary
Stromquist’s 1984 Memos and Systematic Dots Proofs Historical corrections, the three memo arguments, and a reusable conditional counting control
Stromquist’s Twenty-Six-Square Packing Exact verification, comparison with the current record, source attribution, and bounded follow-up
The Best-Known n = 26 Packing Dated literature and source search, exact score normalization, and the limits of the best-known claim

The reports distinguish formal proof, finite numerical checks, and source reports. The document map identifies every maintained guide, dated record, generated view, and superseded document.

The repository supports autonomous research without making process a substitute for evidence. This section gives the operating model at a glance. The operating rules, workflow contracts, and campaign runbook own the full rules.

Evidence uses three assurance labels:

  • reported for a named source claim not checked here;
  • numerically-checked for finite-precision calculations with their precision, rounding, and tolerance recorded; and
  • verified for an exact check, rigorous interval certificate, or complete proof that covers the claim and its preconditions.

Whole results use the separate V/C classifications in epistemics.md. A verified feasible witness proves an upper bound; it does not prove global optimality without a matching verified lower bound.

Finite precision is not enough for a packing with exact contacts. Floating-point arithmetic can establish a strict positive gap, but a tolerance that accepts a true zero-gap contact also accepts a smaller overlap. Exact algebraic signs or outward-rounded intervals are therefore required before a contact-heavy witness becomes formally verified. The synopsis explains the full argument in Why Exactness Is Not Optional.

Two retained examples show the boundary. The Schadt n = 29 decimal pose passes its declared 300-digit numerical check, while the separately promoted interval witness establishes a slightly weaker side rigorously. Trump’s n = 11 witness is verified exactly over a degree-eight number field, including fourteen zero-gap contacts. The per-case records (n = 29, n = 11) state exactly which bound each artifact proves.

Verification answers whether a proposed packing is valid. Proving it optimal is a different problem and requires a matching lower bound. The synopsis’s capability ladder distinguishes what is built, what is ordinary engineering, and what remains mathematically contingent.

Principle Focus Goal
Correctness Soundness Formal validation that third parties can inspect, plus cross-validation of claims and source summaries
Process Discipline The minimum effective structure that keeps consequential decisions, evidence, and handoffs reconstructible
Insight Creativity Freedom to understand the problem, form varied hypotheses, and use all available information and tools
Efficiency Infrastructure Faster iteration through measured improvement of algorithms, systems, tools, and research surfaces

Correctness is the veto: no result advances beyond its evidence, however costly the required check may be. Process is proportional infrastructure, not a second mathematical standard; missing evidence can block promotion, while a preferred form or checkpoint cannot block useful work merely because it looks more disciplined. Insight remains free to propose. Efficiency may simplify process but cannot lower the assurance bar.

The system separates the kind of effort, the lens used to judge it, and the bounded action being executed:

Layer Question Recorded as
Operating principle / focus What quality dimension is preeminent for this phase? correctness ,process ,insight , orefficiency
Workflow What durable result is this phase meant to produce? One of W1–W10, or the narrow maintenance fallback
Slice What bounded action is being performed now, and how will it be checked? Objective, intended artifact, focused validation, and stop condition

Focus and workflow are independent. A W6 experiment may emphasize correctness, insight, or efficiency without changing its promise to execute a preregistered measurement; an efficiency-focused phase does not become W5 unless its durable result is a measured performance decision. A slice is smaller than either: it is one action inside the declared phase.

The durable work objects also have different lifetimes:

Unit Lifetime and role
Packing exploration The self-contained repository: sources, research, code, records, and tools
Campaign The multi-session research program and its shared record contract
Series A campaign-wide tooling regime and comparability boundary
Bead ( think-xxxx ) A durable work item and dependency node, open until the work is settled
Bounded commitment ( BC-NNN ) A planned attempt with entry conditions, acceptable exits, owner, and budget
Agent session An escalated interval of coordinated work containing one or more workflow phases
Workflow phase One declared purpose and focus within a session
Slice One bounded, immediately checkable action within a phase
Exploration / hypothesis A recorded source of ideas / one falsifiable claim with its criterion fixed before measurement
Experiment / run One durable measured round / one lower-level invocation or seed trial
Result / ledger One typed observation or whole-result claim / a generated view over source records

A bead says what needs doing. A bounded commitment says what would count as settling one attempt. A workflow phase says what kind of move is being executed now. One bead may require several commitments, one commitment may span several phases, and one phase may produce zero or several scientific records. The work-unit definitions, campaign runbook, and agent-session guide own the exact contracts.

Choose the workflow whose durable result matches the task. The synopsis owns the complete entry, exit, and transition contracts.

ID Workflow Enter when Durable result Usual handoff
W1 research-survey The sourced state of knowledge is incomplete A sourced survey, source notes, conflicts, and explicit gaps W2
W2 factual-review Existing claims need a correctness-only audit Findings, authorized bounded corrections, or defects; no new theory smuggled into the review W3 or W4
W3 insight-iteration Current evidence needs new explanations or hypotheses Candidate X-NNN /H-NNN items with mechanisms, falsifiers, and information value W6
W4 process-review Work is hard to reconstruct or the discipline itself needs review Process findings, beads, and narrowly scoped contract or check changes W5 or the next owning workflow
W5 efficiency-loop A measured bottleneck limits useful iterations A baseline, profile, equivalence-safe change, and measured decision W6
W6 research-loop A registered hypothesis has a fixed criterion, regime, budget, and instrument contract A frozen instrument and one or more exp-NNN records, raw evidence, verdicts, and a current ledger W2 for promoted or high-risk claims; otherwise W3 or another W6 slice
W7 pipeline-improvement A named packing-pipeline surface or research consumer needs a new, stronger, simpler, or repaired capability A bounded implementation or refactor, executable controls, explicit evidence limits, cost receipt, and readiness decision; no scientific verdict W2 before a materially changed trust boundary reaches W6; otherwise W5 or W6
W8 documentation-pass A period of research has left the reader-facing documents behind what the record now says Reconciled root documents—README, tutorial, synopsis—checked against the artifacts and against each other, with every drift either fixed or logged as a defect; no new claim introduced W2 for any claim the pass could not verify; otherwise the next owning workflow
W9 remediation Confirmed defects or issue backlogs need a systematic repair wave Risk-ranked dispositions, bounded repairs, regression checks, updated defect records, and rerouted blockers; no scientific verdict W10
W10 review-planning-oversight An agenda or consequential session has ended and its results must change the plan Result and stop-reason classifications, actionable dispositions, reader-document review, a reprioritized candidate set, and one selected next entry The selected workflow; W9 or W8 when remediation or documentation work wins

Use general-improvement only for repository maintenance that fits none of W1–W10. Routine work records a workflow, bounded objective, intended artifact, and focused check. Use a versioned agent-session record only when work crosses multiple workflow phases, coordinates independent delegates, or needs durable recovery state.

defects.md is generated from packing/defects.yaml. It records every known defect in this toolchain, what caught it, the consequence, the correction, and the regression that now guards it. Two lessons govern review:

  • Results that look unusually good receive the strongest challenge because many soundness defects have pointed in that direction.
  • The automated gate checks only rules someone encoded. No soundness defect in the log was caught by it.

Current counts and detector statistics belong only in the generated defect log and the synopsis defect section. Corrections follow conventions.md §7: preserve the original record, add a dated correction that states what remains valid, and route any changed conclusion to the artifact that owns it.

W6 is the measured experiment loop rather than an umbrella for every session:

W3 insight iteration → registered hypothesis → W6 measured round → evidence and verdict
          ↑                                                        │
          └──────── successor questions ← W2 factual review ←──────┘

The hypothesis, criterion, regime, budget, and stop rule are fixed before measurement. The round records every outcome and stops at the criterion or clock. Promoted, novel, disputed, or otherwise high-risk claims receive an independent W2 pass before they move forward; routine rounds whose recorded guards already decide the criterion may return directly to W3 or another W6 slice.

The tbd queue owns durable work and dependencies. Campaign agendas order bounded commitments; hypothesis and experiment records own scientific claims and measurements; commits own code; escalated agent-session records own phase and recovery state. The key record IDs are X-NNN for explorations, H-NNN for hypotheses, exp-NNN for experiments, BC-NNN for bounded commitments, T-NNN for registered results, and D-NNN for defects. conventions.md owns the complete ID registry.

The campaign’s bounded research cycle defines clocks, result routing, budgets, and stop rules. Changing agents changes the driver, not the record or the evidence required for a claim.

Document Definitive responsibility
This README High-level orientation and the relationship among the layers
SYNOPSIS.md Current technical state, full workflow contracts, work-unit vocabulary, and handoff
epistemics.md Whole-result V/C/S/N classifications and their executable boundary
conventions.md IDs, filenames, artifact shape, evidence fields, provenance, and corrections
operating-rules.md How sessions choose, divide, validate, and hand off work
Campaign runbook Hypothesis and experiment mechanics, clocks, budgets, verdicts, and routing
W9 remediation pass Systematic defect and issue-backlog triage, repair waves, and terminal dispositions
W10 review, planning, and oversight Post-agenda result classification, document review, reprioritization, and next-entry selection
Agent-session guide Escalation threshold, workflow phases, recovery state, and session closeout
Agendas Mutable ordering and readiness of bounded commitments
development.md Engineering boundaries, commands, tests, and validation tiers

conventions.md owns identifiers, filenames, artifact discipline, evidence fields, provenance, corrections, and the boundary between machine checks and review. epistemics.md owns whole-result classifications. operating-rules.md owns how sessions are conducted, and development.md owns the engineering and validation workflow.

.
├── TUTORIAL.md             First-principles orientation for a newcomer
├── SYNOPSIS.md             Current technical state, results, terminology, and handoff
├── conventions.md          Artifact, identifier, evidence, and correction rules
├── epistemics.md           Whole-result verification and confirmation rubric
├── operating-rules.md      Session conduct and workflow rules
├── development.md          Python setup, engineering boundaries, and validation
├── defects.md              Generated view of packing/defects.yaml
├── docs/project/           Reports, reviews, specs, postmortems, and dated handoffs
├── docs/project/research/  The research reports listed above
├── packing/                Code, data, and the research record
│   ├── campaign/           Hypotheses, experiments, sessions, agendas, and ledger
│   ├── frontier/           Per-case claims, evidence, generated views, and results
│   ├── witnesses/          Witness/v2 interchange and retained examples
│   ├── golden/             Calibration endpoint snapshots
│   ├── atlas/              Known-best, prospective, enumerated, and rendering artifacts
│   ├── resources/          Retained literature and source-faithful transcriptions
│   ├── src/                Maintained sqpack package
│   ├── cases/              Case- and theorem-specific retained code
│   ├── devtools/           Checkers, adapters, generators, and mutation controls
│   ├── benchmarks/         Explicit performance probes
│   ├── tests/              Behavior, command, and architecture contracts
│   ├── sqsearch/           Rust screening annealer
│   ├── defects.yaml        Structured defect log
│   ├── defects.schema.yaml Defect-log contract
│   └── frankensim-probe/   Focused experiments against FrankenSim
├── vendor/kpress/          Vendored kpress submodule: the page's rendering layer
├── AGENTS.md               Project instructions for agents
├── CLAUDE.md               Bridge to AGENTS.md
├── Makefile                Markdown formatting, hooks, and skill mirroring
├── lefthook.yml            Pre-commit Markdown formatter hook
├── package.json            Tooling-only lefthook package
└── package-lock.json       Tooling lockfile

An optional, Git-ignored attic/ holds intake and scratch files. Sources used by durable research are retained under packing/resources/.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @walter stromquist 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/new-results-in-squar…] indexed:0 read:27min 2026-09-09 ·