{"slug": "qwen3-8-flash-passed-17-17-checks-on-a-deepswe-v1-1-task-in-our-agent-harness", "title": "Qwen3.8-Flash passed 17/17 checks on a DeepSWE v1.1 task in our agent harness", "summary": "Qwen3.8-Flash, running inside the closed-source host-enforced AIC software engineering runtime, passed all 17/17 canonical checks on the DeepSWE v1.1 updo-policy-alerting task at an operator-reported model/API cost of $0.31 and an observed time of 56m 55s, according to an evidence bundle published in the AIC-Evals repository. The same-task comparison shows GPT-6 Astra at 14/20 and GPT-5.6 Sol at 12/20 under the official mini-swe-agent harness, while Claude Opus 5 scored 8/20 and Gemini 3.8 Flash, Claude Fable 5, Claude Sonnet 5, and Claude Opus 4.8 each scored 0 passes. The repository states the figures are task-difficulty context only and not a controlled leaderboard, because the AIC run used a different harness and operator protocol and a single post-fix attempt cannot estimate a stable pass rate.", "body_md": "Verified benchmark artifacts, model patches, and integrity receipts for AIC, a closed-source, host-enforced AI software engineering runtime.\n\nThis repository publishes inspectable evidence from selected AIC evaluation runs. It does **not** contain the AIC product, source code, internal system prompts, private role artifacts, held-out tests, or credentials.\n\n**Qwen3.8-Flash + AIC passed all 17/17 canonical checks on the DeepSWE v1.1 `updo-policy-alerting` task for an operator-reported model/API cost of $0.31.**\n\nThe comparison above establishes task-difficulty context. AIC and the official model rows used different harnesses, operator protocols, and sampling designs, so it is not a controlled leaderboard.\n\n[Open the evidence bundle](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix) · [Inspect the frozen patch](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix/model.patch) · [Review official-result context](https://github.com/taiheqi718-art/AIC-Evals/blob/main/comparisons/deepswe-v1.1/updo-policy-alerting) · [Read rubric notes](https://github.com/taiheqi718-art/AIC-Evals/blob/main/comparisons/deepswe-v1.1/updo-policy-alerting/RUBRIC-NOTES.md) · [Read the FAQ](https://github.com/taiheqi718-art/AIC-Evals/blob/main/FAQ.md)\n\n| Benchmark | Task | Model | Attempt | F2P | P2P | Reward | Model cost | Observed time | Effective time | \n|---|---|---|---|---|---|---|---|---|---|\n| DeepSWE v1.1 | [`updo-policy-alerting`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix) | Qwen3.8-Flash | 1 (post-fix) | **17/17** | **123/123** | **1.0** | **$0.31** | **56m 55s** | **< 56m 55s** | \n\n**Timing note:** The observed AIC time includes six non-passing acceptance-probe executions—three `probe_error` and three `behavior_failed` outcomes—and the associated corrective turns. These included probe-construction errors and assertions superseded by corrected evidence, not canonical-verifier failures. The exact retry-adjusted duration cannot be isolated reliably, so effective time is reported only as a conservative upper bound: **< 56m 55s**.\n\nThe official DeepSWE v1.1 data includes repeated `mini-swe-agent` trials for this task. Most of the models below have four trials at each of five reasoning-effort levels:\n\n| Model / system | Harness | Same-task scored passes | \n|---|---|---|\n| **Qwen3.8-Flash + AIC** | AIC | **1/1** post-fix attempt | \n| GPT-6 Astra | Official `mini-swe-agent` | **14/20** | \n| GPT-5.6 Sol | Official `mini-swe-agent` | **12/20** | \n| Claude Opus 5 | Official `mini-swe-agent` | **8/20** | \n| Gemini 3.8 Flash | Official `mini-swe-agent` | **0/8** across two published effort levels | \n| Claude Fable 5 | Official `mini-swe-agent` | **0/20** | \n| Claude Opus 4.8 | Official `mini-swe-agent` | **0/19** scored; 1 provider error excluded | \n| Claude Sonnet 5 | Official `mini-swe-agent` | **0/20** | \n\nThese numbers provide task-difficulty context only. They are **not** an apples-to-apples ranking: the AIC run used a different harness and operator protocol, and one AIC attempt cannot estimate a stable pass rate. See the [effort-by-effort comparison, official source links, snapshot hashes, and derivation](https://github.com/taiheqi718-art/AIC-Evals/blob/main/comparisons/deepswe-v1.1/updo-policy-alerting).\n\nTwo verifier-selected event-ordering details are more specific than the public prose. They are documented neutrally in the task's [rubric notes](https://github.com/taiheqi718-art/AIC-Evals/blob/main/comparisons/deepswe-v1.1/updo-policy-alerting/RUBRIC-NOTES.md); the delivered candidate implements both and passes 17/17.\n\n- the exact public task instruction given to the runtime;\n- the frozen candidate patch produced by the run;\n- the verifier's aggregate `reward.json` ;\n- public-safe result metadata and integrity hashes;\n- a scorecard and its deterministic local renderer;\n- the upstream license applicable to the patched project.\n\nRaw test reports, test names, held-out test source, private rubrics, model transcripts, internal AIC role artifacts, and machine-local configuration are intentionally excluded.\n\nCost and timing are reported with an explicit scope. Model cost is operator-reported model/API spend, not total infrastructure cost. Observed AIC time runs from prompt acceptance to the `delivered` terminal state; end-to-end time additionally includes the post-delivery canonical verifier. Effective time excludes known non-passing probe and correction overhead, but is shown only as a strict upper bound because an exact counterfactual duration is not recoverable.\n\n| Layer | Public artifact | What it establishes | \n|---|---|---|\n| Task | [`task.md`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix/task.md) | Exact public instruction supplied to AIC | \n| Candidate | [`model.patch`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix/model.patch) | Frozen delivered implementation | \n| Result | [`reward.json`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix/reward.json) | Aggregate canonical verifier outcome | \n| Provenance | [`evidence.json`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix/evidence.json) | Candidate, verifier, cost, timing, and retained-evidence identities | \n| Integrity | [`manifest.json`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/runs/deepswe-v1.1/updo-policy-alerting/qwen3.8-flash/attempt-01-post-fix/manifest.json) | SHA-256 inventory of every published run file | \n| Context | [`official-results.json`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/comparisons/deepswe-v1.1/updo-policy-alerting/official-results.json) | Machine-readable same-task derivation and official source hashes | \n| Interpretation | [`RUBRIC-NOTES.md`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/comparisons/deepswe-v1.1/updo-policy-alerting/RUBRIC-NOTES.md) | Two verifier-selected behaviors more specific than the public prose | \n\nWhen the public AIC desktop and CLI clients are released, developers will be encouraged to evaluate models and tasks and submit public-safe evidence bundles through pull requests.\n\nMaintainer-published runs and community submissions will remain separate. Community results will disclose complete comparable attempt series and carry an evidence label—self-attested, artifact-checked, or, when supported by the clients, receipt-verified. A successful single attempt will not be presented as Pass@1 or a stable pass rate.\n\nThe intended bundle, privacy rules, attempt-disclosure policy, and review process are described in [CONTRIBUTING.md](https://github.com/taiheqi718-art/AIC-Evals/blob/main/CONTRIBUTING.md). The clients should generate these bundles automatically so contributors do not need to expose AIC internals or protected verifier material.\n\nA published run shows that the frozen patch identified by its SHA-256 digest produced the recorded verifier result under the stated environment. AIC itself is proprietary and is not distributed here, so this repository is an artifact record—not a fully reproducible copy of the orchestration system.\n\nAttempt numbering is scoped to a materially stable AIC baseline. The first published result is post-fix attempt 1; earlier internal development and recovery rounds used materially different harness revisions and are not counted in this series. This repository does not yet represent a complete attempt ledger or a full-benchmark score. These results are independent publications and are not official leaderboard submissions.\n\nSee [METHODOLOGY.md](https://github.com/taiheqi718-art/AIC-Evals/blob/main/METHODOLOGY.md) for the evidence protocol, [FAQ.md](https://github.com/taiheqi718-art/AIC-Evals/blob/main/FAQ.md) for common interpretation questions, and [DISCLAIMER.md](https://github.com/taiheqi718-art/AIC-Evals/blob/main/DISCLAIMER.md) for scope and limits.\n\nFor citation metadata, see [`CITATION.cff`](https://github.com/taiheqi718-art/AIC-Evals/blob/main/CITATION.cff). Cite the repository together with the immutable run directory used by your analysis.\n\n```\nruns/\n  <benchmark>/\n    <task>/\n      <model>/\n        attempt-<nn>/\n          README.md\n          task.md\n          model.patch\n          patch.meta.json\n          reward.json\n          evidence.json\n          manifest.json\n          scorecard.png\ncomparisons/\n  <benchmark>/\n    <task>/\n      README.md\n      official-results.json\n      RUBRIC-NOTES.md\nassets/\n  social-preview.png\n  social-preview.html\n  render-social-preview.ps1\nsubmissions/\n  <benchmark>/\n    <task>/\n      <model>/\n        <generated-run-id>/\n```\n\nOriginal documentation, metadata, and renderer code in this repository are licensed under the [MIT License](https://github.com/taiheqi718-art/AIC-Evals/blob/main/LICENSE). AIC itself is not included and is not licensed by this repository. Candidate patches remain subject to the upstream project's license, included with each run where applicable.", "url": "https://wpnews.pro/news/qwen3-8-flash-passed-17-17-checks-on-a-deepswe-v1-1-task-in-our-agent-harness", "canonical_source": "https://github.com/taiheqi718-art/AIC-Evals", "published_at": "2026-09-16 11:50:34+00:00", "updated_at": "2026-09-16 12:13:03.460010+00:00", "lang": "en", "topics": ["ai-products", "ai-agents", "ai-tools", "developer-tools", "ai-research"], "entities": ["Qwen3.8-Flash", "AIC", "DeepSWE v1.1", "AIC-Evals", "GPT-6 Astra", "GPT-5.6 Sol", "Claude Opus 5", "mini-swe-agent"], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-flash-passed-17-17-checks-on-a-deepswe-v1-1-task-in-our-agent-harness", "markdown": "https://wpnews.pro/news/qwen3-8-flash-passed-17-17-checks-on-a-deepswe-v1-1-task-in-our-agent-harness.md", "text": "https://wpnews.pro/news/qwen3-8-flash-passed-17-17-checks-on-a-deepswe-v1-1-task-in-our-agent-harness.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-flash-passed-17-17-checks-on-a-deepswe-v1-1-task-in-our-agent-harness.jsonld"}}