cd /news/artificial-intelligence/akka-tests-spec-driven-ai-delivery-a… · home › topics › artificial-intelligence › article
[ARTICLE · art-145417] src=infoq.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Akka Tests Spec-Driven AI Delivery Across 65 Open Source Projects

Akka ported 65 open source projects using a spec-driven AI workflow, spending 99.3 hours and 9.41 billion tokens on the initial tranche and reporting a lines-of-code or performance improvement in 57 of the 65 ports. In the experiment, Anthropic's Sonnet averaged 61 minutes per port versus 120 minutes for Opus, while Opus used about 40% fewer tokens, and Akka found that structured specifications with claims, evidence, and typed behavior improved first-pass implementations. Akka CEO Tyler Jewell's LinkedIn post on the results drew engineer commentary questioning whether the lines-of-code reductions came from dead code removal or target-language differences and whether model capability or specification constraints drove porting efficiency.

by read2 min views4 publishedOct 5, 2026
Akka Tests Spec-Driven AI Delivery Across 65 Open Source Projects
Image: source

Akka used 65 open source projects to test a spec driven workflow for AI assisted software porting, measuring specification structure, context, model and effort selection, automated validation, token consumption, and runtime performance. The initial tranche took 99.3 hours and consumed 9.41 billion tokens, with Akka reporting a lines of code or performance improvement in 57 of the 65 ports.

The experiment used two tranches. Akka analyzed all 65 projects, generating specifications and implementing up to 10% of each project's surface area, then selected 10 for complete implementation based on system characteristics and measurable results. The delivery harness cycled through discovery, specification, porting, benchmarking, and improvement. Discovery analyzed code, models, schemas, and runtime behavior, while Claude with Akka Specify handled implementation, testing, and review. A common benchmark runner compared tests, code size, and latency.

Akka delivery harness workflow(Source: Akka Blog Post) Akka found that structured specifications with claims, evidence, and typed behavior improved first-pass implementations, while gaps in context files remained around cross-component decisions. Follow-up areas include interface enumeration, test ingestion, provenance tracking, differential testing, and adversarial testing.GitHub Spec Kit similarly structures coding agent workflows around specification, planning, tasks, implementation, and convergence. In Akka's experiment, Sonnet averaged 61 minutes per port versus 120 minutes for Opus, while Opus used about 40% fewer tokens. Higher effort settings increased consumption without consistently improving efficiency.

The finding prompted discussion among engineers following the research. Aaditya, commenting on Tyler Jewell's CEO of Akka LinkedIn post, wrote

Smaller model's behavior matched modernization work he had observed, where the small model follows the spec while a larger model may improvise.

Aaditya also questioned

Whether reductions in lines of code resulted primarily from dead code removal or from differences in the target language.

Rick Bryce, Head of Marketing at Avahi, raised a related point in the same discussion and suggested that constraints could influence the result. The comments add questions around whether model capability, specification constraints, or both account for differences in porting efficiency.

Cheaper model giving the tighter port is the finding worth chasing

Akka model and effort efficiency chart (Source: Akka Blog Post) Validation used the original unit and integration tests alongside auditors checking serialization, security, error handling, PII, idempotency, and architectural boundaries. Akka added guardrails as failures exposed new issues, while reporting that additional exit conditions increased porting costs. Performance varied by project: applications, frameworks, and libraries generally improved, while infrastructure and tooling showed median degradation. Akka reported a 143,333 times improvement for Dify, but noted that the compared workloads differed, while Netflix Metaflow was approximately 100 times slower.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @akka 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/akka-tests-spec-driv…] indexed:0 read:2min 2026-10-05 · —