cd /news/ai-agents/passing-tests-expired-sqlite-leases-… · home › topics › ai-agents › article
[ARTICLE · art-143342] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Passing Tests, Expired SQLite Leases: An AI Coding Benchmark

A developer's experiment on a durable TypeScript and SQLite reminder queue found that both the Astra solo and Astra + Luna implementations passed their original acceptance suites while still containing a concurrency defect: they sampled the clock before SQLite granted the write lock, so lease decisions could be made using a stale timestamp. In a retrospective audit with real independent SQLite connections and a controlled clock, Astra solo passed 7/7 diagnostic checks while Astra + Luna passed only 4/7, though the fixed-rate token-cost estimate for Astra + Luna was 47.3% lower at the cost of 46.3% longer elapsed time. The fix is to acquire the transaction first and sample time only after the write lock is held.

by read5 min views6 publishedOct 1, 2026

A background job can pass its acceptance suite and still have a concurrency bug at the exact point where it waits for the database. In a small experiment I ran on a durable TypeScript and SQLite reminder queue, both Astra + Luna implementations passed the original checks. A later audit found that both could make lease decisions using a timestamp sampled before SQLite granted the write lock.

This is a narrow engineering result. A useful review question is whether a time-sensitive decision happens before or after a transaction wait.

The task was to build a durable reminder queue that survives restarts, retries failed deliveries, and handles competing workers. I compared four configurations, with two measured runs per configuration. In paired configurations, Astra planned and reviewed while another model implemented. The table below focuses on Astra solo and Astra + Luna because that is where the cost/time comparison and shared defect appeared.

Configuration Original acceptance Seven later diagnostic checks Mean fixed-rate estimate per attempt Mean elapsed time
Astra solo 2/2 runs accepted 7/7 and 7/7 38.681850 units 543.302 s
Astra + Luna 2/2 runs accepted 4/7 and 4/7 20.384872 units 795.081 s

The fixed-rate token-cost estimate for Astra + Luna was 47.3% lower than Astra solo; elapsed time was 46.3% longer. Those units are an estimate from token counts multiplied by fixed historical rates. They are not a bill, a measured subscription deduction, or a subscription-quota saving.

The diagnostic audit was retrospective. Three of the seven checks probe the same clock-after-lock defect. The 4/7 result is limited evidence about these two implementations on this one task.

A lease usually combines an owner token with an expiration time. A worker claims a job, performs work, then tries to complete or fail it. The database must serialize claims and state transitions so two workers cannot act on the same live lease.

SQLite can make a writer wait for another transaction. If application code reads the clock before requesting the write lock, the value can age during that wait. A simplified, illustrative ordering looks like this:

const now = clock.now();
await beginImmediateTransaction(db); // may wait for another writer
await claimDueJob(db, now);
await commit(db);

If the lock arrives after the previous lease deadline, this code can still evaluate the claim using the earlier time. Depending on the queue's rules, it may create a lease that is already expired, or accept a completion/failure based on ownership that expired while the transaction waited.

For a lock-protected decision, acquire the transaction first and sample time only after the write lock is held:

await beginImmediateTransaction(db);
const now = clock.now();
await claimDueJob(db, now);
await commit(db);

Treat the snippet as an ordering illustration and adapt it to the real transaction API. Keep the time comparison and relevant state update inside the same transaction, preserve owner-token checks, and account for clock semantics. Sampling after lock acquisition removes this particular pre-wait staleness; it does not prove that every timing or lease bug is solved.

The ORCH-1 audit exercised this boundary with real, independent SQLite connections and a controlled clock. The simplified outline below follows that audit; adapt the barriers and connection setup to your own queue:

In the public audit, the clock advanced from 0 to 10 while the operation waited. With a lease duration of 5, a fresh claim should end at 15. Completion and failure should reject ownership that expired at 5. Both Astra + Luna runs instead returned a claim ending at 5 and accepted expired ownership.

The barrier and controllable clock matter. A timing test based only on real sleeps can be flaky and may never hit the boundary in CI. The assertions should distinguish the operations: a new claim may reclaim work after its prior lease expires, but should calculate the new lease from time sampled after lock acquisition; complete and fail should verify the owner token and reject expired ownership using current post-lock time. Assert those state invariants, not an incidental number of milliseconds.

In this experiment, both paired implementations read the clock before waiting for SQLite's write lock. Three probes for claim, complete, and fail each exposed that same timing defect. Astra's review did catch other concrete issues: it found string-preservation problems in SQLite storage, which Luna then fixed. The review still missed the lock-wait boundary, and the supplied tests did not exercise it.

The cost/time tradeoff is real for these measured attempts, but it should not be read as a ranking of models. The test used one TypeScript/SQLite task and two runs per setup. There was no Sol-only control. Before the Astra + Luna runs, the CLI version and executor-selection protocol changed. The additional audit was designed after inspecting the implementations, and its three clock checks share one defect class.

Astra's planning and review accounted for 95.6% of the paired workflow's fixed-rate estimate in these runs. Even a hypothetical free executor would remove only the remaining 4.4% if all observed work stayed fixed. That is a useful reminder to measure the whole workflow, not just the executor price. It is not a general statement about orchestration economics.

For a production queue, I would put lock-wait cases in the acceptance suite from the start, test through separate database connections, and make the clock controllable. I would also keep the benchmark's acceptance result separate from the later diagnostic result so a retrospective test does not rewrite what was originally accepted.

The next comparison should add a Sol-only control under the same client and protocol, use more tasks, and freeze the expanded checks before running the candidates.

AI-generated draft: AI prepared this article from my published report and project notes. The linked sources contain the experiment results and methodology; this article makes no additional benchmark claims.

── more in #ai-agents 4 stories · sorted by recency
── more on @astra 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/passing-tests-expire…] indexed:0 read:5min 2026-10-01 · —