cd /news/ai-agents/agents-now-plan-tasks-across-your-co… · home › topics › ai-agents › article
[ARTICLE · art-145286] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Agents Now Plan Tasks Across Your Codebase

Raycast's AI Chat now runs multi-step tasks autonomously, chaining extension calls, executing code and retrying on failure without per-request interruptions, while shifting billing from rate limits to usage-based pricing with bring-your-own-subscription support for Claude and ChatGPT. Sourcegraph's Agentic Batch Changes delegates multi-repo migrations to Claude or Codex with outcome-based pricing per merged changeset, and Ollama's /v1/systemone endpoint returns calibrated probability distributions across predefined categories instead of text. The roundup concludes Raycast is a "Ship" and the other two warrant evaluation, with setup and model-coverage caveats noted.

read5 min views1 publishedOct 5, 2026

The throughline this week is agents moving from single-shot completions to multi-step execution with real consequences—writing PRs, merging changesets, routing tickets without a human in the loop. What's shifting isn't the capability ceiling; it's where the orchestration layer lives. Tools that used to hand you a diff now hand you a merged PR.

Raycast's AI Chat can now run multi-step tasks autonomously—chaining extension calls, executing code, retrying on failure—without per-request interruptions. Billing moves from rate limits to usage-based, and bring-your-own-subscription support (Claude, ChatGPT) means you can route through existing credits.

This matters because the friction in AI-assisted dev work has never been the individual completion—it's been stitching completions together manually. Raycast sits at the OS layer, which makes it a plausible orchestration host for tasks like code review triage, email filtering, or draft generation that previously required custom scripts or n8n-style wiring.

The Projects feature preserves conversation state across sessions, which is the detail that makes multi-step work actually viable. Without persistent context, agentic flows collapse on anything longer than a single working session.

Verdict: Ship. If you're already on Raycast, connect your Claude or ChatGPT account under Settings → AI → Models & Providers before billing resets. The usage-based shift is transparent at current price points. Start with one repetitive task you're already doing manually—triage or summarization—before building anything complex on top of it.

Sourcegraph's Agentic Batch Changes delegates multi-repo migrations to Claude or Codex: the agent writes migration scripts where it can, flags judgment calls where it can't, and tracks merge status across all repos from a single view. No spreadsheets. No oversized PRs that block unrelated changes.

The coordination tax on large migrations is real and underestimated. When you're moving a logging library or patching a CVE across 200 repos, the bottleneck isn't writing the change—it's knowing which repos landed it, which are blocked, and which have diverged enough that the script broke silently. That's the problem this targets.

Outcome-based pricing (pay per merged changeset) removes the token-cost uncertainty that makes LLM-heavy workflows hard to budget. That's a meaningful structural difference from per-token billing when you're touching hundreds of files.

Verdict: Evaluate. Requires Sourcegraph Cloud indexing your codebase and integration with your VCS, so there's real setup cost. If your org is already on Sourcegraph, this is worth piloting on a security fix or a well-scoped library migration where the blast radius is understood. Don't start with anything that requires repo-specific business logic—validate the agent's judgment on mechanical changes first.

Ollama's /v1/systemone endpoint returns structured probability distributions across predefined categories instead of text. Single request, calibrated confidence scores, no prompt engineering, no output parsing. Two models available: Nimble and Tev1.

Classification workflows built on text-completion models accumulate hidden costs: prompt tuning, output parsing, handling hallucinated categories, latency from multi-turn correction loops. A dedicated decision endpoint with calibrated probabilities is architecturally cleaner for ticket triage, content routing, or any hard-label classification task.

The catch is model coverage. Two models with limited documentation on what label spaces they handle well is a narrow surface. You'll need to validate that Nimble or Tev1 actually fits your category taxonomy before replacing anything in production.

Verdict: Evaluate. Requires Ollama v0.35.0+ and reformatted requests using the /v1/systemone schema. Worth prototyping for ticket routing or support triage if you're already running Ollama locally. Validate classification accuracy against your actual label distribution before committing—don't assume the confidence scores are calibrated for your domain without testing.

Two European open models land on Cloudflare's Workers AI: EuroLLM covering 35 languages, Apertus covering 1,500—including low-resource languages (Romansh, Swiss German, regional African and Asian languages) where frontier models underperform. Both ship with full weights and training data. Apertus is explicitly designed for GDPR and EU AI Act compliance.

For developers building on EU infrastructure or serving multilingual user bases, the dependency on US-hosted frontier models creates both legal friction and quality gaps. Rare and regional languages are where closed models fail quietly—returning fluent-sounding but semantically degraded output that's hard to catch without native speaker review. Apertus addresses a real quality gap, not just a compliance checkbox.

If you're already on Workers AI, integration requires no code changes beyond swapping model identifiers.

Verdict: Ship for multilingual or EU-regulated workloads where you're currently routing to a closed frontier model. Evaluate for everything else—the 1,500-language breadth of Apertus is compelling, but verify output quality on your target languages before replacing a working pipeline.

DeepSeek V4 Flash now accepts image input (JPEG, PNG, GIF, WebP) through Vercel's AI Gateway, with tool use, reasoning, and caching behavior matching the text-only version. No new dependencies if you're already on the ai SDK.

Consolidating vision and text inference at the same endpoint matters for latency-sensitive applications—screenshot analysis, chart parsing, document extraction—where switching providers or managing separate vision endpoints adds routing complexity and failure surface. The fact that caching parity carries over is the detail worth noting: cached responses on repeated image inputs keep costs predictable.

The experimental flag (-exp suffix on the model name) means the API surface can change without deprecation notice.

Verdict: Evaluate. Worth prototyping now for vision tasks where you're already on Vercel's AI Gateway. Configure a text-only fallback for any production path before the experimental flag resolves. Don't build a critical pipeline on a surface that can shift.

Gemini Omni 1.1 Flash now analyzes 10 seconds of prior video context for seamless scene extension, supports keyframe interpolation for deterministic output, generates 360p previews 60% faster, and adds 4K upscaling. Available via Gemini API with previous_interaction_id for chained extensions.

For developers building generative video pipelines, the fast 360p draft loop is the practical win: iterate on narrative continuity and timing cheaply before committing to a full-resolution generation. Keyframe specification makes outputs reproducible, which matters the moment you need to coordinate video generation with downstream editing or review steps. Verdict: Ship for video generation workflows where you're already on the Gemini API. Update your prompt structure to include previous_interaction_id for scene extensions and pass video references under 3 seconds. The 360p preview speed improvement alone reduces iteration cost enough to justify updating existing pipelines.

If this kind of technically precise coverage is useful, Dev Signal lands in your inbox every week—no hype, just what's worth your time and what isn't. Subscribe if you want the next 100 issues.

── more in #ai-agents 4 stories · sorted by recency
── more on @raycast 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agents-now-plan-task…] indexed:0 read:5min 2026-10-05 · —