cd /news/artificial-intelligence/look-for-long-horizon-agents-for-fro… · home topics artificial-intelligence article
[ARTICLE · art-85693] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Look for Long Horizon Agents for frontier Labs, located in Mountain View (Remote)

Bespoke Labs is hiring a contract researcher to design and evaluate reinforcement learning environments and benchmarks for long-horizon agentic tasks, which require hours, days, or weeks of coherent multi-step reasoning. The role involves building RL environments and verifiers grounded in real-world tasks, developing evaluation benchmarks, and analyzing failure modes in long-horizon rollouts. Candidates must demonstrate concrete long-horizon agent or RL experience, such as contributions to SWE-bench, Vending-Bench, or similar projects.

read1 min views2 publishedAug 4, 2026

Bespoke Labs is looking for a researcher to help design and evaluate RL environments and benchmarks for long-horizon agentic tasks — the kind that take an agent hours, days, or weeks of coherent multi-step reasoning to complete, not single-turn prompts.

What you'll do

Design and build long-horizon RL environments and verifiers grounded in real-world tasks (code, tool-use, or enterprise workflows)

Develop evaluation benchmarks that measure agent coherence, planning, and reliability over extended trajectories

Analyze failure modes in long-horizon rollouts (drift, reward hacking, loss of task state) and propose fixes

Collaborate with the broader team on open datasets and reproducible eval recipes

Must-have (hard requirement) Demonstrated long-horizon agent/RL experience — this is non-negotiable. You should be able to point to specific work involving multi-step, multi-day, or sequential-reasoning agent systems (e.g., contributions to environments like SWE-bench, Vending-Bench, FrontierSWE, DeepSWE, OpenReward, Gymnasium, or equivalent original research/production work). Applications without concrete long-horizon evidence will not be considered.

Strong Python; comfort with RL training/eval frameworks (e.g., Verifiers, Gymnasium-style APIs, or custom environment tooling)

Track record of publishing or shipping work others can verify (GitHub, papers, benchmarks, or production systems)

Nice to have

Experience with reward-hacking detection or "fuzzy" quality verifiers beyond pass/fail correctness

Background in multi-agent coordination or agent memory systems

Prior contributions to open-source RL environment or agent-eval projects

Logistics

Type: Contract, remote

Location: Remote (any timezone considered; some overlap with US/India hours preferred)

Compensation: Based on experience — happy to discuss

How to apply

Send a short note plus links to your relevant long-horizon work (GitHub, papers, benchmarks, or production systems you've shipped) to [https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5]. No long-horizon evidence, no need to apply — we will reject on this filter first.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @bespoke labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/look-for-long-horizo…] indexed:0 read:1min 2026-08-04 ·