Look for Long Horizon Agents for frontier Labs, located in Mountain View (Remote) Bespoke Labs is hiring a contract researcher to design and evaluate reinforcement learning environments and benchmarks for long-horizon agentic tasks, which require hours, days, or weeks of coherent multi-step reasoning. The role involves building RL environments and verifiers grounded in real-world tasks, developing evaluation benchmarks, and analyzing failure modes in long-horizon rollouts. Candidates must demonstrate concrete long-horizon agent or RL experience, such as contributions to SWE-bench, Vending-Bench, or similar projects. Bespoke Labs is looking for a researcher to help design and evaluate RL environments and benchmarks for long-horizon agentic tasks — the kind that take an agent hours, days, or weeks of coherent multi-step reasoning to complete, not single-turn prompts. What you'll do Design and build long-horizon RL environments and verifiers grounded in real-world tasks code, tool-use, or enterprise workflows Develop evaluation benchmarks that measure agent coherence, planning, and reliability over extended trajectories Analyze failure modes in long-horizon rollouts drift, reward hacking, loss of task state and propose fixes Collaborate with the broader team on open datasets and reproducible eval recipes Must-have hard requirement Demonstrated long-horizon agent/RL experience — this is non-negotiable. You should be able to point to specific work involving multi-step, multi-day, or sequential-reasoning agent systems e.g., contributions to environments like SWE-bench, Vending-Bench, FrontierSWE, DeepSWE, OpenReward, Gymnasium, or equivalent original research/production work . Applications without concrete long-horizon evidence will not be considered. Strong Python; comfort with RL training/eval frameworks e.g., Verifiers, Gymnasium-style APIs, or custom environment tooling Track record of publishing or shipping work others can verify GitHub, papers, benchmarks, or production systems Nice to have Experience with reward-hacking detection or "fuzzy" quality verifiers beyond pass/fail correctness Background in multi-agent coordination or agent memory systems Prior contributions to open-source RL environment or agent-eval projects Logistics Type: Contract, remote Location: Remote any timezone considered; some overlap with US/India hours preferred Compensation: Based on experience — happy to discuss How to apply Send a short note plus links to your relevant long-horizon work GitHub, papers, benchmarks, or production systems you've shipped to https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5 https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5 . No long-horizon evidence, no need to apply — we will reject on this filter first.