{"slug": "look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote", "title": "Look for Long Horizon Agents for frontier Labs, located in Mountain View (Remote)", "summary": "Bespoke Labs is hiring a contract researcher to design and evaluate reinforcement learning environments and benchmarks for long-horizon agentic tasks, which require hours, days, or weeks of coherent multi-step reasoning. The role involves building RL environments and verifiers grounded in real-world tasks, developing evaluation benchmarks, and analyzing failure modes in long-horizon rollouts. Candidates must demonstrate concrete long-horizon agent or RL experience, such as contributions to SWE-bench, Vending-Bench, or similar projects.", "body_md": "Bespoke Labs is looking for a researcher to help design and evaluate RL environments and benchmarks for long-horizon agentic tasks — the kind that take an agent hours, days, or weeks of coherent multi-step reasoning to complete, not single-turn prompts.\n\nWhat you'll do\n\nDesign and build long-horizon RL environments and verifiers grounded in real-world tasks (code, tool-use, or enterprise workflows)\n\nDevelop evaluation benchmarks that measure agent coherence, planning, and reliability over extended trajectories\n\nAnalyze failure modes in long-horizon rollouts (drift, reward hacking, loss of task state) and propose fixes\n\nCollaborate with the broader team on open datasets and reproducible eval recipes\n\nMust-have (hard requirement)\n\nDemonstrated long-horizon agent/RL experience — this is non-negotiable. You should be able to point to specific work involving multi-step, multi-day, or sequential-reasoning agent systems (e.g., contributions to environments like SWE-bench, Vending-Bench, FrontierSWE, DeepSWE, OpenReward, Gymnasium, or equivalent original research/production work). Applications without concrete long-horizon evidence will not be considered.\n\nStrong Python; comfort with RL training/eval frameworks (e.g., Verifiers, Gymnasium-style APIs, or custom environment tooling)\n\nTrack record of publishing or shipping work others can verify (GitHub, papers, benchmarks, or production systems)\n\nNice to have\n\nExperience with reward-hacking detection or \"fuzzy\" quality verifiers beyond pass/fail correctness\n\nBackground in multi-agent coordination or agent memory systems\n\nPrior contributions to open-source RL environment or agent-eval projects\n\nLogistics\n\nType: Contract, remote\n\nLocation: Remote (any timezone considered; some overlap with US/India hours preferred)\n\nCompensation: Based on experience — happy to discuss\n\nHow to apply\n\nSend a short note plus links to your relevant long-horizon work (GitHub, papers, benchmarks, or production systems you've shipped) to [[https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5](https://experts.bespokelabs.ai/expert/apply/mts-long-horizon-coding-tasks-ER000018?src=JOsj6ol5)]. No long-horizon evidence, no need to apply — we will reject on this filter first.", "url": "https://wpnews.pro/news/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote", "canonical_source": "https://dev.to/ahmed_khan_6c5f55092f881b/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote-3b5p", "published_at": "2026-08-04 06:58:08+00:00", "updated_at": "2026-08-04 07:11:18.721739+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-research"], "entities": ["Bespoke Labs", "SWE-bench", "Vending-Bench", "FrontierSWE", "DeepSWE", "OpenReward", "Gymnasium"], "alternates": {"html": "https://wpnews.pro/news/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote", "markdown": "https://wpnews.pro/news/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote.md", "text": "https://wpnews.pro/news/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote.txt", "jsonld": "https://wpnews.pro/news/look-for-long-horizon-agents-for-frontier-labs-located-in-mountain-view-remote.jsonld"}}