cd/entity/Harbor· home› entities› Harbor
grep -l @harbor /news/*.json | wc -l → 45

Harbor

mentions 45 type Organization page 1/3 feed RSS

// recent coverage 45 mentions

00:00
2026-10-06
gethrbr.com
ai-agents

How to keep the same rules in Claude Code, Codex and Cursor

Writing rules once in a repo-root AGENTS.md file lets Codex and Cursor read them directly, while Claude Code reads AGENTS.md only when no CLAUDE.md exists or via an @AGENTS.md import line, according t…

22:54
2026-09-30
github.com
artificial-intelligence

Context Language Models (CLMs)

Researchers from the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs introduced Context Language Models (CLMs), which treat context as a file the model can update itself, …

00:00
2026-09-28
huggingface.co
ai-agents

Welcome RL Environments to the hub

Hugging Face added an RL Environments filter and framework tags to its Hub, letting reinforcement-learning environment datasets be discovered and loaded directly from dataset repositories rather than …

16:39
2026-09-22
horizonanalyticslabs.com
ai-research

Benchmarks are more broken than we could have imagined

An audit by Horizon of 20 public task datasets in the Harbor hub found 29 confirmed broken tasks out of 5,241 scanned, with failures that often made models look better rather than worse, according to …

15:05
2026-09-21
github.com
ai-tools

Practical model evaluation and compression tools

Developer 0xSero released model-toolkit, a GitHub repository of standalone Python tools for evaluating, observing, pruning, and quantizing language models, drawn from the REAP and EXL3 experiments inc…

18:22
2026-09-17
vercel.com
ai-research

Run Terminal-Bench and other Harbor evals on Vercel Sandbox

Vercel now supports running Harbor evaluations, including Terminal-Bench, SWE-bench, tau3-bench and OSWorld, on Vercel Sandbox, with each trial executing in its own isolated Firecracker microVM when u…

00:00
2026-09-15
islo.dev
ai-agents

How ARIMLABS runs 500,000 agent executions a day on Islo

ARIMLABS now runs roughly 500,000 agent executions per day on Islo's cloud computers, up from a small pilot, according to CEO Mykyta Mudryi. The company moved its Harbor-based long-horizon agent evalu…

00:00
2026-09-14
stripe.dev
ai-tools

Harbor: Stripe’s AI-assisted prototyping tool

Stripe staff engineer Cristian Rivera created Harbor, an internal AI-assisted prototyping tool, according to a Stripe post authored by technical writer Sai Samant. Rivera works on Stripe's Web Presenc…

16:35
2026-09-01
dev.to
ai-agents

How to Design AI Evaluations You Can Actually Trust

Google's Developer Relations team published a suite of Agent Skills for Google products on GitHub and outlined five rules for designing trustworthy AI evaluations. The rules emphasize understanding th…

08:00
2026-08-22
infoq.com
artificial-intelligence

AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks

AWS released aws-bench, an open-source benchmark to evaluate AI agents on real AWS tasks, using disposable AWS accounts and automated verifiers. The benchmark, built on Harbor, supports agents like Cl…

16:21
2026-08-21
frontierroles.com
ai-infrastructure

Member of Technical Staff, Mercor Enterprise Platform — Mercor

Mercor, a profitable Series C AI data company valued at $10 billion, is hiring a Member of Technical Staff for its Enterprise Platform in San Francisco, offering a salary of $130k–500k/yr. The role in…

18:36
2026-08-19
unsloth.ai
artificial-intelligence

Unsloth Dynamic 3.0 GGUFs

Unsloth released Dynamic v3.0 GGUFs for Qwen3.8-27B, claiming more than 10% better top-1% accuracy at the same size compared to every other provider. The new quants, which work with llama.cpp and Unsl…

18:10
2026-08-18
cline.ghost.io
ai-research

Open-sourcing evals for open-weight agents

Cline, an AI coding assistant, is open-sourcing its evaluation framework for open-weight coding agents, revealing that its requests run 20–30% heavier on tokens than the most efficient harnesses. The …

07:00
2026-08-18
dev.to
artificial-intelligence

Designing AI Evals: Clarity Now and Visualization Next

A developer from Google Cloud's DevRel team demonstrates how to design objective evaluations for AI agent skills using open-source frameworks like Inspect AI and Harbor. The investigation uses Gemini …

00:10
2026-08-16
github.com
artificial-intelligence

Big Pickle on SWE Atlas – Codebase QnA

OpenCode Zen's free stealth model big-pickle resolved 50.8% (63/124) of Scale AI's SWE Atlas Codebase QnA benchmark tasks on 2026-08-11, outperforming all official Mini-SWE-Agent scaffold entries and …

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics