cd/entity/Harbor· home entities Harbor
grep -l @harbor /news/*.json | wc -l → 31

Harbor

mentions 31 type Organization page 1/2 feed RSS

// recent coverage 31 mentions

08:00
2026-08-22
infoq.com
artificial-intelligence

AWS Releases Aws-Bench to Evaluate Agents on Cloud Tasks

AWS released aws-bench, an open-source benchmark to evaluate AI agents on real AWS tasks, using disposable AWS accounts and automated verifiers. The benchmark, built on Harbor, supports agents like Cl…

16:21
2026-08-21
frontierroles.com
ai-infrastructure

Member of Technical Staff, Mercor Enterprise Platform — Mercor

Mercor, a profitable Series C AI data company valued at $10 billion, is hiring a Member of Technical Staff for its Enterprise Platform in San Francisco, offering a salary of $130k–500k/yr. The role in…

18:36
2026-08-19
unsloth.ai
artificial-intelligence

Unsloth Dynamic 3.0 GGUFs

Unsloth released Dynamic v3.0 GGUFs for Qwen3.8-27B, claiming more than 10% better top-1% accuracy at the same size compared to every other provider. The new quants, which work with llama.cpp and Unsl…

18:10
2026-08-18
cline.ghost.io
ai-research

Open-sourcing evals for open-weight agents

Cline, an AI coding assistant, is open-sourcing its evaluation framework for open-weight coding agents, revealing that its requests run 20–30% heavier on tokens than the most efficient harnesses. The …

07:00
2026-08-18
dev.to
artificial-intelligence

Designing AI Evals: Clarity Now and Visualization Next

A developer from Google Cloud's DevRel team demonstrates how to design objective evaluations for AI agent skills using open-source frameworks like Inspect AI and Harbor. The investigation uses Gemini …

00:10
2026-08-16
github.com
artificial-intelligence

Big Pickle on SWE Atlas – Codebase QnA

OpenCode Zen's free stealth model big-pickle resolved 50.8% (63/124) of Scale AI's SWE Atlas Codebase QnA benchmark tasks on 2026-08-11, outperforming all official Mini-SWE-Agent scaffold entries and …

00:48
2026-08-15
frontierroles.com
ai-infrastructure

Member of Technical Staff, Enterprise Evals Platform — Mercor

Mercor, a profitable Series C AI data company valued at $10 billion, is hiring a Member of Technical Staff for its Enterprise Evals Platform in San Francisco, offering a salary of $220,000–$425,000 pe…

06:24
2026-07-29
brandonbarker.me
artificial-intelligence

Headroom cut 39% of my tokens and raised my Claude bill

A controlled test of the token-saving tool Headroom on 25 senior software engineering tasks using Anthropic's Claude found that while it cut input tokens by 39%, it did not reduce costs due to cache b…

12:13
2026-07-25
letta.com
ai-agents

Trajectory: A Standard Format for Agent Experience Data

Letta AI introduced the Trajectory package, a standardized data format for agent experience that reduces token counts by up to 5.6× compared to native session formats, enabling agents to learn across …

13:39
2026-07-24
frontierbench.ai
ai-agents

Frontier-Bench

The team behind Terminal-Bench and Harbor released Frontier-Bench v0.1, a continuous benchmark with 74 tasks across 7 domains that measures agent abilities at the frontier, with the best models achiev…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics