# Measure how engineers and teams work with AI

> Source: <https://www.eversynced.com/pheebs/>
> Published: 2026-10-08 20:59:13+00:00

“$9,960 of last month’s model spend went to a bigger model than the work needed.”

Model spend that buys nothing.

An open-source AI telemetry tool created by Eversynced. It sits quietly inside the AI coding agents Claude Code, Cursor, and Codex via hooks, capturing lightweight interaction signals: the shape of the session, not its contents.

`npm install -g pheebs`  Works with Claude CodeCursorCodex

Pheebs records the shape of the session, never its contents.

**What you did**, on your machine

`src/billing/checkout.test.ts`  `src/billing/session.ts`  `npm test -- checkout`  **What Pheebs recorded**, interaction signals

The events roll up into observations about your own work, the kind a colleague might make if they had been sitting next to you for a month and could remember all of it.

“61% of AI-written lines in the payments service shipped with no check.”

Where AI code ships unchallenged.

“Half of your team’s AI edits never had a test, typecheck, or build run behind them.”

Whether AI output gets verified.

“One in three follow-up prompts was fixing something the AI broke, not moving the work forward.”

Rework hiding inside the speedup.

“The review skill you shipped is used weekly by 78% of engineers. The migration skill never caught on.”

Whether the enablement investment landed.

“Nobody on the team runs tests inside the agent loop. That’s a missing harness, not a skills gap.”

Fix the setup, or coach the people.

Most teams run the largest model for everything, because nothing tells them what the work needed. Pheebs sizes the work in every session and compares it with the model that ran. The gap between the two is a savings opportunity, priced in dollars from your team's real sessions.

Tasks are sized

Every task prompt gets a scope, from a one-file change to open-ended design. A session is judged on its hardest prompt.

Misses count both ways

An over-provisioned session burns budget silently. An under-powered one shows up as repair prompts.

An audit, not a router

Pheebs never intercepts a prompt or switches a model on anyone's behalf. It reads the gap and prices it. The decision stays yours.

Pheebs registers lifecycle hooks in the agent's own config, plus OpenTelemetry export where the tool supports it. From then on it fires in the background.

A hook fires

Session starts and ends, prompts, skill and slash-command expansions, sub-agent spawns, tool calls and failures, compaction, background tasks.

Lightweight fields are extracted

Event type, durations, counts, models, trigger types. A prompt becomes a character count and, when the prompt intent classifier is enabled, an intent label. The text itself is never stored.

Identity and repo are resolved

The developer is the id behind your Pheebs token, stamped by the backend, or a truncated hash of your git email when no token is set. The codebase is org/repo from the git remote.

Logged locally, then sent

Every event is stored in a local JSONL log. With a token set, it also goes to the backend.

OpenTelemetry rides along

Claude Code and Codex export native OTel metrics and logs through the Pheebs proxy.

A fresh install ships with no backend endpoint, so it writes only to `~/.pheebs/logs/`.

`org/repo` from the git remote.  `npm test` is read in process and recorded as `tool_intent: test_run`.  Pheebs captures interaction patterns.

Pheebs holds a base URL and a token, and nothing else. Give it both and the events are posted to your backend as well.

Both need to be set, or nothing is posted. Unset either one and sending stops. The local JSONL stays the durable copy either way.

You can find a reference backend in the Pheebs repo.

What your endpoint answers

Proficiency = repertoire + judgement signals.

Pheebs proposes a specific way of looking at the data it captures: the AI Proficiency Model. The model defines the practices and signals that matter when working with AI agents, and gives structure to what would otherwise be a stream of raw events.

Layer 1 reads repertoire: which AI harness capabilities show up in an engineer's work. Layer 2 reads judgement: what happens to AI output before it ships, and whether the model that ran fit the work.

AI harness repertoire

Twenty-four practices across six competencies. Each practice is detectable from telemetry: either its detector fired or it didn't.

Models

Which models are in play: model choice, effort settings, plan mode, and autonomy modes.

Artifacts

The reusable config that shapes the agent: skills, sub-agents, slash commands, and context files.

MCP

Live connections to external systems: tickets, databases, browsers, and documentation.

Evals

Verification wired into the agent loop: tests, typecheck, lint, build, and review passes.

Context management

Deliberate use of the context window: compaction and the save, resume, clear lifecycle.

Orchestration

More than one agent at a time: sub-agents, parallel work, worktrees, hooks, and plugins.

AI judgement signals

Judgement on both sides of the AI loop: whether output gets verified, challenged, and refined before it ships (inspired by the Discernment competency from Anthropic's AI Fluency Index), and whether the model chosen fit the work. Continuous rates computed from session telemetry.

Verification coverage

The share of AI edits followed by a verification action: a test run, typecheck, lint, build, or a check against a spec.

Pushback rate

How often the engineer challenges AI output instead of accepting it. Questioning collapses exactly when output looks polished.

Refinement-to-repair ratio

Whether follow-up prompts refine intent (healthy iteration) or repair breakage (rework).

Wholesale-accept rate

Sessions with no pushback, no repair, and no verification, weighted by lines changed. The composite red flag: polished output, no questions asked.

Model-fit rate

The share of sessions whose model class matched the size of the work. Misses count both ways: over-provisioned burns budget, under-powered shows up as repairs.

| Engineer | Verification | Pushback | Refine : repair | Wholesale | Model-fit | 
|---|---|---|---|---|---|
| Michael | 39% | 9% | 1.1 : 1 | 41% | 58% | 
| Dwight | 55% | 17% | 1.6 : 1 | 26% | 63% | 
| Jim | 44% | 12% | 1.2 : 1 | 38% | 51% | 
| Pam | 71% | 22% | 2.4 : 1 | 15% | 35% | 
| Angela | 62% | 15% | 1.9 : 1 | 21% | 66% | 
| Kevin | 26% | 4% | 0.7 : 1 | 55% | 41% | 
| Team median | 50% | 14% | 1.4 : 1 | 33% | 55% | 

Either run the backend, hold the data, and set up the reporting or hire us to do it for you.

Self-hosted

Stand up backend. Telemetry goes from your developers' machines to your infrastructure. We never see it.

Managed

The same open-source client, pointed at a backend we operate, with the proficiency model rendered as reports and dashboards. That is our AI Enablement Assessment service: a 30-day telemetry sprint that ends in an executive debrief and a plan for the gaps.

`init` detects the agents you already have, configures the ones you pick, and asks where to send events. Leave that blank and Pheebs stays local.

Every Eversynced engineer is instrumented with Pheebs. It powers the measurement layer of our AI delivery framework, and the reporting built on top of it ships with the AI Enablement Assessment we run for client teams.
