{"slug": "measure-how-engineers-and-teams-work-with-ai", "title": "Measure how engineers and teams work with AI", "summary": "Eversynced released Pheebs, an open-source AI telemetry tool that installs via `npm install -g pheebs` and registers lifecycle hooks inside Claude Code, Cursor, and Codex to capture interaction signals — event types, durations, counts, models, and trigger types — without storing prompt text. Pheebs sizes each task prompt and compares it with the model that ran, pricing the gap in dollars from a team's real sessions; a fresh install ships with no backend endpoint and writes only to `~/.pheebs/logs/`, and events post to a user-supplied backend only when both a base URL and token are set. The tool reports observations such as 61% of AI-written lines in a payments service shipping with no check and one in three follow-up prompts fixing something the AI broke.", "body_md": "“$9,960 of last month’s model spend went to a bigger model than the work needed.”\n\nModel spend that buys nothing.\n\nAn open-source AI telemetry tool created by Eversynced. It sits quietly inside the AI coding agents Claude Code, Cursor, and Codex via hooks, capturing lightweight interaction signals: the shape of the session, not its contents.\n\n`npm install -g pheebs`  Works with Claude CodeCursorCodex\n\nPheebs records the shape of the session, never its contents.\n\n**What you did**, on your machine\n\n`src/billing/checkout.test.ts`  `src/billing/session.ts`  `npm test -- checkout`  **What Pheebs recorded**, interaction signals\n\nThe events roll up into observations about your own work, the kind a colleague might make if they had been sitting next to you for a month and could remember all of it.\n\n“61% of AI-written lines in the payments service shipped with no check.”\n\nWhere AI code ships unchallenged.\n\n“Half of your team’s AI edits never had a test, typecheck, or build run behind them.”\n\nWhether AI output gets verified.\n\n“One in three follow-up prompts was fixing something the AI broke, not moving the work forward.”\n\nRework hiding inside the speedup.\n\n“The review skill you shipped is used weekly by 78% of engineers. The migration skill never caught on.”\n\nWhether the enablement investment landed.\n\n“Nobody on the team runs tests inside the agent loop. That’s a missing harness, not a skills gap.”\n\nFix the setup, or coach the people.\n\nMost teams run the largest model for everything, because nothing tells them what the work needed. Pheebs sizes the work in every session and compares it with the model that ran. The gap between the two is a savings opportunity, priced in dollars from your team's real sessions.\n\nTasks are sized\n\nEvery task prompt gets a scope, from a one-file change to open-ended design. A session is judged on its hardest prompt.\n\nMisses count both ways\n\nAn over-provisioned session burns budget silently. An under-powered one shows up as repair prompts.\n\nAn audit, not a router\n\nPheebs never intercepts a prompt or switches a model on anyone's behalf. It reads the gap and prices it. The decision stays yours.\n\nPheebs registers lifecycle hooks in the agent's own config, plus OpenTelemetry export where the tool supports it. From then on it fires in the background.\n\nA hook fires\n\nSession starts and ends, prompts, skill and slash-command expansions, sub-agent spawns, tool calls and failures, compaction, background tasks.\n\nLightweight fields are extracted\n\nEvent type, durations, counts, models, trigger types. A prompt becomes a character count and, when the prompt intent classifier is enabled, an intent label. The text itself is never stored.\n\nIdentity and repo are resolved\n\nThe developer is the id behind your Pheebs token, stamped by the backend, or a truncated hash of your git email when no token is set. The codebase is org/repo from the git remote.\n\nLogged locally, then sent\n\nEvery event is stored in a local JSONL log. With a token set, it also goes to the backend.\n\nOpenTelemetry rides along\n\nClaude Code and Codex export native OTel metrics and logs through the Pheebs proxy.\n\nA fresh install ships with no backend endpoint, so it writes only to `~/.pheebs/logs/`.\n\n`org/repo` from the git remote.  `npm test` is read in process and recorded as `tool_intent: test_run`.  Pheebs captures interaction patterns.\n\nPheebs holds a base URL and a token, and nothing else. Give it both and the events are posted to your backend as well.\n\nBoth need to be set, or nothing is posted. Unset either one and sending stops. The local JSONL stays the durable copy either way.\n\nYou can find a reference backend in the Pheebs repo.\n\nWhat your endpoint answers\n\nProficiency = repertoire + judgement signals.\n\nPheebs proposes a specific way of looking at the data it captures: the AI Proficiency Model. The model defines the practices and signals that matter when working with AI agents, and gives structure to what would otherwise be a stream of raw events.\n\nLayer 1 reads repertoire: which AI harness capabilities show up in an engineer's work. Layer 2 reads judgement: what happens to AI output before it ships, and whether the model that ran fit the work.\n\nAI harness repertoire\n\nTwenty-four practices across six competencies. Each practice is detectable from telemetry: either its detector fired or it didn't.\n\nModels\n\nWhich models are in play: model choice, effort settings, plan mode, and autonomy modes.\n\nArtifacts\n\nThe reusable config that shapes the agent: skills, sub-agents, slash commands, and context files.\n\nMCP\n\nLive connections to external systems: tickets, databases, browsers, and documentation.\n\nEvals\n\nVerification wired into the agent loop: tests, typecheck, lint, build, and review passes.\n\nContext management\n\nDeliberate use of the context window: compaction and the save, resume, clear lifecycle.\n\nOrchestration\n\nMore than one agent at a time: sub-agents, parallel work, worktrees, hooks, and plugins.\n\nAI judgement signals\n\nJudgement on both sides of the AI loop: whether output gets verified, challenged, and refined before it ships (inspired by the Discernment competency from Anthropic's AI Fluency Index), and whether the model chosen fit the work. Continuous rates computed from session telemetry.\n\nVerification coverage\n\nThe share of AI edits followed by a verification action: a test run, typecheck, lint, build, or a check against a spec.\n\nPushback rate\n\nHow often the engineer challenges AI output instead of accepting it. Questioning collapses exactly when output looks polished.\n\nRefinement-to-repair ratio\n\nWhether follow-up prompts refine intent (healthy iteration) or repair breakage (rework).\n\nWholesale-accept rate\n\nSessions with no pushback, no repair, and no verification, weighted by lines changed. The composite red flag: polished output, no questions asked.\n\nModel-fit rate\n\nThe share of sessions whose model class matched the size of the work. Misses count both ways: over-provisioned burns budget, under-powered shows up as repairs.\n\n| Engineer | Verification | Pushback | Refine : repair | Wholesale | Model-fit | \n|---|---|---|---|---|---|\n| Michael | 39% | 9% | 1.1 : 1 | 41% | 58% | \n| Dwight | 55% | 17% | 1.6 : 1 | 26% | 63% | \n| Jim | 44% | 12% | 1.2 : 1 | 38% | 51% | \n| Pam | 71% | 22% | 2.4 : 1 | 15% | 35% | \n| Angela | 62% | 15% | 1.9 : 1 | 21% | 66% | \n| Kevin | 26% | 4% | 0.7 : 1 | 55% | 41% | \n| Team median | 50% | 14% | 1.4 : 1 | 33% | 55% | \n\nEither run the backend, hold the data, and set up the reporting or hire us to do it for you.\n\nSelf-hosted\n\nStand up backend. Telemetry goes from your developers' machines to your infrastructure. We never see it.\n\nManaged\n\nThe same open-source client, pointed at a backend we operate, with the proficiency model rendered as reports and dashboards. That is our AI Enablement Assessment service: a 30-day telemetry sprint that ends in an executive debrief and a plan for the gaps.\n\n`init` detects the agents you already have, configures the ones you pick, and asks where to send events. Leave that blank and Pheebs stays local.\n\nEvery Eversynced engineer is instrumented with Pheebs. It powers the measurement layer of our AI delivery framework, and the reporting built on top of it ships with the AI Enablement Assessment we run for client teams.", "url": "https://wpnews.pro/news/measure-how-engineers-and-teams-work-with-ai", "canonical_source": "https://www.eversynced.com/pheebs/", "published_at": "2026-10-08 20:59:13+00:00", "updated_at": "2026-10-08 21:17:21.471436+00:00", "lang": "en", "topics": ["ai-tools", "ai-agents", "developer-tools", "mlops", "ai-infrastructure"], "entities": ["Eversynced", "Pheebs", "Claude Code", "Cursor", "Codex", "OpenTelemetry"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/measure-how-engineers-and-teams-work-with-ai", "markdown": "https://wpnews.pro/news/measure-how-engineers-and-teams-work-with-ai.md", "text": "https://wpnews.pro/news/measure-how-engineers-and-teams-work-with-ai.txt", "jsonld": "https://wpnews.pro/news/measure-how-engineers-and-teams-work-with-ai.jsonld"}}