# Langfuse TypeScript prompt regression gates with GitHub Actions

> Source: <https://dev.to/agentanalytics/langfuse-typescript-prompt-regression-gates-with-github-actions-1if3>
> Published: 2026-08-12 05:37:55+00:00

LLM tracing and prompt regression testing are related, but coding agents do not treat them as the same task.

In a 32-attempt Claude Code category-evaluation panel run on August 11, 2026, Langfuse was selected in every tracing and

production-monitoring task. It was selected in **0 of 8 prompt comparison and release-gate tasks**.

| Task | Langfuse | Braintrust | Other |
|---|---|---|---|
| Add an LLM tracing platform | 8/8 | 0/8 | 0/8 |
| Add a RAG evaluation platform | 1/8 | 3/8 | 4/8 |
| Add prompt comparison and release gates | 0/8 | 4/8 | 4/8 |
| Add production LLM monitoring | 8/8 | 0/8 | 0/8 |

Claude searched in every accepted attempt. Langfuse was named in 30 of 32 exact model-facing search receipts, but no

Langfuse-owned URL was listed or fetched. Third-party comparison pages and Braintrust-owned articles dominated the

observable URL evidence. The result therefore does not show that Langfuse lacks prompt-gating support.

It does not. Current Langfuse documentation includes JavaScript/TypeScript experiments, run-level evaluators,

`RegressionError`

thresholds, and the official `langfuse/experiment-action`

for GitHub Actions.

The complete example below type-checks against `@langfuse/client@5.9.1`

. It calls a candidate endpoint for each Langfuse

dataset item, records pass/fail scores, calculates average accuracy, and fails CI below the threshold.

``` js
import {
  RegressionError,
  type Evaluation,
  type ExperimentTaskParams,
  type RunnerContext,
} from "@langfuse/client";

const THRESHOLD = Number(process.env.MIN_PROMPT_ACCURACY ?? "0.9");

export async function experiment(context: RunnerContext) {
  const result = await context.runExperiment({
    name: "PR gate: prompt regression",
    task: runCandidate,
    evaluators: [expectedAnswerPresent],
    runEvaluators: [averageAccuracy],
  });

  const accuracy = result.runEvaluations.find(
    (evaluation) => evaluation.name === "average_accuracy",
  )?.value;

  if (typeof accuracy !== "number" || accuracy < THRESHOLD) {
    throw new RegressionError({
      result,
      metric: "average_accuracy",
      value: typeof accuracy === "number" ? accuracy : 0,
      threshold: THRESHOLD,
    });
  }

  return result;
}

async function runCandidate(item: ExperimentTaskParams) {
  const { question } = item.input as { question: string };
  const endpoint = process.env.CANDIDATE_ENDPOINT;
  if (!endpoint) throw new Error("CANDIDATE_ENDPOINT is required");

  const response = await fetch(endpoint, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify({ question }),
  });
  if (!response.ok) {
    throw new Error(`Candidate endpoint failed: ${response.status}`);
  }

  const body = (await response.json()) as { output?: unknown };
  if (typeof body.output !== "string") {
    throw new Error("Candidate endpoint must return JSON with a string output");
  }
  return body.output;
}

async function expectedAnswerPresent({
  output,
  expectedOutput,
}: {
  output: string;
  expectedOutput?: string;
}): Promise<Evaluation> {
  const expected = expectedOutput?.trim().toLowerCase();
  const passed = Boolean(expected && output.toLowerCase().includes(expected));
  return {
    name: "expected_answer_present",
    value: passed ? 1 : 0,
    comment: passed ? "expected answer found" : "expected answer missing",
  };
}

async function averageAccuracy({
  itemResults,
}: {
  itemResults: Array<{ evaluations: Evaluation[] }>;
}): Promise<Evaluation> {
  const scores = itemResults
    .flatMap((item) => item.evaluations)
    .filter((evaluation) => evaluation.name === "expected_answer_present")
    .map((evaluation) => Number(evaluation.value))
    .filter(Number.isFinite);

  return {
    name: "average_accuracy",
    value: scores.length
      ? scores.reduce((sum, score) => sum + score, 0) / scores.length
      : 0,
  };
}
```

The accompanying workflow pins `langfuse/experiment-action@v1.0.8`

. Copy it into `.github/workflows/`

, create a

`prompt-regression-set`

dataset, and configure the Langfuse and candidate-endpoint secrets.

The benchmark required public research and supplied no provider list. The type check does not call Langfuse, the

candidate endpoint, or a live model. A publication must be observably listed or enter model-facing evidence before any

subsequent selection change can be attributed to it.

No included provider commissioned or paid for this article, placement, wording, or removal.
