Wire promptfoo into GitHub Actions so prompt changes get tested against a golden dataset before merge.
What you'll build #
A GitHub Actions check that runs a promptfoo eval suite against a golden dataset on every pull request that touches your prompts, fails the check when outputs regress, and posts a before/after comparison as a PR comment.
Prerequisites #
- Node.js 22.22.0 or newer. promptfoo enforces this; the project recommends Node 24 LTS. Check with
node --version. - promptfoo 0.122.2 (the version this tutorial was verified against). Anything in the 0.122.x line works the same way.
- A GitHub repository you can add secrets and workflows to.
- An OpenAI API key. promptfoo supports 60+ providers, so swap in Anthropic, Azure, or a local model if you prefer; the action has a matching
*-api-keyinput for each. - Commands are written for macOS/Linux. On Windows, use WSL.
1. Install promptfoo and scaffold the project #
npm install -g promptfoo
promptfoo --version # 0.122.2 at time of writing
mkdir prompt-ci && cd prompt-ci
git init
mkdir prompts
You can also run everything through npx promptfoo@latest without installing, which is exactly what the GitHub Action does on the runner.
2. Write the prompt under test #
The example is a support-ticket triage prompt that must return strict JSON. Structured output is where silent regressions hurt most: a "small wording tweak" that makes the model chatty breaks every downstream parser.
Create prompts/triage.txt:
You are a support ticket triage bot for a SaaS product.
Classify the ticket into exactly one category: billing, bug, account, or feature_request.
Rate urgency as low, medium, or high.
Respond with JSON only, no prose: {"category": "...", "urgency": "..."}
Ticket: {{ticket}}
{{ticket}} is a Nunjucks variable that promptfoo fills from each test case.
3. Define the golden dataset #
Create promptfooconfig.yaml in the repo root. The tests block is your golden dataset: real inputs with known-correct expectations. Start with cases that have burned you before.
description: Ticket triage regression suite
prompts:
- file://prompts/triage.txt
providers:
- id: openai:chat:gpt-5.6-luna
config:
reasoning_effort: low
max_completion_tokens: 200
defaultTest:
assert:
- type: is-json
tests:
- vars:
ticket: I was charged twice for my subscription this month.
assert:
- type: javascript
value: JSON.parse(output).category === 'billing'
- vars:
ticket: The export button crashes the app and I have a demo in an hour!
assert:
- type: javascript
value: |
const o = JSON.parse(output);
return o.category === 'bug' && o.urgency === 'high';
- vars:
ticket: Would love a dark mode someday, no rush.
assert:
- type: javascript
value: |
const o = JSON.parse(output);
return o.category === 'feature_request' && o.urgency === 'low';
defaultTest applies the is-json assertion to every case, so each test checks parseability plus its own field-level expectations. Deterministic assertions (is-json, contains, regex, javascript) are free and fast; save model-graded ones like llm-rubric for later.
4. Run the suite locally #
export OPENAI_API_KEY=sk-your-key
promptfoo eval
eval runs every prompt Γ provider Γ test combination and exits with code 100 if any test fails (code 1 is reserved for operational errors like bad config). That exit code is the whole CI trick: a failing eval fails the job, same as a broken unit test. Run promptfoo view to inspect results in a local web UI.
5. Add the GitHub Actions workflow #
Store the API key as a repo secret first:
gh secret set OPENAI_API_KEY
Then create .github/workflows/promptfoo.yml using the official promptfoo-action:
name: Prompt evaluation
on:
pull_request:
paths:
- 'prompts/**'
- 'promptfooconfig.yaml'
jobs:
evaluate:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v4
- name: Cache promptfoo results
uses: actions/cache@v4
with:
path: .promptfoo-cache
key: ${{ runner.os }}-promptfoo-${{ hashFiles('promptfooconfig.yaml', 'prompts/**') }}
restore-keys: |
${{ runner.os }}-promptfoo-
- name: Run promptfoo eval
uses: promptfoo/promptfoo-action@v1
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
openai-api-key: ${{ secrets.OPENAI_API_KEY }}
config: promptfooconfig.yaml
prompts: 'prompts/**'
cache-path: .promptfoo-cache
fail-on-threshold: 100
Three details matter here. The paths filter means the job only spends API tokens when prompt files or the config actually change. pull-requests: write lets the action post its results comment; without it the eval runs but the comment silently fails. fail-on-threshold: 100 fails the check unless every test passes; lower it (say 90) once your suite is big enough that one flaky case shouldn't block a merge. The cache step keeps identical prompt+input pairs from being re-billed across runs.
Commit everything and push:
git add .
git commit -m "Add promptfoo regression suite"
git push origin main
Verify it works #
Locally, a passing run ends like this (table trimmed):
ββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β ticket β [openai:chat:gpt-5.6-luna] β
ββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββ€
β I was charged twice forβ¦ β [PASS] {"category":"billing"β¦β
ββββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββ
Successes: 3
Failures: 0
Errors: 0
Pass Rate: 100.00%
echo $? prints 0. Now prove the gate catches a regression. Branch off, edit prompts/triage.txt to end with Explain your reasoning, then give the JSON. (a realistic "improvement" that breaks JSON-only output), and open a PR. Within a few minutes you should see the Prompt evaluation check go red, and a bot comment on the PR summarizing pass/fail counts with a link to the before/after comparison in the web viewer. Revert the line and the check goes green.
Troubleshooting #
OpenAI API key is not set Set the OPENAI_API_KEY environment variable or add apiKey to the provider config. Locally, export it in the same shell running eval. In CI, the secret name in secrets.OPENAI_API_KEY must match what you created with gh secret set, and it has to be passed through the openai-api-key action input. Note that model-graded assertions like similar and llm-rubric call OpenAI as the default grader, so this error can appear even when your tested provider is a different vendor.
Resource not accessible by integration in the action logs. The workflow token can't post the PR comment. Add the permissions block from step 5. On PRs from forks, GITHUB_TOKEN is read-only by design; set disable-comment: true for fork PRs and rely on the check status alone.
The check passes instantly without calling the model. The action evaluates prompt files changed in the PR. If your edit only touched promptfooconfig.yaml, or your prompts glob doesn't match the changed paths, nothing gets evaluated. Fix the glob, or set use-config-prompts: true to always use the prompts listed in the config, or force-run: true to evaluate regardless of the diff.
npm warn EBADENGINE Unsupported engine followed by a crash. Your Node is older than 22.22.0. Upgrade to Node 24 LTS, or pin the runner with actions/setup-node before the promptfoo step.
Next steps #
Move the golden dataset out of YAML once it grows: tests: file://tests.csv loads cases from a spreadsheet your support team can edit. Add model-graded checks (llm-rubric, similar, factuality) for prompts without a single right answer, and latency and cost assertions to catch performance regressions, not just quality ones. Since providers live in config, the same suite also gates model upgrades: add openai:chat:gpt-5.6-terra as a second provider and the PR comment becomes a side-by-side model comparison. The CI/CD integration docs cover GitLab, Jenkins, and bare-CLI setups if you're not on GitHub Actions.
Sources & further reading #
-
[Getting started](https://www.promptfoo.dev/docs/getting-started/) β promptfoo.dev -
[GitHub Actions integration](https://www.promptfoo.dev/docs/integrations/github-action/) β promptfoo.dev -
[promptfoo-action README](https://github.com/promptfoo/promptfoo-action) β github.com -
[Assertions and metrics reference](https://www.promptfoo.dev/docs/configuration/expected-outputs/) β promptfoo.dev -
[Command line reference](https://www.promptfoo.dev/docs/usage/command-line/) β promptfoo.dev -
[Troubleshooting](https://www.promptfoo.dev/docs/usage/troubleshooting/) β promptfoo.dev
Priya NairΒ· AI & Developer Experience Writer
Priya covers AI frameworks, developer productivity tooling, and the startup ecosystem across South and Southeast Asia, bringing a researcher's rigour and a practitioner's empathy to every story. She is deeply sceptical of benchmarks and asks hard questions so her readers don't have to.
Discussion 1 #
Does promptfoo handle cost tracking across runs, or do you end up just eating the bill every time someone tweaks a prompt in a PR? Running evals on every commit sounds amazing until you're suddenly explaining the AWS bill to finance.