# How Fyxer Used AI to Run 360 Growth Tests With Just Four Engineers

> Source: <https://industrycontents.com/ai-experimentation-workflow-fyxer/>
> Published: 2026-10-07 16:00:00+00:00

7 min read

Fyxer’s AI reportedly cut test research from days to hours. Two published totals still disagree.

A growth engineer at [Fyxer](https://www.fyxer.com/) could once spend days researching an idea and a week getting a test into production. By 2025, the UK company had pushed parts of that work into a shared AI experimentation workflow. Research could take hours. Some builds could fit into an afternoon.

The question was practical. Fyxer, which sells an AI assistant for email and meetings, was growing quickly and had more growth ideas than four engineers could test. Could the team increase experiment velocity without adding a matching layer of process and headcount?

A [case study published by experimentation vendor GrowthBook](https://www.growthbook.io/blog/how-a-team-of-4-used-a-b-testing-to-help-fyxer-grow-from-1m-to-35m-arr-in-1-year) says the answer was a system that connected AI coding tools, shared instructions, internal data and the experimentation platform. GrowthBook, which sells A/B testing software, reports that the company ran 541 experiments in 12 months, with the four-person growth engineering team responsible for 360 of them. Fyxer growth engineer [Kameron Tanseli gave a lower company total of 514](https://www.linkedin.com/posts/kameron-tanseli_in-2025-fyxer-ran-514-experiments-thats-activity-7425498675856142336-yFdj) in his own account and confirmed the team’s 360 tests.

## The backlog was the growth constraint

Fyxer had plenty of hypotheses. The gap sat between an idea and a trustworthy readout. GrowthBook describes a familiar sequence. Engineers research the user problem, inspect data, write a specification, build the variant, wire the test, check the results, remove stale code and document what happened. Any one step can leave a promising idea waiting in a queue.

The company skipped the shortcut of a single prompt that generates finished tests. It worked at making each handoff legible to AI tools. Shared Claude skills captured recurring instructions. Model Context Protocol (MCP) connections, an open standard that lets AI tools reach other software, exposed internal systems and GrowthBook’s API. Engineers used Cursor, an AI code editor, to preview changes, and simple experiments could sometimes be built in one pass with Claude Opus, Codex or Tembo.

That distinction matters. The central experiment was an operating one. Give AI enough context to handle the repeatable parts of AI A/B testing, then keep engineers responsible for the product decision and release. Kuaishou tried a related setup when it [let an AI agent run growth tests](https://industrycontents.com/kuaishou-a-b-agent-growth-tests/), and its first two ideas hurt the platform.

## Fyxer turned the loop into reusable AI work

The workflow covered more than code generation. GrowthBook says a tool called Dot acted as an AI analyst over BigQuery, Google’s cloud data warehouse, and Slack. Cursor automations prepared pull requests, found stale experiment code and updated documentation. Claude could query experiment metrics through the GrowthBook API when a number moved unexpectedly.

Here sits the structural lesson for a growth team. The useful unit of automation was a chain of defined tasks with known data sources and outputs. Research produced a sharper brief. The brief fed the build. The build carried an experiment assignment. The result flowed back into analysis and cleanup.

GrowthBook says this cut research from days to hours and development from roughly a week to an afternoon. Those are vendor-published estimates from outside a disclosed time study. Even so, they name the metric the team wanted to move. That metric was time to a decision.

## Four engineers owned separate growth verticals

Fyxer paired the tooling with a clear ownership model. Tanseli says each of the four growth engineers ran a separate vertical and owned discovery, prioritisation and delivery. That cut the coordination cost of moving a test through several specialist teams.

The team could reportedly keep five or six experiments running in parallel. Its 360 tests work out to 90 per engineer across the year. Throughput alone leaves rigour unproven, yet it shows why reusable context and automated experiment analysis mattered. Without them, extra test volume would have piled up analysis, cleanup and documentation debt.

GrowthBook reports a 25 percent win rate. The case study also lists several business outcomes. Free-to-paid conversion rose from 5 percent to 35 percent after a credit-card gate, annual plan selection grew 2.3 times, trial starts from personal email addresses climbed 65 percent and 33 percent of referral invitations were accepted. The source leaves out sample sizes, confidence intervals and full test windows for these examples, so treat them as signals that nobody has independently audited.

## Test volume exposed bad ideas quickly

A high-tempo programme earns its keep when it kills weak ideas as efficiently as it finds winners. The published account describes a team that ran several hypotheses through one shared instrumentation and decision process. Each experiment avoided the overhead of a custom project.

That framing serves readers better than the headline ARR growth. ARR, or annual recurring revenue, is the yearly value of a company’s subscriptions. GrowthBook connects the experimentation programme with Fyxer’s reported move from $1 million to $35 million in ARR. The evidence cannot separate how much of that increase came from experiments, market demand, sales, pricing or other factors. It supports a claim about testing capacity and stops short of a causal claim about revenue.

The same caution applies to the 25 percent win rate. A win rate depends on what enters the backlog, how a win is defined and whether guardrail metrics can overturn a local lift. The public material documents none of those definitions in full.

## The published totals disagree

The strongest sourcing problem sits in the top-line count. GrowthBook’s April 2026 case study says Fyxer ran 541 experiments in 12 months. Tanseli’s first-hand LinkedIn post says 514 experiments in 2025. Both say the growth engineering team ran 360.

The gap may reflect a date window, a correction or a simple transposition. Neither source explains it. The defensible number for this story is therefore the team total both accounts share, and it is also the number most relevant to the operating lesson.

One more disclosure belongs here. GrowthBook sells the experimentation software used in the programme and published the detailed case. Tanseli’s post gives useful first-hand corroboration for the staffing, workflow and team count, yet nobody has independently audited the experiments or outcomes.

## Run a bottleneck audit before chasing a test count

A comparable company can try the most portable lever without a machine learning team. Start by timing one experiment from idea to decision. Break the elapsed time into research, data access, specification, implementation, quality assurance, analysis, cleanup and documentation. Pick the slowest repeatable step and give an AI tool a narrow job, a controlled data source and a required output.

For many teams, the first useful move will be more modest than Fyxer’s stack. It could be a shared instruction file that produces experiment briefs from one template, a read-only connection that answers standard metric questions, or an automation that flags tests whose code should be removed. The growth metric is cycle time. The guardrail asks whether the team can still explain assignments, exclusions and decisions without asking the model to reconstruct them.

Fyxer’s case also carries conditions that make the result hard to copy. Engineers owned end-to-end verticals, the company had a sizeable flow of product traffic, internal data was accessible to AI tools and the team was willing to maintain shared context and integrations. A company with fragmented ownership or unreliable event data would automate the ambiguity along with the work.

The deeper point is that an AI experimentation workflow compounds only when it leaves the system easier to understand after each test. Faster code is useful. Faster learning depends on whether the next person can see what was tested, which metric moved and why the team made its decision.

### Sources

- [Fyxer](https://www.fyxer.com/)
- [GrowthBook case study, published April 4, 2026](https://www.growthbook.io/blog/how-a-team-of-4-used-a-b-testing-to-help-fyxer-grow-from-1m-to-35m-arr-in-1-year)
- [Kameron Tanseli’s first-hand LinkedIn account](https://www.linkedin.com/posts/kameron-tanseli_in-2025-fyxer-ran-514-experiments-thats-activity-7425498675856142336-yFdj)
- [Kameron Tanseli](https://kamrn.com/)

Compare the evidence and operating lessons from more AI experiments in our [Growth Signal Index](https://industrycontents.com/ic-lab/growth-signal-index/), or browse benchmark tooling in the [Industry Contents Lab](https://industrycontents.com/ic-lab/).
