Cloud performance testing gives you what a laptop can’t: load spread across many machines, a clean baseline, and results your whole team can see. You set the virtual user (VU) count. Postman handles the fleet, the tuning, and the merging, and you get one set of numbers you can trust.
Most teams still start performance testing on a laptop, and that’s fine. This post covers what the cloud adds, why that matters more as AI agents become some of your heaviest API consumers, and how to put it in your pipeline.
Why cloud performance testing beats your laptop #
On a local run, the machine running the performance test becomes part of the test. Your CPU, Wi-Fi, VPN, and uplink all sit between the load and the API. When numbers look bad, you can’t tell if the API is slow or your laptop is. A laptop also runs out of room: pushing more load means raising OS limits and tuning TCP, which can have side effects and may need admin rights you don’t have. And two runs on two days, on two networks, aren’t a fair comparison.
A cloud run sends traffic from Postman’s infrastructure instead. You get:
- A clean baseline. The load generator isn’t competing with your email client.
- A real network path. Requests arrive over the public internet, the way remote users and agents reach you.
- Repeatability. The same runner setup every time, so a change in results means a change in your API.
- Scale. Postman’s cloud ramps to millions of virtual users. A one-hour soak at 2,000 concurrent users isn’t a stretch.
- Shared results. Every run, including runs started from CI, lands in Postman with its history and trend. Your team looks at the same numbers.
Load comes from the region of your Postman account: the US by default, or the EU on EU Data Residency plans.
What Postman’s cloud takes care of #
Cloud runs use the same collection you run locally. The difference is the work you no longer do. (For how it works under the hood, see Performance Testing in the Cloud with Postman: What We Built.) Here’s what you own if you build the setup yourself, and what you don’t with Postman.
Scaling out On your own: you size the machines, set up a fleet, split the virtual users across them, and get every machine to start at the same time. With Postman: you set --vu-count. Postman picks the machines and spreads the load across them.
Machine tuning On your own: every load generator needs higher file-descriptor and port limits and tuned TCP settings. Those go into the image and have to stay the same across the fleet. On a laptop or a locked-down machine, you may not be allowed to change them at all. With Postman: the load generators come ready.
Merging results On your own: you collect the output from every machine and combine it. Correct percentiles take extra work, and you need a dashboard to show them. With Postman: metrics from every machine come back as one run, with live and final latency percentiles, throughput, and error rate.
One test, one tool On your own: load scripts live in a separate tool, apart from the API collections your team already maintains. With Postman: the collection, auth flow, dataset, and pass condition you use for local performance testing run in the cloud and from CI. No rewrite. A few things stay local-only: data files, mTLS client certificates, and Local Vault secrets. For cloud runs, use a dataset and Shared Vault secrets instead.
Setup and teardown On your own: you write separate scripts to seed and reset test data around the load test. With Postman: --setup-collection runs once before the load starts and --teardown-collection runs once after it ends, whatever the outcome. They run as their own stages, outside the load phase.
Fixed egress IPs On your own: you run NAT gateways or reserved IPs, and you update allowlists whenever the fleet changes. With Postman: on Enterprise, --runner postman-cloud-static-ip sends load from a dedicated cluster with a fixed egress IP that you allowlist once.
Nothing to keep running On your own: you patch images, upgrade the test runner, shut machines down after each test, and pay for capacity that sits idle between tests. With Postman: you pay only for the VU-hours you use.
Why AI agents make performance testing matter more #
Human traffic has a shape. People click, read, and click again. Agent traffic doesn’t.
Agents are fast becoming a main audience for APIs, and they don’t behave like people. An agent working on one task can call your API many times in a row. It can call several endpoints at once. It doesn’t take breaks, and it doesn’t keep office hours. One user request can turn into dozens of API calls, and many users doing this at once can produce load your capacity planning never saw.
That makes the shape of the load as important as its size. Postman performance testing has four load profiles (–load-profile), and each one maps to a question about agent traffic:
- fixed: Can the API hold steady concurrency over time? The VU count stays constant. This is your baseline for agents that run all day.
- ramp-up: Where does it start to degrade as more agents come online? VUs climb from 25% to 100%, then hold. Use it to find the knee in the curve before adoption finds it for you.
- spike: What happens when a lot of agents start at once, such as a scheduled job firing or a launch going out? VUs start at 10%, jump to 100%, then drop back to 10%.
- peak: Can it survive staying near maximum load? VUs climb from 20% to 100%, hold, then come back down to 20%. Agents don’t go home, so peak load can last for hours.
Each profile is a fixed pattern across the run’s duration, which you set in minutes. The profiles are VU-based: rps is something you can gate on, not a rate you drive load at. If your team talks about load tests and stress tests, Performance testing vs. load testing vs. stress testing explains how the terms relate.
Agents also don’t forgive slow or flaky responses. A person waits. An agent times out, retries, or picks another path. So your p95 and your error rate become product behavior.
Why does the cloud matter here? Agent load can be large, steady, and long. A laptop hits its own limits first. A cloud run spreads the load across machines and keeps going for as long as the test needs.
One thing to plan for. Cloud runs come from Postman’s IP ranges. If your API allowlists by IP, or rate limits per IP, use the static-IP runner (Enterprise) and allowlist it. Otherwise your performance test measures your firewall instead of your API.
Make performance tests realistic with datasets #
Sending the same request a thousand times is misleading: caches answer fast and hot rows stay in memory.
Attach a dataset so each virtual user sends different real inputs. Datasets work the same way wherever the run happens, which makes them the right choice for cloud and CI runs. You can connect a dataset to a live MySQL, PostgreSQL, or SQL Server database on Team and Enterprise plans, so you’re not exporting stale CSVs. Custom JDBC sources need Enterprise. Use a test or staging copy, or a read-only account with anonymized data.
For agent traffic, include what agents actually send: long prompts, large payloads, unusual parameter combinations, and a few malformed requests.
How do I add performance testing to CI with a p95 gate? #
A performance test you run by hand gets run before big launches and not much else. A test in your pipeline runs every time.
Cloud runs fit CI well. A CI runner is small and shared, so it can’t generate much load, and its numbers move with whatever else is running on it. A cloud run takes the load off the runner. The build only waits for the result.
If you haven’t set up the CLI yet, start with [Working with the Postman CLI](https://blog.postman.com/working-with-the-postman-cli/). Then sign in with an API key, add a pass condition, and the test becomes a gate:
`postman login --with-api-key "$POSTMAN_API_KEY"`
`postman performance run <collectionId> \ --runner postman-cloud \ --load-profile ramp-up \ --vu-count 50 \ --duration 10 \ --pass-if "less_than(p95, 500)" \ --output ndjson`
This runs the collection from Postman’s cloud with 50 virtual users for 10 minutes. It exits with code 1, and fails the job, if p95 latency goes over 500 ms. The check happens after the run, so a bad build still sends its full load before the gate fails. Start with a low VU count and a short duration on anything live.
You can gate on avg, p90, p95, p99, error_rate, or rps, with less_than, less_than_eq, greater_than, or greater_than_eq. –output ndjson streams results as newline-delimited JSON, which suits CI logs and coding agents where the terminal dashboard can’t render.
When local performance testing is still the right call #
Local performance testing is fast and free on every plan, and it’s the better choice when you’re building the test itself, when the API is on localhost or behind a VPN or firewall, when it needs mTLS, or when you’re checking a dev build for obvious problems like a slow query or a memory leak. Don’t pay for more than you need.
Start local. Move to the cloud when the answer needs to be one you’d bet a launch on.
When is a script-first tool a better fit? #
Postman is the shortest path when your API tests already live in Postman collections and you want performance testing without rewriting them. A code-first load tool may fit better if you need:
- Second-by-second ramp shaping, such as 50 to 3,000 clients in exactly 30 seconds.
- An arrival-rate model that drives a target requests-per-second instead of a VU count.
- The load test itself, not just the pipeline step, to be a plain script file in your repo.
Try it #
Take a collection you already run locally. Attach a dataset. Run it once locally and once in the cloud, then compare the two. The gap is what your laptop was hiding. Once you trust the numbers, add the --pass-if line to your pipeline.
FAQ #
Can Postman run load tests from the cloud, not just my machine? Yes. postman performance run <collectionId> --runner postman-cloud sends load from Postman’s managed cloud infrastructure. Without --runner, load comes from the machine that runs the command.
Does Postman performance testing need the desktop app? No. The Postman CLI runs performance tests headlessly in any CI system. The desktop app is one way to configure a test, not a requirement for running it.
How many virtual users can Postman generate? Postman’s cloud ramps to millions of virtual users. Cloud runs need at least 10. A local run is limited by the machine it runs on.
Can I fail a CI build when p95 latency is too high? Yes. Add --pass-if "less_than(p95, 500)". The command exits with code 1 when the condition isn’t met, which fails the CI job. This works for local and cloud runs.
Is Postman performance testing free? Local performance runs are unlimited on every plan, including Free. Cloud runs cost $0.04 per VU-hour on Solo, Team, and Enterprise, with pay-as-you-go turned on.
Can I allowlist Postman’s load generators in my firewall? Yes, on Enterprise. –runner postman-cloud-static-ip sends load from a static-IP cluster in your account’s region (US, or EU on EU Data Residency plans).
Where do cloud runs send load from? From your Postman account’s region: the US by default, or the EU on EU Data Residency plans. Each run uses one runner, so you can’t combine local and cloud load in the same run.