The summer after my sophomore year I interned at the College Board on a performance engineering team, and stayed on through the school year as a network operations analyst. The internship was where performance testing first clicked for me as its own discipline — not “is the response correct?” but “how does the service hold up when a few hundred thousand students hit it at once?”
We built our performance tests as JMeter test plans and let Jenkins run them headless against each environment. JMeter is purpose-built for this: you model virtual users, shape how they arrive, and get timers, assertions, and percentile reporting out of the box. The same plan you build in the GUI runs unattended in CI on every build.
This post is a from-scratch introduction. By the end you’ll understand how a JMeter test plan is put together, how to shape it into the four scenarios we ran, and how to let Jenkins run and grade them.
The mental model
A JMeter test plan is a tree, and for a first test you only need a few nodes:
- Thread Group — your virtual users. This is where the load shape lives: how many users, how fast they ramp up, how long they run.
- HTTP Request — the actual call each user makes.
- Assertion — what counts as a failure (a bad status, or a response that took too long).
- Listener / results file — where samples get written for reporting.
The key insight that took me a while: the type of performance test — load, spike, soak, endurance — is almost entirely a property of the Thread Group. Same request, same assertions; you’re just changing how the virtual users arrive and how long they stay.
A first test plan
You build the plan in the GUI, but under the hood it’s XML (.jmx), and the Thread
Group is where the interesting knobs are. Rather than hard-code them, read them from
properties so the same plan can be driven from the command line:
<ThreadGroup testname="Scores API" enabled="true">
<stringProp name="ThreadGroup.num_threads">${__P(threads,50)}</stringProp>
<stringProp name="ThreadGroup.ramp_time">${__P(ramp,60)}</stringProp>
<boolProp name="ThreadGroup.scheduler">true</boolProp>
<stringProp name="ThreadGroup.duration">${__P(duration,300)}</stringProp>
<elementProp name="ThreadGroup.main_controller" elementType="LoopController">
<boolProp name="LoopController.continue_forever">false</boolProp>
<intProp name="LoopController.loops">-1</intProp> <!-- loop until scheduler stops it -->
</elementProp>
</ThreadGroup>${__P(threads,50)} reads the threads property, defaulting to 50 if it isn’t set.
With the scheduler on and loops set to -1, each user just keeps making requests
until the duration runs out — so you control the whole run with three numbers.
Two additions make the numbers actually mean something:
- A timer (a Gaussian Random Timer, say). Real users pause between actions; without think time you’re not modeling traffic, you’re measuring how fast your load generator can spin a loop. If you want to hold a fixed request rate regardless of thread count, the Throughput Shaping Timer does that.
- A CSV Data Set Config to feed each request different data — student IDs, search terms — so you’re not hammering one cached row and getting numbers that look great and mean nothing.
The four scenarios
Same plan, different Thread Group settings. What changes is the load shape and the question you’re asking.
| Scenario | Question it answers | How you shape it |
|---|---|---|
| Load | How does it behave at expected/peak concurrency? | steady threads, short ramp, moderate duration — your baseline |
| Spike | Can it absorb a sudden surge and recover? | jump from a low baseline to a burst, then back down |
| Soak | Does it leak or degrade under sustained load? | moderate threads held for hours |
| Endurance | Is it stable at peak over a long stretch? | peak threads held for hours (overnight) |
Load, soak, and endurance are the same Thread Group with different threads and duration — the difference is intent. Soak and endurance overlap a lot; I think of
soak as “moderate load, are we leaking?” and endurance as “peak load, does latency
drift over hours?” Both catch the bugs a five-minute run never will: a connection pool
that slowly starves, a cache that never evicts, a heap that creeps up until the GC
can’t keep pace.
Spike is the one the stock Thread Group can’t really express, because it’s a shape,
not a single thread count. That’s what the JMeter Plugins (jp@gc) custom thread
groups are for — the Ultimate Thread Group lets you lay out the profile directly: hold
10 users for a minute, jump to 200 for a minute, drop back to 10 and watch how long it
takes to recover. The recovery is usually the interesting part.
Running it headless
Never run a load test from the GUI — the GUI is for building and debugging. For an actual run, go non-GUI:
# -n non-GUI, -t the plan, -l raw results, -e -o an HTML dashboard.
jmeter -n -t scores.jmx
-Jthreads=100 -Jramp=60 -Jduration=600
-l results.jtl -e -o dashboard/The -J flags set the properties the plan reads through ${__P(...)}, so one plan
covers load, soak, and endurance just by changing those numbers. results.jtl is the
raw per-sample data; dashboard/ is a browsable report with percentiles and graphs.
Wiring it into Jenkins
This is where it stops being a one-off and becomes something the whole team benefits from. A single parameterized pipeline runs any scenario, grades it, and publishes the report:
pipeline {
agent any
parameters {
choice(name: 'PROFILE', choices: ['load', 'spike', 'soak', 'endurance'])
}
stages {
stage('Run JMeter') {
steps {
script {
// load/soak/endurance share one plan via -J props; spike has
// its own plan because the shape lives in the thread group.
def cmd = [
load: 'jmeter -n -t scores.jmx -Jthreads=100 -Jramp=60 -Jduration=600',
soak: 'jmeter -n -t scores.jmx -Jthreads=50 -Jramp=120 -Jduration=7200',
endurance: 'jmeter -n -t scores.jmx -Jthreads=100 -Jramp=300 -Jduration=28800',
spike: 'jmeter -n -t scores-spike.jmx'
]
sh "${cmd[params.PROFILE]} -l results.jtl -e -o dashboard"
}
}
}
}
post {
always {
// Performance plugin: reads the .jtl, fails/destabilizes the build on
// the error thresholds, and keeps a trend across builds.
perfReport errorFailedThreshold: 1, errorUnstableThreshold: 0, sourceDataFiles: 'results.jtl'
publishHTML target: [reportDir: 'dashboard', reportFiles: 'index.html', reportName: 'JMeter Dashboard']
}
}
}A few things fall out of this nicely:
- Load runs on every merge — it’s short, so it catches regressions early and cheaply.
- Soak and endurance run on a schedule (
triggers { cron('H 2 * * *') }), overnight, because nobody wants to block a PR for eight hours. - The Performance plugin gives you a trend graph across builds. Watching p95 drift up over a month is often how you catch a slow leak before it pages someone — and never grade on the average, since averages hide exactly the slow requests you care about.
- Keep the target environment fixed. If staging gets resized between runs, your numbers are noise. Consistent hardware is what makes the trend line mean anything.
Trade-offs
A single JMeter box only generates so much load before the load generator itself
becomes the bottleneck — at that point you move to distributed mode (one controller
driving several worker nodes) or a cloud runner. But for getting a team
performance-aware inside the CI they already have, with real pass/fail thresholds on
every build, a handful of .jmx plans and a Jenkins job go a long way. It’s how I got
introduced to all of this, and it’s still where I’d tell someone to start.