A performance number without its conditions is difficult to use. Ten milliseconds may be excellent or unacceptable; an average can hide the slow requests that matter; a local result may say little about production work. The first task is therefore not selecting a profiler. It is defining the observation.

Start with a question

Frame the work as a comparison that could change a decision. For example: does parsing dominate the time for this input size? Does the new cache reduce repeated reads after warm-up? Does memory return to a stable level after the batch completes?

A good question provides a boundary. It suggests which signal to collect and, just as importantly, which signals can wait. This keeps measurement from becoming a tour of every number a tool can display.

Preserve the conditions

Record the input, software revision, machine conditions, run count, and whether the system was cold or warm. These notes do not need to be elaborate. They need to be sufficient for the next result to mean the same thing.

Run enough times to see the shape of the variation. If the spread is wide, investigate the spread before celebrating a small improvement in the center. Noise is not an inconvenience to average away; it may be the most informative part of the result.

Use a repeatable measurement record

A compact record makes a result comparable without turning every experiment into a report. Keep it beside the code or change that motivated the work.

Question
The decision this comparison is meant to inform
Baseline
The current behavior and the exact revision that produced it
Input
Size, shape, source, and any sanitization or sampling
Environment
Machine, operating mode, important limits, and competing work
Preparation
Cold or warm state, cache handling, and the warm-up procedure
Runs
Sample count, order, and whether baseline and candidate were interleaved
Metric
The unit, aggregation, and threshold that would change the decision
Raw result
The individual observations or a file that preserves them
Conclusion
What changed, under which conditions, and what remains unknown

Read the distribution before the headline

For a small number of runs, list the individual values. A median, range, and sample count usually communicate more than a heavily rounded average. Percentiles become useful when there are enough observations to support them; a p99 from a few dozen requests mostly reports one sample with an impressive label.

For latency, keep the center and the tail separate. For throughput, record the duration and concurrency that produced the rate. For memory, distinguish peak allocation from the level that remains after work completes. For CPU, state whether the number is process time, one-core utilization, or whole-machine utilization.

  • Wide spread: look for warm-up, background work, throttling, garbage collection, retries, or mixed input classes.
  • Very small gain: compare the improvement with run-to-run variation before calling it real.
  • Only one metric improves: check whether time moved into memory, network work, startup, or a slower tail.
  • Production differs: compare inputs and conditions first; do not immediately add more local tuning.

Keep the raw observation beside the conclusion. A result that cannot be reinterpreted later has a short useful life.

Change one thing

Make the smallest change that tests the current explanation, then repeat the same measurement. Large rewrites can improve a number while making it impossible to know why. A narrow change produces knowledge even when it fails.

The final note should state what improved, under which conditions, and what tradeoff was introduced. “Faster” is incomplete. “Reduced the warm-path median while leaving tail latency unchanged” is a result that another decision can build on.

Before accepting a result

  • The question names a decision, not merely a metric to improve.
  • Baseline and candidate use the same input, preparation, and observation window.
  • The result is larger than the unexplained variation—or is explicitly reported as inconclusive.
  • Raw observations remain available beside the summary.
  • The conclusion names the conditions where it applies and one important condition not tested.
  • A second run after a restart or clean setup reproduces the direction of the result.