Insights / Technical notes

Investigating Minecraft Lag: Quiet Metrics and Shared-Storage Diagnosis

From quiet TPS/MSPT collection to correlating JFR and OS I/O observations. This article separates a completed cause investigation from performance improvements that have not been tested.

  • Minecraft
  • Monitoring
  • Performance
Investigating Minecraft Lag: Quiet Metrics and Shared-Storage Diagnosis
Table of contents
  1. Capture the first profile while lag is occurring
  2. Measure averages and brief pauses separately
  3. Collect quietly, and record missing data
  4. Correlate bounded JVM, OS, and save observations
  5. Treat the next step as a separate test

Minecraft lag can come from different sources: server tick processing, brief save stalls, networking, or client rendering. This anonymized investigation covers several Paper servers without revealing internal hostnames or configuration.

Capture the first profile while lag is occurring

Record the reported lag time, player count, and whether saving was happening; compare the profile with MSPT from that period. Take a short normal-period sample too. Distinguishing more processing from more waiting helps choose between plugin tuning and storage investigation.

PaperMC:Profiling during the problem

Measure averages and brief pauses separately

Record TPS, average, maximum and p95 MSPT, and the cumulative number of slow ticks. A good average can hide brief stalls. Host-wide CPU usage alone cannot distinguish single-threaded work from I/O wait.

Collect quietly, and record missing data

In this case, a small plugin wrote measurements from Paper’s public API to local JSON, and a collector aggregated them into monitoring time series. File I/O ran outside the game’s main thread; replacing a temporary file avoided readers seeing a partial update. This did not suppress ordinary console logs.

Check timestamps and whether collection succeeded. Do not treat an old value from a stopped collector as healthy, or replace missing data with zero TPS.

Correlate bounded JVM, OS, and save observations

While the issue was reproducible, JFR, OS wait observations, and block I/O were collected in separate, bounded windows. Those observations were compared with continuous MSPT data and periodic I/O pressure to distinguish GC activity from waits during synchronous saves. The captures were not all taken at the same time. For a general Paper starting point, see the official spark profiling guide. Here, JFR and OS observations were added to investigate waits during saves.

Across a normal 240-second window and a stalled 220-second window, per-second write-await medians were 1.60 ms and 43.71 ms; the maxima were 6.00 ms and 123.57 ms. These are two observation windows, not before-and-after results or a general benchmark.

Along with synchronous saves and filesystem journal waits, delays also appeared in device requests from multiple applications, narrowing the candidate to the shared-storage path. A completion event in a Linux block tracepoint may represent only part of a request, so unmatched requests were not mixed into a statistic for all I/O. Capture duration and volume were bounded, and measurement overhead was considered.

Separate observations from hypotheses Continuous metrics and bounded samples use different observation windows. Storage-migration effects remain untested.
  1. Quiet continuous metrics Record TPS/MSPT and missing data without adding routine console output.
  2. Bounded investigation Collect JFR, OS, and block I/O separately and correlate them. Shared storage is a candidate cause, not a confirmed fault.

Treat the next step as a separate test

After the investigation, bounded tracing ended and always-on collection was confirmed healthy. A comparison over several cycles under similar use after moving storage has not been run. This record prepares a follow-up with backup, restore, and rollback plans; it does not reduce save durability or assign blame to GC or a particular plugin without evidence.

For the distinction between monitoring signals and diagnosis, see OpenClaw monitoring and incident investigation. For recovery testing, see R2 and restic backup monitoring.