Learn Labs
9. The Trouble with Distributed Systems

9.4 Process Pauses

The broken leader-lease loop:

while (true) {
    request = getIncomingRequest();
    // Ensure that the lease always has at least 10 seconds remaining
    if (lease.expiryTimeMillis - System.currentTimeMillis() < 10000) {
        lease = lease.renew();
    }
    if (lease.isValid()) {
        process(request);     // ← what if the thread pauses for 15 s HERE?
    }
}

Two bugs:

  1. It relies on SYNCHRONIZED CLOCKS: the expiry time was set by a different machine and is compared to the local system clock. If clocks are out of sync by more than a few seconds, the code will start doing strange things.
  2. Even using only the local monotonic clock: the code assumes very little time passes between checking the time and processing the request. If the thread stops for 15 seconds around lease.isValid, the lease will likely have expired by the time the request is processed, and another node will already have taken over. THERE IS NOTHING TO TELL THIS THREAD IT WAS PAUSED, so it won't notice until the next loop iteration — BY WHICH TIME IT MAY HAVE ALREADY DONE SOMETHING UNSAFE.

4.1 Nine reasons a thread pauses for a long time

  1. LOCK/QUEUE contention for a shared resource. “Often Worse on machines with more CPU cores, and contention problems can be difficult to diagnose.”
  2. Garbage collection. “Stop-the-world” pauses Sometimes lasted several minutes in the past; still noticeable with modern GCs.
  3. Vm SUSPEND/RESUME. Pausing all processes and saving memory to disk. “Can occur At any time and last for An arbitrary length of time.” Used for Live migration of VMs between hosts without a reboot; the pause length depends on the rate at which processes are writing to memory.
  4. End-user devices. Laptops and phones suspend and resume arbitrarily — e.g. when the user Closes the lid.
  5. Context switches. The OS switching to another thread, or the hypervisor to a different VM, can pause the running thread At any arbitrary point in the code. CPU time spent in other VMs is Steal time. Under heavy load, it may take some time before the paused thread runs again.
  6. Synchronous disk I/O. “In many languages, disk access can happen Surprisingly, even if the code doesn’t explicitly mention file access — for example, the Java classloader lazily loads class files when first used, which could happen At any time.” I/O pauses and GC pauses May even conspire to combine their delays. And if the disk is network-attached (EBS), I/O latency is subject to Network variability.
  7. Page faults / swapping. A simple memory access may fault in a page from disk. Under memory pressure this may require swapping a different page Out. In the extreme, the OS spends most of its time swapping and Gets little actual work done — thrashing. “To avoid this, Paging is often disabled on server machines — if you would rather Kill a process to free memory than risk thrashing.”
  8. Sigstop. Ctrl-Z in a shell. Immediately stops the process from getting CPU cycles until SIGCONT. “Even if your environment does not normally use SIGSTOP, It might be sent accidentally by an operations engineer.”

The problem is similar to making multithreaded code on a single machine thread-safe: YOU CAN'T ASSUME ANYTHING ABOUT TIMING. But the tools we have for that — mutexes, semaphores, atomic counters, lock-free data structures, blocking queues — DON'T DIRECTLY TRANSLATE, because a distributed system has NO SHARED MEMORY, ONLY MESSAGES SENT OVER AN UNRELIABLE NETWORK.

A node must assume its execution can be paused for a significant length of time AT ANY POINT, EVEN IN THE MIDDLE OF A FUNCTION. During the pause, the rest of the world keeps moving and may even declare the paused node dead. Eventually the node may continue running, WITHOUT EVEN NOTICING THAT IT WAS ASLEEP until it checks its clock sometime later.

4.2 Could we eliminate pauses? — hard real-time systems

Where failure to respond by a deadline causes serious damage: computers controlling aircraft, rockets, robots, cars. "If your car's onboard sensors detect a crash, you wouldn't want the airbag release to be delayed because of an inopportune GC pause."

⚠️ In embedded systems, REAL-TIME means the system is carefully designed and tested to meet specified timing guarantees IN ALL CIRCUMSTANCES. This is in contrast to the vaguer use of "real-time" on the web (servers pushing data to clients, stream processing).

What real-time requires at every level:

  • A real-time operating system (RTOS) scheduling processes with a guaranteed allocation of CPU time in specified intervals
  • Library functions must document their WORST-CASE EXECUTION TIMES
  • Dynamic memory allocation may be restricted or disallowed entirely (real-time GCs exist, but the application must not give the collector too much work)
  • An enormous amount of testing and measurement

This severely restricts the range of programming languages, libraries, and tools. Developing real-time systems is VERY EXPENSIVE, and they are most commonly used in safety-critical embedded devices.

Also: "REAL-TIME" IS NOT THE SAME AS "HIGH-PERFORMANCE" — in fact, real-time systems may have LOWER THROUGHPUT, since they prioritize timely responses above all else.

For most server-side data processing systems, real-time guarantees are simply NOT ECONOMICAL OR APPROPRIATE. Consequently, these systems must suffer the pauses and clock instability that come from operating in a non-real-time environment.

4.3 Limiting the impact of garbage collection

GC has genuinely improved: "A properly tuned collector will now usually pause processes for NO MORE THAN A FEW MILLISECONDS." Java offers CMS, G1, ZGC, Epsilon, Shenandoah, each optimized for different memory profiles; Go offers a simpler concurrent mark-and-sweep collector that attempts to optimize itself.

Four mitigation strategies:

StrategyDetail
Use a language without a GCSwift (automatic reference counting), Rust and Mojo (object lifetimes tracked via the type system, so the compiler determines how long memory must be allocated)
Reduce garbagePool and reuse objects rather than discarding them; allocate data OFF-HEAP
Treat GC as a planned outage ✔If the runtime can warn the application that a GC pause is coming, the application STOPS SENDING NEW REQUESTS to that node, waits for outstanding requests to finish, and THEN performs the GC while no requests are in progress. This HIDES GC PAUSES FROM CLIENTS and reduces high percentiles of response time
Restart before a full GCUse the GC only for SHORT-LIVED objects (fast to collect) and RESTART PROCESSES PERIODICALLY before they accumulate enough long-lived objects to require a full GC. One node at a time, with traffic shifted away first — as in a ROLLING UPGRADE

On this page