A load average of 15 on a server whose CPU is basically idle isn't a broken sensor, it's a number most people misread. This one walks through what load average is actually counting, when the kernel's OOM killer steps in and how it scores a victim, why a container can get killed by its own memory limit while the host still has gigabytes free, and which of the three usual fixes - swap, container limits, upgrading RAM - solves the problem versus which one just postpones it.
A load average of 18 on a box whose CPU sits nearly idle gets misread constantly: the number looks big, so the instinct is "must need more vCPUs," and that's the first thing people buy. Half of those cases have nothing to do with cores. They get fixed by an extra gigabyte of RAM or a memory limit on a container that got out of hand. This piece stays on memory specifically: what load average is actually counting, when the OOM killer (the kernel mechanism that force-kills processes so the whole system doesn't lock up with zero memory left) steps in, and how to tell its work apart from a plain CPU shortage. If your case looks more like disk, network, or CPU steal, start with the general triage in "Diagnosing a Slow Linux Server: Where to Start" - this article covers memory only, and goes deeper on it than that one does.
Short version. Load average isn't a CPU percentage - it's an exponentially weighted average of how many processes are running on the CPU, waiting for it, or stuck in state D (uninterruptible sleep), usually from disk or swap. The OOM killer fires once the kernel can't free enough memory even after dropping the page cache: either RAM and swap on the whole host are both gone, or a container or systemd unit with a memory limit hits its own memory.max while the host still has plenty free. The victim gets picked by a badness score computed fresh each time, not by some fixed ranking. Proof lives in dmesg or journalctl -k: "Out of memory" for a host-wide kill, "Memory cgroup out of memory" for a container - two different events, worth not confusing. Swap buys time before the kill; container limits contain the blast radius; upgrading RAM is the only fix if the shortage is chronic rather than a one-off.
The three numbers uptime prints look like a CPU percentage, but they aren't one. They're exponentially damped moving averages of how many tasks sat in the scheduler's run queue over the last 1, 5, and 15 minutes, and Linux puts two very different kinds of tasks in that queue, unlike BSD. One kind is actually running or waiting for a free core. The other is stuck in state D (uninterruptible sleep) - a system call that can't be interrupted until it finishes, typically a disk read, a network filesystem call, or a swap operation.
See it yourself:
uptime
w shows the same numbers plus who's logged in and what they're running right now - useful when you need to know who kicked something off.The common mistake follows directly from this: comparing load straight against the core count and calling it "the CPU is maxed out." That's only true when the queue is full of tasks actually ready to compute. A load of 20 on an idle CPU isn't a broken sensor; it's twenty processes stuck waiting on something slow, not queued for a core. The first check is the %Cpu(s) line in top: high id (idle) alongside high wa (I/O wait) rules the processor out entirely.
Disk isn't the only thing that parks processes in state D. Memory shortage does it too: once RAM runs low, the kernel starts paging (shuffling memory pages between RAM and disk), and every memory access a process makes turns into a disk operation instead of an instant RAM read. The process technically "waits on disk," but the actual cause sits in memory.
vmstat tells the two apart in one shot, since it shows both the CPU queue and swap activity:
vmstat 1 5
vmstat -y 1 5.r is the number of processes genuinely ready to run right now. Sustained above your core count (nproc) means real CPU contention.b is the number of processes in state D. High b alongside low r isn't a CPU story - find what they're waiting on.si and so are swap-in and swap-out speed in KB/s. Zero means no swap traffic right now; sustained nonzero values mean active paging, which is usually what turns "a bit slow" into "feels completely frozen."Then check memory directly:
free -h
Look at available, not free - it's the memory that can actually go to new processes without touching swap, including the part of buff/cache (the filesystem cache) the kernel will hand back on demand. available near zero alongside nonzero si/so is thrashing, not a CPU shortage. Swap can be occupied even without active traffic right now; check that separately:
swapon --show
Empty output means there's no swap configured at all, and a memory shortage with no swap to fall back on goes straight to the OOM killer.
The OOM killer is part of the Linux kernel, not a separate service. It steps in when some process needs memory and the kernel genuinely can't find any: the page cache is already dropped, everything that could be paged out already has been, and swap is either full or doesn't exist. At that point the kernel has exactly one move left - kill one or more processes outright to free memory before the whole system locks up.
It doesn't pick a victim alphabetically or by "newest process first." For every process, at the moment memory runs out, the kernel computes a badness score - roughly the share of system memory that process holds, adjusted by oom_score_adj (a range from -1000, which excludes a process from consideration entirely, to +1000, which makes it the top candidate). Whatever has the highest badness right then gets killed - the score is recalculated fresh at every OOM event, not read off some fixed ranking computed earlier. Check current values for a specific process with:
cat /proc/<pid>/oom_score
cat /proc/<pid>/oom_score_adj
In practice, a heavy database process with a large memory footprint is a frequent target on sheer size alone, even when the actual leak is happening in a much smaller neighboring process. That's one reason not to trust "it killed the thing that was actually causing the problem" on faith, without the check covered next.
The classic symptom: a process just disappears, with no error in its own logs. It never got the chance to write one, since SIGKILL gives it no warning and no cleanup. Look for proof in the kernel's log, not the application's:
sudo dmesg -T | grep -i "killed process"
Here's a real line, captured September 25, 2026 by deliberately triggering the OOM killer on a HIP test server (“stress-ng --vm 1 --vm-bytes 2500M” with no swap) - the structure matches what secondary sources predicted; your numbers and process name will differ:
Out of memory: Killed process 1068570 (stress-ng-vm) total-vm:2620784kB, anon-rss:1237080kB, file-rss:512kB, shmem-rss:0kB, UID:0 pgtables:2504kB oom_score_adj:1000
1068570 (stress-ng-vm) is the killed process's PID and name, timestamped by -T.total-vm is the process's entire virtual address space, usually far bigger than what it actually occupied in RAM - not the number that matters here.anon-rss is the memory the process genuinely held in RAM (anonymous pages, not memory-mapped files) - the closer proxy for "how much it actually used."oom_score_adj is the same value that shifted this process's odds of being picked.dmesg's buffer is size-limited and clears on reboot. If the server has already restarted since the kill, look in the systemd journal instead, which survives a reboot when persistent logging is on:
journalctl -k -b | grep -i "out of memory"
-b limits output to the current boot; -b -1 checks the previous one if you're hunting for something that happened before the last reboot.
This is where developers, more than classic sysadmins, tend to get tripped up. A cgroup (control group) is a kernel mechanism that caps how much CPU, memory, and other resources a group of processes can use; every Docker container and every systemd unit with a defined limit runs inside one. The --memory flag on docker run (or mem_limit in a Compose file) and MemoryMax= in a systemd unit file do the exact same thing at the kernel level: they set a memory.max ceiling for that specific group of processes.
The part that gets missed most often: a container with a memory limit can be killed by its own cgroup's OOM killer well before the host itself runs short of anything. The kernel tracks each cgroup's limit independently, and if the processes inside one try to cross its own ceiling, it kills a process inside that group, even if free -h on the host shows half its RAM sitting idle at that exact moment. This isn't a system-wide OOM at all, just a local version of it, scoped to a single group of processes.
The log signature differs too. A host-wide kill starts with "Out of memory: Killed process," the exact wording covered above. A cgroup-limit kill starts with "Memory cgroup out of memory: Killed process," and usually sits next to a line naming a path like oom_memcg=/system.slice/docker-<container id>.scope (confirmed live on September 25, 2026 via docker run --memory=50m: current Docker on Ubuntu 24.04 uses systemd as the default cgroup driver, so the path runs through system.slice rather than the separate /docker/<id> hierarchy some older sources show, written for the cgroupfs driver), identifying which container or unit owned the limit. Once you see the second wording, don't waste time hunting for a server-wide leak; go straight to whatever container or service the path points at.
Docker gives you a more direct check than grepping dmesg by hand:
docker inspect <container> --format '{{.State.OOMKilled}}'
true means the container died from its own cgroup's memory limit.false with exit code 137 means it got SIGKILL (128 + signal 9 = 137) for some other reason; exit code 137 alone proves nothing without this check.For a systemd unit with MemoryMax=, the surface symptom looks similar - systemctl status shows something like "Main process exited, code=killed, status=9/KILL" - but that alone doesn't confirm memory either, until you find the matching "Memory cgroup out of memory" line in journalctl -k. There's also a softer threshold worth knowing about: alongside the hard memory.max, cgroup v2 has memory.high, and crossing that lower threshold doesn't kill anything - it makes the kernel start aggressively reclaiming and throttling the group's memory instead, so the container gets visibly sluggish first and only gets killed later, at memory.max, if things don't improve. It's the same "slow before dead" pattern as host-level swap thrashing, just scoped to one container.
Worth knowing about separately: systemd-oomd, a userspace service that watches PSI metrics (pressure stall information, the share of time a cgroup "loses" to memory pressure) and kills processes preemptively, ahead of the kernel's own OOM killer. It ships enabled by default on Ubuntu Desktop, but generally not on Server images, so on a typical VPS you'll almost certainly see the classic kernel OOM killer, not systemd-oomd; if you installed it yourself anyway, its decisions live in journalctl -u systemd-oomd.
If a single VPS runs several containers with no explicit limits, the exact setup in the piece on running multiple services with Docker Compose, that missing limit is precisely why killing one container ends up looking like a mysterious server-wide memory shortage: without a cap, that container competes for the whole host's memory on equal footing with everything else, and the system-wide OOM killer's badness score could just as easily land on some innocent neighbor instead.
Property | Host OOM killer | Cgroup OOM killer (Docker/systemd) |
|---|---|---|
What ran out | RAM and swap for the whole host | the memory.max ceiling of one specific process group |
Host state at the time | memory genuinely exhausted | host can have gigabytes free |
Line in dmesg / journalctl -k | "Out of memory: Killed process..." | "Memory cgroup out of memory: Killed process..." + an oom_memcg=/system.slice/docker-<id>.scope path |
Who becomes the victim | the highest-badness process across the entire server | a process inside the same group that hit its limit |
How to confirm | free -h, vmstat, /proc/<pid>/oom_score | docker inspect --format '{{.State.OOMKilled}}', systemctl status of the unit |
The order that separates memory from CPU and disk in about two minutes: uptime or w for load itself, vmstat 1 for the r/b columns and swap activity, free -h for how much headroom is left, dmesg -T | grep -i "killed process" for OOM history. Here's how to read the combinations.
What the terminal shows | Likely cause |
|---|---|
load high, CPU nearly idle, r low, b high, si/so = 0 | waiting on disk or network, not memory - see the disk-diagnostics hub |
load high, CPU nearly idle, b high, si/so nonzero, available in free -h near zero | swap thrashing: memory is short right now |
load close to nproc, us/sy high in top | a genuine CPU shortage, not memory |
a process vanished with no error, dmesg has "Out of memory: Killed process" | the system-wide OOM killer fired |
a container keeps restarting, docker inspect shows OOMKilled: true, free -h on the host shows plenty free | a cgroup memory limit (--memory / MemoryMax=) killed it, not a host-wide shortage |
I won't pretend these three are interchangeable options on a menu - each has an honest, narrow job, and mixing them up costs you.
Swap is a stopgap, not a cure. It postpones the moment the OOM killer arrives, trading a hard crash for slow degradation, which beats losing a process outright but doesn't fix the underlying shortage: if a process's working set is reliably bigger than RAM, paging runs constantly, and that constant paging often hurts more than the swap helps. Keep swap as insurance against rare short spikes - a one-off build, a package upgrade - not as a permanent stand-in for RAM you don't have. If there's no swap yet:
sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
fallocate -l 2G reserves a fixed-size file for swap - on a plan with a 10-20 GB disk, don't go much bigger than a couple of gigabytes; it's still disk space.chmod 600 keeps other users from reading the file, since it can end up holding fragments of other processes' memory.mkswap formats the file for swap use, swapon turns it on immediately, before any reboot.To keep swap active after a reboot, add a line to /etc/fstab:
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
What you should see. swapon --show afterward lists /swapfile with the size you set - empty output means it didn't turn on.
Container memory limits aren't red tape, they're damage control. Leaving containers unbounded is less code to write, but worse in an actual incident: without a cap, one leaking service competes for the entire host's memory on equal footing with everything else, and when the host-wide OOM killer's badness score fires, it can just as easily pick an unrelated neighbor over the container that's actually leaking. An explicit --memory on Docker or MemoryMax= in a systemd unit turns "the whole server went down for no obvious reason" into "exactly one service died, exactly the one that should have," and pairing it with a restart policy (restart: unless-stopped in Compose, Restart=on-failure in systemd) fixes that without you.
Upgrade RAM when the shortage is chronic, not occasional. If swap thrashing shows up under ordinary daily load rather than once a week during a backup job, neither swap nor container limits solve that - they just make a bad situation more predictable. Swap running constantly at that point isn't insurance anymore, it's a chronic condition that hits response times directly. Moving to a plan with more memory, the Scale tab on a hiplet's page in HIP's panel and the Memory tier built specifically for this, doesn't delay the problem, it removes the cause. Expect the panel to ask for the server to power off while the new plan applies; that's routine when a plan changes, not a glitch.
Load average counts more than processes ready to run on the CPU; it also counts processes stuck in state D (uninterruptible sleep), typically blocked on disk, a network filesystem, or swap. With a load of 15 and a mostly idle CPU, check the si/so columns in vmstat 1 - it's often active swapping from a memory shortage, not a disk problem.
Run sudo dmesg -T | grep -i "killed process" or journalctl -k -b | grep -i "out of memory". A line like "Out of memory: Killed process 1068570 (stress-ng-vm)" with a timestamp and process name is direct proof. For containers, also check docker inspect --format '{{.State.OOMKilled}}'.
At the moment memory runs out, the kernel computes a badness score for every process - roughly its share of used memory, adjusted by oom_score_adj from -1000 to +1000. It kills whichever process has the highest badness right then; the score is recalculated fresh at every OOM event, not read off a fixed ranking set earlier.
A system-wide kill happens once the host has exhausted RAM and swap entirely. A cgroup-OOM kill happens earlier and more locally, as soon as a process inside a container or systemd unit with a memory limit hits its own memory.max, even with gigabytes free on the host. The log line differs: "Out of memory" versus "Memory cgroup out of memory."
It delays it, it doesn't fix the shortage: once a process's working set stays bigger than RAM, the kernel pages constantly, every memory access routes through disk, and the system looks frozen even though memory technically still exists. Treat swap as insurance against short, rare spikes, not a standing substitute for RAM you don't have.
The kernel killed the container's main process with SIGKILL because it exceeded its cgroup's memory limit, usually returning exit code 137. Check with docker inspect <container> --format '{{.State.OOMKilled}}'. Exit code 137 alone, without this check, proves nothing - any SIGKILL produces the same code.
si/so in vmstat is almost always swap thrashing, not a shortage of cores.dmesg -T or journalctl -k; for containers, check docker inspect --format '{{.State.OOMKilled}}'.--memory, systemd MemoryMax=) kill inside the limit even with the host wide open - the log line starts with "Memory cgroup out of memory," a separate event from a system-wide kill.MemoryMax= and OOMPolicy= belong.