How to Diagnose OOM Killer Events in Linux

By 

•

Published on

•

9 min read

Rising memory usage bars above a Linux log with a highlighted Killed process entry

A service disappears without warning. There is no crash report, just a process that was running and now is not, sometimes followed by systemd restarting it. Before assuming an application bug, check whether it was killed because of memory pressure.

When the kernel cannot reclaim enough memory to satisfy an allocation, it may invoke the out-of-memory (OOM) killer. The killed process cannot catch the signal or write a final error message, so you need to look outside its own logs.

This guide explains how to confirm an OOM event, read what the kernel recorded, and choose a fix based on the cause.

Quick Reference

TaskCommand
Search the kernel ring buffersudo dmesg -T | grep -iE 'oom|out of memory|killed process'
Search this boot’s kernel journalsudo journalctl -k -g 'oom|out of memory|killed process'
Search the previous bootsudo journalctl -k -b -1 -g 'oom|out of memory|killed process'
Check userspace OOM killssudo journalctl -u systemd-oomd --since '2 hours ago'
Show a running process’s OOM scorecat /proc/PID/oom_score
Show current memory and swap usefree -h

Replace PID with a running process ID. Run the log checks on the affected host; a container may not have access to the host’s kernel log.

What the OOM Killer Does

Linux can allow applications to reserve more virtual memory than the system can back with RAM and swap. This behavior depends on the overcommit policy. Reserving an address range does not mean that the application is using that much physical memory.

An OOM event occurs when the kernel cannot satisfy an allocation after attempting to reclaim memory. This can affect the whole host or be restricted to a service or container that has reached its memory limit. An OOM kill does not necessarily mean that every byte of RAM and swap on the host was in use.

Under the usual selection policy, the kernel considers eligible processes within the affected memory domain and scores them using their memory use and oom_score_adj settings. It does not know which application matters most to your business. The chosen process receives SIGKILL, which it cannot handle or ignore.

Confirm an OOM Event

Start with the kernel log. The dmesg command prints the kernel ring buffer, and -T adds human-readable timestamps. Search for both OOM messages and the line identifying the killed process:

Terminal
sudo dmesg -T | grep -iE 'oom|out of memory|killed process'

An abbreviated example looks like this:

output
[Wed Sep 30 09:14:22 2026] node invoked oom-killer: gfp_mask=0x140cca(GFP_HIGHUSER_MOVABLE|__GFP_COMP), order=0, oom_score_adj=0
[Wed Sep 30 09:14:22 2026] Out of memory: Killed process 2417 (node) total-vm:4185672kB, anon-rss:3923116kB, file-rss:0kB, shmem-rss:0kB
[Wed Sep 30 09:14:22 2026] oom_reaper: reaped process 2417 (node), now anon-rss:0kB, file-rss:0kB, shmem-rss:0kB

The first line names the process whose allocation invoked the OOM killer. That process is not necessarily the victim or the main source of memory pressure. The Killed process line identifies the victim, here PID 2417 running node.

The total-vm value is its virtual address space, not its RAM use. The anon-rss value shows anonymous resident memory, about 3.7 GiB in this example. file-rss and shmem-rss report resident file-backed and shared memory. The final line records the OOM reaper’s progress reclaiming the victim’s memory.

On systemd systems, you can also search the journal. This is useful when the ring buffer no longer contains the event:

Terminal
sudo journalctl -k -g 'oom|out of memory|killed process'

The -k option selects kernel messages from the current boot. The -g option filters messages with a regular expression; an all-lowercase pattern matches without regard to case. The wider pattern also catches messages such as Memory cgroup out of memory.

If the machine rebooted after the incident, search the previous boot:

Terminal
sudo journalctl -k -b -1 -g 'oom|out of memory|killed process'

This works only if logs from that boot were retained. See the journalctl guide for boot selection and persistent logging. The timestamps from dmesg -T can be inaccurate after suspend and resume, so use journal timestamps when comparing events across logs.

For a systemd service, check its journal around the same time. Replace app.service with the affected service:

Terminal
sudo journalctl -u app.service --since '2 hours ago'

Look for the process exit and any restart at the time of the OOM message. A status=9/KILL entry records a SIGKILL exit, and a shell exit status of 137 is also consistent with that signal. Neither establishes who sent it. A manual kill or another supervisor can produce the same result.

Read the Process Table in the Log

The OOM report normally includes a process table when vm.oom_dump_tasks is enabled. A cgroup OOM report is restricted to that cgroup’s eligible tasks. The filtered commands above hide most of this context, so read the full kernel journal around the event:

Terminal
sudo journalctl -k --since '2026-09-30 09:13:00' --until '2026-09-30 09:16:00'

Use the date and time of your incident, and add -b -1 if it happened during the previous boot. An abbreviated table might contain these rows:

output
[  pid  ]   uid  tgid total_vm      rss ... oom_score_adj name
[   2417]  1000  2417  1046418   980779 ...            0 node
[   1183]   999  1183   210334    45211 ...            0 mysqld

The total_vm and rss columns are in pages. On a system with 4 KiB pages, the Node.js process’s rss of 980779 corresponds to 3923116 KiB, matching the kill message. Do not assume the page size is always 4 KiB; check it with:

Terminal
getconf PAGESIZE

The result is the page size in bytes. Use the table to identify large consumers, but do not treat the largest value as proof of a leak. Several processes may have exhausted the available memory together, or a legitimate workload may have exceeded a configured limit.

Check Service and Container Memory Limits

A service can hit a cgroup memory limit while the host still has available RAM. Look for Memory cgroup out of memory or constraint=CONSTRAINT_MEMCG in the kernel report, along with the affected cgroup path.

For a systemd service, inspect its current memory use, configured limits, and cgroup path:

Terminal
systemctl show app.service -p MemoryCurrent -p MemoryHigh -p MemoryMax -p ControlGroup

Numeric memory values are in bytes. MemoryMax=infinity means the unit has no explicit hard limit of its own; a parent slice can still impose one.

On cgroup v2, use the reported ControlGroup path to inspect event counters. For example, if it is /system.slice/app.service, run:

Terminal
sudo cat /sys/fs/cgroup/system.slice/app.service/memory.events

The kernel’s cgroup documentation explains these counters. oom counts occasions when the cgroup reached its limit and an allocation was about to fail; oom_kill counts processes killed by the kernel OOM killer. The counters include descendants and have no timestamps. An oom_kill value alone does not distinguish a host OOM from a cgroup-limit OOM. Correlate it with the logs, and remember that counters disappear when the cgroup is removed.

Check systemd-oomd

Some systems also run systemd-oomd, a userspace service that monitors memory pressure and swap use. It can kill processes in a selected cgroup before the kernel reaches an OOM condition, so there may be no kernel Out of memory message.

Check its journal around the time the application disappeared:

Terminal
sudo journalctl -u systemd-oomd --since '2 hours ago'

A kill entry identifies the cgroup selected by systemd-oomd. Match that path and timestamp to the affected service or user session. If the service is not installed or enabled, this check does not apply. Its selection policy is separate from the kernel’s oom_score_adj mechanism; see the systemd-oomd documentation for details.

Inspect and Adjust OOM Scores

For a running process, /proc exposes its current OOM score and adjustment. To inspect your current shell safely, use $$, which expands to its PID:

Terminal
cat /proc/$$/oom_score
cat /proc/$$/oom_score_adj

For another process, replace $$ with its current PID. The killed process from the log is already gone, and its PID may later be reused. You cannot recover its old score by reading the same /proc/PID path after the event.

The oom_score_adj setting ranges from -1000 to 1000. Positive values make selection more likely, while negative values make it less likely. A value such as -500 does not guarantee that the process will survive. -1000 excludes the process from kernel OOM victim selection, which shifts pressure to other workloads and should not be a routine fix.

If you need to lower a systemd service’s score, create a drop-in with:

Terminal
sudo systemctl edit app.service

Add the adjustment under [Service]:

/etc/systemd/system/app.service.d/override.confini
[Service]
OOMScoreAdjust=-500

Save the file. systemctl edit reloads systemd’s configuration, but the setting applies to newly started processes. Restart the service during a suitable maintenance window, since this interrupts it:

Terminal
sudo systemctl restart app.service

This makes the preference persistent across service restarts. It does not reduce memory use or prevent systemd-oomd from choosing the service’s cgroup.

Prevent the Next OOM Event

Start by checking current memory and swap use:

Terminal
free -h

Read the available column, not the free column. These are current values, so they may look healthy after a large process has been killed. The free command guide explains the columns.

If one application’s memory use keeps growing under a steady workload, investigate a possible leak. If usage rises with traffic or job size, reduce worker concurrency, batch size, or cache limits before assuming a bug.

For systemd services on cgroup v2, MemoryHigh= applies reclaim pressure and throttling, while MemoryMax= provides a hard limit. Exceeding MemoryMax= can cause an OOM kill inside the service. Choose limits from measured usage and leave room for the rest of the host. The systemd resource-control documentation recommends MemoryHigh= as the main control and MemoryMax= as the last line of defense.

If the host does not have enough capacity for its normal workload, add RAM or reduce the workload. Adding swap space can absorb temporary pressure from swappable memory, but it does not fix a leak or remove a service’s memory limit. Sustained swapping can also make the machine slow.

Troubleshooting

dmesg reports “Operation not permitted”
Run it with sudo on the host. Kernel log access is often restricted, and root inside a container may still lack permission. Ask the host administrator for logs if you cannot access them.

The previous boot has no journal entries
List retained boots with sudo journalctl --list-boots. Logs may have been stored only in memory or removed by retention limits. Enabling persistent logging helps future investigations but cannot recover discarded entries.

The process was killed, but the search returns nothing
Check the right boot and time range, then inspect systemd-oomd and supervisor logs. On systems using a syslog daemon, retained kernel messages may also be in /var/log/kern.log, /var/log/messages, or rotated copies. Missing logs do not prove that no OOM occurred, and a SIGKILL exit alone does not prove that one did.

Conclusion

Save the full log around an OOM event before it rotates, including the process table and cgroup information. To watch memory use while reproducing the workload, see our guides on the top command and vmstat .

Linuxize Weekly Newsletter

A quick weekly roundup of new tutorials, news, and tips.

About the authors

Dejan Panovski

Dejan Panovski

Dejan Panovski is the founder of Linuxize, an RHCSA-certified Linux system administrator and DevOps engineer based in Skopje, Macedonia. Author of 1000+ Linux tutorials with 20+ years of experience turning complex Linux tasks into clear, reliable guides.

View author page