Domain insights

Your VPS Has Free Disk Space, So Why Did It Crash? RAM, Swap, and the Linux OOM Killer

Learn why a Linux VPS can crash with plenty of disk space, how RAM and swap work, what the OOM killer does, and how to diagnose memory pressure.

VPS and Linux
Your VPS Has Free Disk Space, So Why Did It Crash? RAM, Swap, and the Linux OOM Killer

The alert arrives at 2:17 in the morning.

Your website is down.

You log in to the VPS and check the obvious thing:

df -h

The disk is only 38 percent full.

So why did PHP stop? Why did the database disappear? Why does the system log say a process was "killed"?

Because disk space and working memory are different resources.

A server can have hundreds of gigabytes of free storage and still run out of RAM. When memory pressure becomes severe, Linux has to make hard decisions. It can reclaim memory, move some inactive pages to swap, refuse allocations, or, when the situation becomes critical, invoke the Out Of Memory killer.

The OOM killer sounds dramatic because it is.

Its job is to sacrifice a process so the operating system has a chance to survive.

If you run a VPS, a dedicated server, a web stack, a database, containers, or a busy control panel, understanding this behavior is one of the most useful troubleshooting skills you can learn. It turns a mysterious "server crash" into a measurable resource problem.

This guide keeps the explanation practical.

Compare the memory allocation in our VPS plans against the combined needs of your application, database and background jobs. Our VPS guide and hosting comparison help you choose a suitable starting platform.

Chapter 1: Disk space is the warehouse, RAM is the workbench

Imagine a workshop.

The warehouse can hold ten thousand boxes. That is your disk.

The workbench can hold only ten boxes at a time. That is your RAM.

You may have enormous storage capacity, but the work stops if the bench is completely covered and every new task needs more working space.

Applications use RAM for active data.

A web server keeps processes, connections, buffers, and caches in memory.

PHP workers need memory while executing requests.

A database keeps indexes, query buffers, connections, temporary structures, and cached pages in memory.

The operating system also uses RAM.

So does your monitoring agent, control panel, antivirus tool, backup software, container runtime, DNS service, mail server, and everything else running on the machine.

Disk free space does not solve a shortage on the workbench.

Chapter 2: What Linux does with RAM

Linux tries to use memory efficiently.

This creates another common misunderstanding.

An administrator runs:

free -h

and sees a small number under free.

Panic follows.

But Linux intentionally uses otherwise idle RAM for useful caches. Memory used for file cache can often be reclaimed when applications need it.

The more useful number to watch is usually available memory rather than expecting most RAM to sit completely unused.

A simplified output may look like:

total        used        free      shared  buff/cache   available
Mem:            8Gi         5Gi       400Mi       300Mi        2.6Gi        2.2Gi
Swap:           2Gi       200Mi        1.8Gi

Only 400 MiB is labeled free, but about 2.2 GiB may still be available for new workloads.

That server is not automatically in trouble.

Memory troubleshooting is about pressure, growth, and behavior over time, not one scary number taken out of context.

Chapter 3: What swap actually does

Swap is disk space Linux can use as an extension for memory pages that do not need to remain in physical RAM every moment.

Swap can be a partition or a file.

Check it with:

swapon --show

and:

free -h

Red Hat describes swap as temporary storage for inactive processes and data when physical memory is full. It can help prevent out-of-memory conditions, but it is slower than RAM and should not be treated as a replacement for properly sized memory.

That point matters.

A little swap can give a server breathing room during a short spike.

A server constantly pushing large amounts of active data into swap may become painfully slow.

The machine has not solved its memory problem. It has changed the symptom from "out of memory" to "everything feels frozen."

Disk capacity, available RAM and swap are different resources during a memory incident.
Disk capacity, available RAM and swap are different resources during a memory incident.

Chapter 4: The moment Linux runs out of options

Applications ask the kernel for memory.

Most of the time, the kernel satisfies those requests.

When memory gets tight, Linux can reclaim caches and attempt other memory management strategies. If swap exists, inactive memory pages may be moved there.

But resources are finite.

Eventually, the system may reach a point where it cannot satisfy an important memory allocation.

This is where the Out Of Memory mechanism becomes relevant.

Linux can invoke the OOM killer to select a process and terminate it, freeing memory so the rest of the system may continue.

The kernel documentation describes an OOM selection process that uses heuristics to decide which task is a suitable victim. Linux also exposes controls such as oom_score_adj that can influence how likely a process is to be selected.

A container or systemd service can also reach its own cgroup memory limit while the host still has available RAM. On systems using systemd-oomd, a userspace memory-pressure policy can stop a cgroup before the kernel reports a global OOM; inspect service/container limits and the systemd-oomd journal as well as kernel messages when the evidence points in that direction.

For most hosting administrators, the key idea is not the scoring algorithm.

The key idea is this:

When Linux kills a service during an OOM event, the killed process may be the victim of the shortage, not the original cause of the shortage.

That distinction changes how you investigate.

Chapter 5: The innocent victim problem

Suppose a PHP application has a memory leak.

Traffic rises.

Hundreds of PHP workers grow over time.

MySQL is using a reasonable amount of memory.

Eventually, the server runs out.

The OOM killer chooses mysqld as a victim because killing it would release a large amount of memory quickly.

Your monitoring then reports:

Database service stopped.

It is tempting to conclude:

MySQL caused the incident.

Maybe it did.

Maybe it did not.

The database might simply have been the largest useful target available when the system was already in trouble.

This is why an OOM incident requires timeline analysis.

You need to know which workloads were growing before the kill, not only which process died at the end.

A fresh memory snapshot cannot reconstruct an earlier incident on its own. Use the timestamps and service history described in our VPS logs guide and keep monitoring over time.

Chapter 6: How to confirm an OOM event

On a modern system using systemd, start with the kernel journal.

Try:

sudo journalctl -k | grep -Ei 'out of memory|oom|killed process'

You can also use:

sudo dmesg -T | grep -Ei 'out of memory|oom|killed process'

Depending on permissions, log rotation, and system configuration, one source may be more useful than the other.

An OOM event often leaves messages indicating that the kernel invoked the OOM killer and terminated a process.

Do not stop after finding the line.

Record:

  • The timestamp
  • The killed process
  • The process ID
  • The memory-related values shown in the event
  • The services active at that time
  • Traffic or job activity around the same minute
  • Cron jobs, backups, scans, imports, or batch processes running nearby
  • Container or control panel events

The final kill is the last page of the story.

Your job is to find the first page.

Chapter 7: The first five commands I would run

If the server is stable enough to inspect, these commands provide a useful first picture.

1. Current memory

free -h

Look at total, available memory, and swap use.

2. Swap configuration

swapon --show

No output usually means swap is not active.

3. Top memory consumers

ps aux --sort=-%mem | head -20

This gives a quick list of processes using the largest percentage of RAM.

4. Live system activity

vmstat 1 10

vmstat can help show memory, run queue, I/O, and swap activity over several seconds.

Red Hat specifically notes the si and so fields for swap-in and swap-out activity when evaluating swapping behavior.

5. Kernel OOM evidence

sudo journalctl -k | grep -Ei 'out of memory|oom|killed process'

These commands do not replace monitoring, but they help answer the immediate question:

Is this really a memory incident?

Chapter 8: Common causes on hosting servers

Too many PHP workers

Each PHP-FPM process consumes memory.

A configuration that allows far more workers than the VPS can support may behave well under light traffic and collapse during a burst.

If one worker uses 120 MB and you allow 80 workers, the theoretical process pool can become very expensive on a small VPS.

Do not size worker counts from wishful thinking. Measure actual process memory under realistic load.

MySQL or MariaDB sized for a larger machine

Database tuning examples from the internet are dangerous when copied blindly.

A buffer pool or cache size that is sensible on a 64 GB server can suffocate a 4 GB VPS.

Backup and compression jobs

A nightly backup can increase CPU, disk I/O, process count, and memory use at the same time.

If your incidents happen on a schedule, inspect cron jobs and backup windows.

Malware scanners and security agents

Security tools are important, but some scans are resource intensive.

Schedule and size them with the rest of the server workload in mind.

Traffic spikes

A promotion, bot wave, crawler burst, attack, or viral link can multiply concurrent application workers.

The problem may be concurrency rather than one request consuming extreme memory.

Memory leaks

A process slowly grows and never returns memory as expected.

A restart clears the symptom, which is why leaks can hide for weeks.

If memory climbs in a staircase pattern between restarts, investigate the application or daemon instead of adding endless RAM.

Containers without sensible limits

A container platform can make one workload look isolated while it still competes for finite host memory.

Limits, reservations, and monitoring matter.

Chapter 9: Is swap good or bad?

The useful answer is: it depends on the workload and the objective.

Swap is not evil.

Swap is not free RAM either.

A modest swap area can prevent a short spike from becoming an immediate OOM kill. It can give inactive pages somewhere to go and provide the administrator a little time to react.

But a server that constantly depends on heavy swapping is under memory pressure.

For latency-sensitive workloads, swapping can be especially noticeable because storage is much slower than RAM.

The right question is not:

Should Linux ever use swap?

The better questions are:

  • Is swap activity occasional or constant?
  • Are users seeing latency when swapping rises?
  • Is the working set larger than physical RAM?
  • Is one service misconfigured?
  • Is the VPS simply undersized?
  • Would adding RAM solve the root cause, or only delay a leak?

Red Hat's guidance is direct: swap can help when RAM is full, but it should not be considered a replacement for more physical memory.

Chapter 10: Why clearing caches is usually not the real fix

During memory incidents, administrators sometimes search for commands that "free RAM."

That can lead to aggressive cache clearing or repeated service restarts.

A restart may temporarily restore service.

It may also erase useful evidence and hide the growth pattern you needed to diagnose.

Linux already knows how to reclaim cache when applications genuinely need memory.

If the system is repeatedly reaching OOM, the real investigation belongs in:

  • Workload sizing
  • Process counts
  • Application behavior
  • Memory leaks
  • Database configuration
  • Container limits
  • Traffic patterns
  • Scheduled jobs
  • Physical RAM capacity

Treat manual cache clearing as a troubleshooting action that needs a specific reason, not a maintenance ritual.

If a tuned workload still exceeds the resources assigned to it, compare a larger VPS with our dedicated servers. For a website that does not need operating-system control, shared hosting may reduce the administration involved. The choice should follow measured workload requirements.

Chapter 11: Capacity planning in simple numbers

You do not need an advanced performance model to avoid obvious mistakes.

Start with a memory budget.

Suppose a VPS has 8 GB of RAM.

Reserve capacity for the operating system and background services.

Then estimate the major consumers:

Operating system and base services     1.2 GB
Database                                2.5 GB
Web server                              0.4 GB
PHP workers                             2.4 GB
Security and monitoring                 0.5 GB
Safety margin                           1.0 GB

Total:

8.0 GB

That is already tight because real workloads move.

If PHP worker memory doubles under certain requests, or the database needs a temporary spike, the budget breaks.

A safety margin is not wasted RAM. It is what keeps a normal burst from turning into an incident.

If you are still choosing the platform size, our Shared Hosting, VPS, Dedicated and Cloud Compared guide helps frame the infrastructure decision before tuning begins.

Chapter 12: Watch memory as a trend, not an emergency

The best time to investigate memory pressure is before the OOM killer appears.

Track:

  • Used and available RAM
  • Swap use
  • Swap-in and swap-out rate
  • Per-process memory
  • PHP-FPM worker count
  • Database memory
  • Container memory
  • Request concurrency
  • Load average
  • I/O wait
  • Restart events
  • OOM events

Set alerts with enough lead time to act.

An alert at 95 percent may arrive too late if the workload can consume the remaining memory in seconds.

A useful monitoring design often has multiple thresholds and focuses on sustained pressure rather than one brief sample.

For example, a short burst to high memory use may be harmless if available memory recovers immediately. A slow daily climb from 45 percent to 90 percent is more suspicious even though it looks calm.

Chapter 13: What to do after an OOM incident

A good incident response is more than "restart the service."

1. Restore service safely

If a critical process was killed, bring it back in a controlled way.

2. Preserve evidence

Save the relevant kernel logs, service logs, graphs, and timestamps before log rotation removes them.

3. Identify the memory growth source

Compare the hours before the incident, not only the final minute.

4. Review configuration limits

Check application workers, database buffers, Java heaps, container limits, and similar settings.

5. Review scheduled activity

Backups, malware scans, imports, log processing, and cron jobs may have overlapped.

6. Decide whether the problem is configuration or capacity

If normal legitimate traffic requires more memory than the server has, tuning cannot manufacture physical RAM.

Scale the instance or redesign the workload.

7. Test the fix

Do not wait for the next real outage to learn whether your change worked.

Use controlled load testing where appropriate.

Chapter 14: When adding RAM is the right answer

Engineers sometimes resist scaling because they want to "optimize first."

Optimization is valuable.

But not every capacity problem is a software bug.

If a healthy database, expected traffic, required security tools, and properly sized application workers genuinely need 12 GB of working memory, running them on a 4 GB VPS is not efficiency. It is denial with a control panel.

Add RAM when the measured workload justifies it.

Move to a larger VPS, dedicated server, or cloud design when the operational requirements have outgrown the current machine.

The goal is not to keep the smallest server alive at any cost.

The goal is to run the service reliably.

Chapter 15: A practical memory checklist

Before declaring a Linux VPS healthy, verify:

  • free -h shows a reasonable amount of available memory during normal load.
  • Swap is configured according to your workload and provider design.
  • Swap activity is not constantly heavy.
  • Application worker limits fit within the RAM budget.
  • Database memory settings fit the actual machine size.
  • Backups and scans are scheduled with resource overlap in mind.
  • You know how to find OOM messages in the kernel journal.
  • Monitoring records per-process or per-service trends.
  • Alerts fire before the server reaches a critical state.
  • A documented response exists for repeated memory pressure.
  • Capacity is reviewed as traffic grows.

A server should not need luck to stay online.

Frequently asked questions

Can a VPS crash even if the disk has lots of free space?

Yes. Disk storage and RAM are different resources. A system can have terabytes of free disk space and still run out of working memory.

What is swap?

Swap is disk-backed space Linux can use for memory pages when physical RAM is under pressure. It can provide breathing room, but it is slower than RAM.

What is the Linux OOM killer?

It is a kernel mechanism that can terminate a process during severe out-of-memory conditions to free memory and give the system a chance to continue operating.

If MySQL was killed, does that mean MySQL caused the OOM?

Not necessarily. The killed process may be the largest useful victim rather than the original source of the memory growth. Review the timeline before assigning the cause.

How do I check whether OOM happened?

Look in kernel logs with commands such as: sudo journalctl -k | grep -Ei 'out of memory|oom|killed process'

Should I disable the OOM killer?

Do not change OOM behavior casually. Kernel OOM settings affect how the operating system responds under extreme memory pressure and should be changed only for a well understood architecture and recovery strategy.

Should I add swap to every VPS?

There is no universal size or policy for every workload. Swap can be useful, but the correct design depends on memory size, latency needs, provider configuration, workload behavior, and operational objectives.

Will adding more RAM fix a memory leak?

It may delay the failure, but it does not remove the leak. If a process keeps growing without bound, identify and correct the underlying software or configuration issue.