Skip to main content
ERPNext how-to & fixes

ERPNext Bench Update Out of Memory: The Real Fix Ladder (2026)

bench update killed on a small VPS is rarely a hardware problem — split the update into three stages, add swap, and cap Node with NODE_OPTIONS.

MManojAugust 9, 202612 min read
More in ERPNext how-to & fixes#upgrades#devops#troubleshooting
Share

An ERPNext bench update out of memory failure — the process is killed mid-run, exit code 137, no error stack — is almost always the Node/webpack asset build during bench update --build, not the code pull or the DB migrate. On a 2 GB VPS the fix is not automatically "buy a bigger box." Split the update into three stages (pull → migrate → build), size swap correctly, cap the Node heap with NODE_OPTIONS, and if you can, build the assets on a different machine and rsync them over. This ladder gets 2 GB and 4 GB boxes through a v16 upgrade cleanly.

Every "server needs more RAM" invoice we've read from a hosting provider after an ERPNext upgrade started with the same signature: bench update dies during asset build, the reader assumes the whole server is undersized, and buys a machine twice as expensive as they need. The failure is real; the diagnosis is usually wrong. I am Manoj, and I lead ERPNext and Frappe implementation and infrastructure programmes at MithTech, a 35-person Bangalore practice that designs, customises, integrates and operates business-critical software. This is the ladder we walk before recommending a hardware upgrade.

How do I confirm the failure is actually OOM?

Answer

An ERPNext bench update out of memory failure has three tells: the process exits with code 137, dmesg | grep -i oom prints an oom-killer line naming a Node or Python process, and journalctl -u <bench-supervisor-unit> (or the terminal output) stops mid-sentence with no Python traceback. If your bench update ended in a clean Python exception, it was not OOM — read the exception.

Skimmable summary: three signals; you need at least two to be confident it was OOM.

Run the following in order:

# 1. What was the exit code?
echo $?          # ← immediately after the failed bench update

# 2. Did the kernel kill something big?
sudo dmesg -T | grep -iE "oom-killer|killed process" | tail -20

# 3. Which processes were memory-heaviest at the time?
sudo journalctl --since "10 minutes ago" | grep -iE "oom|killed" | tail -20

Signal cross-check:

Exit codedmesg lineWhat happened
137oom-killer naming node / yarn / webpackNode asset build OOM'd
137oom-killer naming python / gunicorn / uwsgiPython side OOM'd (rare — usually pip install of a heavy dep)
137No oom-killer lineContainer/systemd sent SIGKILL for a different reason (unlikely on a bare VPS)
1 or non-137—Not OOM. Read the actual Python traceback and treat it as a normal failure.

If exit was not 137, the fix ladder below does not apply — you have a code-level failure and need to read the traceback.

Add swap before you touch anything else

Skimmable summary: 2–4 GB of swap gets most 2 GB VPS boxes through v16. Do this in three commands.

If your VPS has < 4 GB of RAM and no swap, add 4 GB of swap now. It is the cheapest safety net for the asset build and takes two minutes:

# Create a 4 GB swap file
sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

# Make it persistent across reboots
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

# Verify
free -h

free -h should now show ~4 GB of swap available. This is not a permanent fix — a 2 GB VPS with 4 GB of swap can complete a v16 asset build, but the build is slower because the last few hundred MB are on disk. If the build takes more than 30 minutes and finishes, swap saved you. If it does not finish in an hour, you need one of the ladder steps below.

Do not add more than 8 GB of swap. Beyond that ratio, thrashing costs you more time than a small VPS upgrade would.

Which step of bench update is actually killing you?

Skimmable summary: assets → migrate → pull, in decreasing order of memory pressure.

bench update runs (in order) pull → requirements → patch → build → restart-supervisor — the flag surface is documented in bench/commands/update.py. Each stage has a distinct peak RSS profile:

Typical peak RSSNotes
Pullbench update --pull --no-backup~200 MBGit operations only; almost never OOM
Requirementsbench update --requirements~500 MBPip resolves + installs; can spike on a heavy dep with a C compile
Migratebench update --patch~400 MBRuns schema migrations; scales with number of sites and patch complexity
Buildbench build1.5–3 GBNode webpack/Vite building JS + CSS for every app; the OOM killer's favourite
Restartbench restart~100 MBSends signal to supervisor; negligible

The 1.5–3 GB peak of bench build is the reason a "2 GB VPS was fine yesterday and OOMs on update today" — yesterday nothing was building assets. Today bench update runs the whole pipeline as one process tree, so the peak is patch + build overlapping.

How do you split bench update into memory-safe steps?

Skimmable summary: run the three sub-commands separately, each in its own process tree — peak memory drops from the sum to the max.

Instead of bench update in one shot, run this on any box under 8 GB of RAM:

Take a backup manually

bench --site all backup --with-files. This is what bench update runs first when you do not pass --no-backup. Doing it as a separate step gets the backup safely off the critical path and lets you skip it in the update proper.

Pull code only

bench update --pull --no-backup. This runs git pull across all apps and sets up Python requirements. Peak RSS ~500 MB. Should complete on a 2 GB box without swap.

Run migrations

bench update --patch. This runs the DB schema migrations and site-level patches. Peak RSS ~400 MB. Same profile — a 2 GB box handles this cleanly without swap.

Build assets last, with a Node heap cap

NODE_OPTIONS="--max-old-space-size=1024" bench build. This is the OOM-prone step. Capping the Node heap to 1024 MB constrains any single build worker; combined with swap, a 2 GB box will complete. For a 4 GB box, use --max-old-space-size=2048.

Restart supervisor

bench restart. Reloads gunicorn, workers and scheduler under the new code. Trivial memory footprint.

Each step exits cleanly before the next begins, so peak memory across the whole update drops from sum(stages) to max(stages). On a 2 GB box, that is the difference between "the kernel killed everything" and "the update took 15 minutes."

When should you set NODE_OPTIONS and to what?

Skimmable summary: cap Node's heap to a fraction of RAM so one build worker cannot consume the whole box.

By default, Node 20+ picks a max-old-space-size of roughly 25–50% of system RAM, which on a 4 GB VPS is ~1–2 GB — enough for a single build worker to starve everything else including MariaDB. NODE_OPTIONS="--max-old-space-size=<N>" (in MB) sets a hard ceiling; the flag is documented on the official Node.js CLI reference.

Sizing rule of thumb:

VPS RAMRecommended --max-old-space-sizeWhy
2 GB1024Leaves ~1 GB for the kernel, MariaDB, and one background worker
4 GB2048Leaves 2 GB for the rest of the bench + MariaDB tuning headroom
8 GB4096Rarely needed to cap — but does prevent runaway workers under heavy load
16 GB+Do not capNode's default is fine; capping wastes headroom

Set it inline for one command:

NODE_OPTIONS="--max-old-space-size=1024" bench build

Or persist it in the bench user's shell:

echo 'export NODE_OPTIONS="--max-old-space-size=1024"' >> ~/.bashrc

The persistent form is safer because supervisor-managed background jobs that shell out to Node will also see the cap.

Can you build assets on a different machine?

Skimmable summary: yes — build on a bigger box, rsync sites/assets/ over. Standard practice for constrained production boxes.

If your production VPS is genuinely small (2 GB, no room to add swap, and MariaDB is already using most of it), doing the asset build on a separate host and copying the output over is a clean pattern:

  1. On a build machine (a 4–8 GB VM you spin up for the upgrade — a Hetzner CX22 works fine and costs pennies for the hour):

    • Clone the exact same bench directory (git clone of the bench repo + bench get-app for every app on production).
    • Check out the same versions.
    • bench build there. Peak RSS 1.5–3 GB completes comfortably.
  2. Rsync the built assets to the production box:

rsync -avz --delete \
   --exclude 'assets/frappe/dist' \
   /path/to/build-bench/sites/assets/ \
   user@prod:/path/to/prod-bench/sites/assets/
  1. On production, run bench update --pull --no-backup && bench update --patch && bench restart (skip --build). The assets are already there.

This is standard practice at hosting providers running many customer sites on constrained boxes. It also parallelises: you can build tomorrow's release on the build box while today's is still serving traffic.

The rest of the OOM budget: MariaDB and workers

bench build OOMs are the loudest, but two other processes commonly steal memory during an upgrade:

  • MariaDB's InnoDB buffer pool sized too aggressively for the box. If you tuned innodb_buffer_pool_size to 70% of RAM per the ERPNext MariaDB tuning guide, during an upgrade the build needs some of that back. Drop the buffer pool temporarily (SET GLOBAL innodb_buffer_pool_size=<smaller>) before the build, restore after.
  • The RQ workers. They keep running during bench update unless you stop them. On a 2 GB box, sudo supervisorctl stop <bench-name>: (whole group) before the upgrade and start after can free 200–400 MB.

Neither is required if you have adequate swap, but on a truly constrained box both are cheap wins.

FAQ

The four questions that come up most on this ticket class.

Why does bench update OOM on a 2 GB VPS when the machine was fine yesterday?

Because yesterday nothing was compiling JavaScript. The daily steady-state of ERPNext (gunicorn + workers + MariaDB + Redis) fits in 2 GB. Adding bench build's 1.5–3 GB Node peak on top pushes the machine into OOM. The "the server is too small" conclusion is wrong — the steady state is fine; only the upgrade blows the ceiling. Split the update and add swap and the same box will complete cleanly.

Is --no-backup safe in production?

It's safe if you take the backup yourself in a separate step before starting. --no-backup skips the bench update-managed backup; it does not delete data. Manually running bench --site all backup --with-files before, then running bench update --no-backup, is what most production upgrades actually do — you get a backup at a known time, and the update runs without the backup step pinning extra memory.

Does more swap always help?

Up to about 4× physical RAM, yes. Beyond that, the swap thrashing costs you more time than the equivalent VPS upgrade would. A 2 GB VPS with 8 GB of swap will complete bench build — but in an hour. If your upgrade window is smaller than that, pay for the bigger box for the day and downsize after.

Does exit code 137 always mean OOM on Linux?

Almost always. Exit 128 + N on Linux means the process was killed by signal N; 137 = 128 + 9 where 9 is SIGKILL. The kernel's OOM killer is by far the most common source of an unsolicited SIGKILL. The other sources — a manual kill -9, a Docker container hitting --memory limits, a systemd MemoryMax — leave different fingerprints in dmesg and journalctl. dmesg | grep -i oom gives you the definitive answer in one line.

bench update killed by OOM on your VPS?

We deploy and operate self-hosted ERPNext on Hetzner, AWS, DigitalOcean and bare-metal for Indian SMBs. If your upgrade window is closing and the kernel keeps killing bench build, we can help you get the upgrade through on the box you have — without buying hardware you may not need.

Next step · ERPNext how-to & fixes

Now see support that keeps it fixed

Support and maintenance for ERPNext you already run: fixes, upgrades and the recurring issues your team should not have to chase.

M

Written by

Manoj

Founder of MithTech, an open-source ERP & automation engineering practice. Hands-on ERPNext/Frappe implementation across multi-branch, multi-warehouse Indian operations — GST/TDS/PT compliance, branch-level permissions, and custom Frappe apps that give management real-time visibility.

Free · By email

Get practical ERPNext & automation guides

New implementation guides, cost breakdowns and open-source tips for Indian businesses — occasionally, straight to your inbox. No spam.

Already a MithTech client?

Help the next operator choose.

Most teams evaluating ERPNext have no way to tell who actually delivers. If we’ve run an implementation for you, two lines on Google count for more than anything we can write about ourselves.

Leave a Google review

Only if we’ve actually worked together — Google filters reviews from non-customers, so an honest one is worth more than ten polite ones.

Keep reading

See what this looks like for your business

A 30-minute working session with a principal consultant. We pressure-test the architecture and outline the engagement model that fits your governance and procurement posture. You leave with a written brief.

0
Published on 9 August 2026

Manoj

Comments & ratings

No comments yet. Start a new discussion.