An ERPNext bench update out of memory failure — the process is killed mid-run, exit code 137, no error stack — is almost always the Node/webpack asset build during bench update --build, not the code pull or the DB migrate. On a 2 GB VPS the fix is not automatically "buy a bigger box." Split the update into three stages (pull → migrate → build), size swap correctly, cap the Node heap with NODE_OPTIONS, and if you can, build the assets on a different machine and rsync them over. This ladder gets 2 GB and 4 GB boxes through a v16 upgrade cleanly.
Every "server needs more RAM" invoice we've read from a hosting provider after an ERPNext upgrade started with the same signature: bench update dies during asset build, the reader assumes the whole server is undersized, and buys a machine twice as expensive as they need. The failure is real; the diagnosis is usually wrong. I am Manoj, and I lead ERPNext and Frappe implementation and infrastructure programmes at MithTech, a 35-person Bangalore practice that designs, customises, integrates and operates business-critical software. This is the ladder we walk before recommending a hardware upgrade.
How do I confirm the failure is actually OOM?
Answer
An ERPNext bench update out of memory failure has three tells: the process exits with code 137, dmesg | grep -i oom prints an oom-killer line naming a Node or Python process, and journalctl -u <bench-supervisor-unit> (or the terminal output) stops mid-sentence with no Python traceback. If your bench update ended in a clean Python exception, it was not OOM — read the exception.
Skimmable summary: three signals; you need at least two to be confident it was OOM.
Run the following in order:
# 1. What was the exit code?
echo $? # ← immediately after the failed bench update
# 2. Did the kernel kill something big?
sudo dmesg -T | grep -iE "oom-killer|killed process" | tail -20
# 3. Which processes were memory-heaviest at the time?
sudo journalctl --since "10 minutes ago" | grep -iE "oom|killed" | tail -20
Signal cross-check:
| Exit code | dmesg line | What happened |
|---|---|---|
| 137 | oom-killer naming node / yarn / webpack | Node asset build OOM'd |
| 137 | oom-killer naming python / gunicorn / uwsgi | Python side OOM'd (rare — usually pip install of a heavy dep) |
| 137 | No oom-killer line | Container/systemd sent SIGKILL for a different reason (unlikely on a bare VPS) |
| 1 or non-137 | — | Not OOM. Read the actual Python traceback and treat it as a normal failure. |
If exit was not 137, the fix ladder below does not apply — you have a code-level failure and need to read the traceback.
Add swap before you touch anything else
Skimmable summary: 2–4 GB of swap gets most 2 GB VPS boxes through v16. Do this in three commands.
If your VPS has < 4 GB of RAM and no swap, add 4 GB of swap now. It is the cheapest safety net for the asset build and takes two minutes:
# Create a 4 GB swap file
sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
# Make it persistent across reboots
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
# Verify
free -h
free -h should now show ~4 GB of swap available. This is not a permanent fix — a 2 GB VPS with 4 GB of swap can complete a v16 asset build, but the build is slower because the last few hundred MB are on disk. If the build takes more than 30 minutes and finishes, swap saved you. If it does not finish in an hour, you need one of the ladder steps below.
Do not add more than 8 GB of swap. Beyond that ratio, thrashing costs you more time than a small VPS upgrade would.
Which step of bench update is actually killing you?
Skimmable summary: assets → migrate → pull, in decreasing order of memory pressure.
bench update runs (in order) pull → requirements → patch → build → restart-supervisor — the flag surface is documented in bench/commands/update.py. Each stage has a distinct peak RSS profile:
| Typical peak RSS | Notes | ||
|---|---|---|---|
| Pull | bench update --pull --no-backup | ~200 MB | Git operations only; almost never OOM |
| Requirements | bench update --requirements | ~500 MB | Pip resolves + installs; can spike on a heavy dep with a C compile |
| Migrate | bench update --patch | ~400 MB | Runs schema migrations; scales with number of sites and patch complexity |
| Build | bench build | 1.5–3 GB | Node webpack/Vite building JS + CSS for every app; the OOM killer's favourite |
| Restart | bench restart | ~100 MB | Sends signal to supervisor; negligible |
The 1.5–3 GB peak of bench build is the reason a "2 GB VPS was fine yesterday and OOMs on update today" — yesterday nothing was building assets. Today bench update runs the whole pipeline as one process tree, so the peak is patch + build overlapping.
How do you split bench update into memory-safe steps?
Skimmable summary: run the three sub-commands separately, each in its own process tree — peak memory drops from the sum to the max.
Instead of bench update in one shot, run this on any box under 8 GB of RAM:
Take a backup manually
bench --site all backup --with-files. This is what bench update runs first when you do not pass --no-backup. Doing it as a separate step gets the backup safely off the critical path and lets you skip it in the update proper.
Pull code only
bench update --pull --no-backup. This runs git pull across all apps and sets up Python requirements. Peak RSS ~500 MB. Should complete on a 2 GB box without swap.
Run migrations
bench update --patch. This runs the DB schema migrations and site-level patches. Peak RSS ~400 MB. Same profile — a 2 GB box handles this cleanly without swap.
Build assets last, with a Node heap cap
NODE_OPTIONS="--max-old-space-size=1024" bench build. This is the OOM-prone step. Capping the Node heap to 1024 MB constrains any single build worker; combined with swap, a 2 GB box will complete. For a 4 GB box, use --max-old-space-size=2048.
Restart supervisor
bench restart. Reloads gunicorn, workers and scheduler under the new code. Trivial memory footprint.
Each step exits cleanly before the next begins, so peak memory across the whole update drops from sum(stages) to max(stages). On a 2 GB box, that is the difference between "the kernel killed everything" and "the update took 15 minutes."
When should you set NODE_OPTIONS and to what?
Skimmable summary: cap Node's heap to a fraction of RAM so one build worker cannot consume the whole box.
By default, Node 20+ picks a max-old-space-size of roughly 25–50% of system RAM, which on a 4 GB VPS is ~1–2 GB — enough for a single build worker to starve everything else including MariaDB. NODE_OPTIONS="--max-old-space-size=<N>" (in MB) sets a hard ceiling; the flag is documented on the official Node.js CLI reference.
Sizing rule of thumb:
| VPS RAM | Recommended --max-old-space-size | Why |
|---|---|---|
| 2 GB | 1024 | Leaves ~1 GB for the kernel, MariaDB, and one background worker |
| 4 GB | 2048 | Leaves 2 GB for the rest of the bench + MariaDB tuning headroom |
| 8 GB | 4096 | Rarely needed to cap — but does prevent runaway workers under heavy load |
| 16 GB+ | Do not cap | Node's default is fine; capping wastes headroom |
Set it inline for one command:
NODE_OPTIONS="--max-old-space-size=1024" bench build
Or persist it in the bench user's shell:
echo 'export NODE_OPTIONS="--max-old-space-size=1024"' >> ~/.bashrc
The persistent form is safer because supervisor-managed background jobs that shell out to Node will also see the cap.
Can you build assets on a different machine?
Skimmable summary: yes — build on a bigger box, rsync sites/assets/ over. Standard practice for constrained production boxes.
If your production VPS is genuinely small (2 GB, no room to add swap, and MariaDB is already using most of it), doing the asset build on a separate host and copying the output over is a clean pattern:
-
On a build machine (a 4–8 GB VM you spin up for the upgrade — a Hetzner CX22 works fine and costs pennies for the hour):
- Clone the exact same bench directory (
git cloneof the bench repo +bench get-appfor every app on production). - Check out the same versions.
bench buildthere. Peak RSS 1.5–3 GB completes comfortably.
- Clone the exact same bench directory (
-
Rsync the built assets to the production box:
rsync -avz --delete \
--exclude 'assets/frappe/dist' \
/path/to/build-bench/sites/assets/ \
user@prod:/path/to/prod-bench/sites/assets/
- On production, run
bench update --pull --no-backup && bench update --patch && bench restart(skip--build). The assets are already there.
This is standard practice at hosting providers running many customer sites on constrained boxes. It also parallelises: you can build tomorrow's release on the build box while today's is still serving traffic.
The rest of the OOM budget: MariaDB and workers
bench build OOMs are the loudest, but two other processes commonly steal memory during an upgrade:
- MariaDB's InnoDB buffer pool sized too aggressively for the box. If you tuned
innodb_buffer_pool_sizeto 70% of RAM per the ERPNext MariaDB tuning guide, during an upgrade the build needs some of that back. Drop the buffer pool temporarily (SET GLOBAL innodb_buffer_pool_size=<smaller>) before the build, restore after. - The RQ workers. They keep running during
bench updateunless you stop them. On a 2 GB box,sudo supervisorctl stop <bench-name>:(whole group) before the upgrade andstartafter can free 200–400 MB.
Neither is required if you have adequate swap, but on a truly constrained box both are cheap wins.
FAQ
The four questions that come up most on this ticket class.
Why does bench update OOM on a 2 GB VPS when the machine was fine yesterday?
Because yesterday nothing was compiling JavaScript. The daily steady-state of ERPNext (gunicorn + workers + MariaDB + Redis) fits in 2 GB. Adding bench build's 1.5–3 GB Node peak on top pushes the machine into OOM. The "the server is too small" conclusion is wrong — the steady state is fine; only the upgrade blows the ceiling. Split the update and add swap and the same box will complete cleanly.
Is --no-backup safe in production?
It's safe if you take the backup yourself in a separate step before starting. --no-backup skips the bench update-managed backup; it does not delete data. Manually running bench --site all backup --with-files before, then running bench update --no-backup, is what most production upgrades actually do — you get a backup at a known time, and the update runs without the backup step pinning extra memory.
Does more swap always help?
Up to about 4× physical RAM, yes. Beyond that, the swap thrashing costs you more time than the equivalent VPS upgrade would. A 2 GB VPS with 8 GB of swap will complete bench build — but in an hour. If your upgrade window is smaller than that, pay for the bigger box for the day and downsize after.
Does exit code 137 always mean OOM on Linux?
Almost always. Exit 128 + N on Linux means the process was killed by signal N; 137 = 128 + 9 where 9 is SIGKILL. The kernel's OOM killer is by far the most common source of an unsolicited SIGKILL. The other sources — a manual kill -9, a Docker container hitting --memory limits, a systemd MemoryMax — leave different fingerprints in dmesg and journalctl. dmesg | grep -i oom gives you the definitive answer in one line.
bench update killed by OOM on your VPS?
We deploy and operate self-hosted ERPNext on Hetzner, AWS, DigitalOcean and bare-metal for Indian SMBs. If your upgrade window is closing and the kernel keeps killing bench build, we can help you get the upgrade through on the box you have — without buying hardware you may not need.