Skip to main content
ERPNext how-to & fixes

ERPNext v16 Background Jobs Not Running: A Diagnostic Ladder

Emails not sending, reports not building, scheduled jobs silent — the five reasons ERPNext background jobs actually stop in v16, and the exact bench + redis-cli commands to walk from symptom to fix on Docker and bare-metal.

MManojAugust 9, 202612 min read
More in ERPNext how-to & fixes#v16#troubleshooting
Share

When ERPNext background jobs not running is the ticket, the symptoms are almost always the same in v16: emails queued but never sent, scheduled reports frozen at yesterday's date, assign to notifications silent, imports and long exports that never finish. The cause is one of five things — the scheduler is disabled at the site level, the RQ worker container is dead, the Redis queue is unreachable, the long queue has a poisoned job blocking it, or pause_scheduler was flipped on and forgotten. This post is the diagnostic ladder we walk in that exact order, with the bench and redis-cli commands to prove each one.

A silent scheduler is worse than a loud one. Emails don't bounce, users don't complain until the next AR call, and by then a week of automated flows have quietly not happened. I lead enterprise ERPNext and Frappe implementation programmes at MithTech, a 35-person practice in Bengaluru that designs, customises, integrates and operates business-critical software for manufacturers, distributors, schools, financial services and multi-location operations. This is the runbook we hand every self-hosted client after their v15→v16 upgrade.

Is the scheduler even running? Start with bench doctor

Answer

bench doctor is the fastest single command for background-job triage in ERPNext v16 — it prints the scheduler status of every site on the bench, the number of RQ workers currently attached to Redis, and how many jobs are pending per queue. Any obvious failure lights up here first, before you go container-diving.

Skimmable summary: run bench doctor first; three of the five failure modes fail visibly at this step.

Run it from your bench root — inside the container on Docker, on the host on bare-metal:

bench doctor

You want to see, for each site: Scheduler is enabled, a non-zero worker count under the default, short and long queues, and a low or zero pending count. If any of those three is off, you have your suspect.

Two useful companions:

# What's actually queued right now
bench --site your-site.local show-pending-jobs

# Enable a per-site scheduler that got flipped off during upgrade
bench --site your-site.local enable-scheduler

show-pending-jobs prints the queue name followed by the tasks currently sitting in it. A long queue with dozens of jobs and a default queue that's empty means the workers on long are the ones you want to look at, not everything.

How do you tell a stuck worker from an empty queue?

The single most common misdiagnosis is confusing "no jobs running" with "no jobs queued." They look identical from the desk. frappe.enqueue(...) returns a job id, the desk cheerfully shows "queued," and nothing happens next.

Redis is where the truth lives. The queue's raw state is what RQ (Redis Queue) exposes and what bench wraps. From the bench container:

# Show the raw queue lengths
redis-cli -h redis-queue -p 6379 llen rq:queue:default
redis-cli -h redis-queue -p 6379 llen rq:queue:short
redis-cli -h redis-queue -p 6379 llen rq:queue:long

# Show workers currently registered
redis-cli -h redis-queue -p 6379 smembers rq:workers

If llen is positive but rq:workers is empty, you have jobs in the queue and no worker to pick them up — a dead worker container, or on bare-metal a Supervisor group that failed to restart. If rq:workers has entries but llen keeps growing, workers are alive but slow — usually the long queue chewing through a single expensive job that blocks everything behind it.

Why do ERPNext v16 background jobs actually stop?

Skimmable summary: ranked roughly by how often we see them in v15→v16 upgrade tickets.

SignatureFirst check
1Scheduler disabled per sitebench doctor says "Scheduler is disabled" or scheduler_enabled: false in sites/SITE/site_config.jsonbench --site SITE scheduler status
2RQ worker container deadQueues have length, smembers rq:workers is emptydocker compose ps queue-long queue-short (Docker) or supervisorctl status (bare-metal)
3Redis queue unreachableWorkers restart-loop, logs show ConnectionError to redis-queue:6379redis-cli -h redis-queue ping from a worker
4long queue backlog on one poisoned joblong queue length climbs monotonically, one job stays "started" for hoursredis-cli hgetall rq:job:JOB_ID
5pause_scheduler or maintenance modeEverything looks healthy but no ticks firegrep pause_scheduler sites/SITE/site_config.json

Every row below matches a cause. Walk them in order — the check at the top is cheaper than the check at the bottom.

What is the fix order on frappe_docker?

Confirm the scheduler service is actually up

docker compose ps scheduler should show running with a recent uptime. If it's restarting, docker compose logs --tail=200 scheduler shows why — the two common ones are a MariaDB that isn't ready when the scheduler tries to connect (fix: raise depends_on health), and a stale common_site_config.json from a broken previous install (fix: docker compose down -v in staging only, or re-run the configurator service).

Confirm both worker containers are up

The standard compose.yaml runs two: queue-short (bench worker --queue short,default) and queue-long (bench worker --queue long,default,short). Both need to be running. If only one is up, jobs on the missing queue never move. Restart with docker compose up -d queue-short queue-long, then re-check with docker compose logs --tail=100 queue-long for a ConnectionError.

Check the shared Redis queue

From inside docker compose exec backend bash, run redis-cli -h redis-queue ping. Expect PONG. Anything else — timeout, Connection refused, or a NOAUTH — is a broken FRAPPE_REDIS_QUEUE env var or a Redis container that failed its own healthcheck. Look at docker compose logs redis-queue.

Re-enable the scheduler if it was flipped off

The upgrade to v16 does not automatically re-enable a scheduler that was paused mid-upgrade. From backend: bench --site your-site.local enable-scheduler. Confirm with bench --site your-site.local scheduler status.

Deal with a stuck `long` queue

If steps 1–4 pass and long is still growing, one job is holding the queue. List currently started jobs: redis-cli -h redis-queue keys "rq:job:*" and inspect the started ones with redis-cli -h redis-queue hgetall rq:job:<id>. A stuck import or a mis-written scheduled method that loops silently is the usual culprit. bench --site <site> purge-jobs --queue long clears it — but only after you've noted the job function so you can fix the caller.

What is the fix order on bare-metal bench?

On a Supervisor-managed bare-metal install, the anatomy is the same but the levers are different.

Check the Supervisor group

sudo supervisorctl status | grep -E "(scheduler|worker)" should show four processes running: one scheduler, one each for default, short, long workers (the group is defined in /etc/supervisor/conf.d/<bench>.conf). Any FATAL or EXITED is your cause. Restart the group with sudo supervisorctl restart <bench-name>: (trailing colon = whole group).

Read the tails

Every worker writes to <bench>/logs/worker.log (rotated). tail -F logs/worker.log logs/scheduler.log while you trigger a small job (bench --site <site> execute frappe.utils.background_jobs.enqueue --args '{"method": "frappe.ping"}') tells you whether jobs are being accepted and whether they are being executed.

Reconcile Redis

ps aux | grep redis-server should show two Redis processes on a standard bench: one on the cache port, one on the queue port (configured in <bench>/config/redis_cache.conf and config/redis_queue.conf). If only one is running, the queue Redis died — the scheduler and workers will silently fail. Restart with bench setup redis and re-run Supervisor.

Rule out the config-only flags

Open sites/<site>/site_config.json and check three keys: scheduler_enabled (should be 1 or absent), pause_scheduler (should be 0 or absent), maintenance_mode (should be 0). Any of these flipped will silently stop the scheduler without any error in the logs.

When should you run bench purge-jobs — and when should you not?

bench purge-jobs clears all pending jobs across all queues on the site. It is a legitimate step when a queue is provably stuck and you have identified the offending job — after the fix. It is a terrible first step, because you lose the evidence that would tell you why the queue jammed.

The safe order is: show-pending-jobs → note the function names → redis-cli hgetall rq:job:<id> on any long-started job → fix the caller → then purge-jobs.

bench purge-jobs --queue long scopes the clear to one queue, which is almost always what you want. Losing scheduled emails from default because a data import wedged long is a common self-inflicted wound.

How do you monitor so this stops surprising you?

Skimmable summary: two low-effort signals catch every failure mode above before users file tickets.

Once ERPNext background jobs not running has been diagnosed once, the two signals worth wiring into whatever monitor you already have are:

  1. Scheduled Job Log freshness. Query tabScheduled Job Log for the newest modified timestamp. In v16 the default scheduler tick is 240 seconds; anything older than roughly three ticks (~12 minutes) means the scheduler has stopped enqueueing. This is a database-backed signal, so it works across split containers where a process check cannot.
  2. Queue length trend. Every 60 seconds, redis-cli -h redis-queue llen rq:queue:long and record it. A monotonic climb over an hour is the failure signature — normal queues drain between spikes.

Both are two-line shell scripts feeding your existing metrics stack. The one thing you should not do is set a Frappe Notification to alert you when the scheduler is down, for the same reason a smoke alarm shouldn't need mains power: a scheduler-driven alert cannot tell you the scheduler stopped.

Three adjacent problems that present as "jobs not running" but actually aren't:

  • Slow database making jobs look stuck. A long-running SQL join inside a scheduled method will pin one worker for minutes and starve everything behind it. If jobs are running but very slowly, start with MariaDB tuning for ERPNext — a right-sized InnoDB buffer pool often turns "stuck" into "fine" without touching the queue.
  • Asset 404s after bench build are unrelated to jobs but frequently misattributed — a broken desk that can't load JavaScript makes new jobs impossible to trigger from the UI. If your workers are healthy but nothing gets enqueued from the desk, see the frappe_docker asset breakage guide.
  • The System Health Report shows red but jobs are actually fine. That is the file-lock false negative — link at the top of this post. Do not chase infrastructure fixes for a diagnostic bug.

FAQ

The four questions we get asked most on this after an upgrade.

Why are ERPNext background jobs not running right after my v15 to v16 upgrade?

The most common reason is the scheduler being paused mid-upgrade and not re-enabled by the bench migrate step. Run bench --site <site> scheduler status; if it says disabled, bench --site <site> enable-scheduler fixes it in one command. The second most common is a stale common_site_config.json where redis_queue still points at the old bench's Redis socket path — restart Supervisor or the queue containers after any config edit.

How do I know if the RQ worker is running or just registered?

redis-cli -h redis-queue smembers rq:workers lists worker IDs Redis knows about. Combine with redis-cli -h redis-queue hgetall rq:worker:<worker-id> — a healthy worker updates its last_heartbeat every few seconds. A worker in Redis with a stale heartbeat is registered but not actually running; a common failure mode when a container crashed without cleanly unregistering.

Can I have separate workers for long jobs so they don't block the queue?

Yes, and on any deployment with regular imports or large reports you should. The standard frappe_docker compose already isolates queue-long from queue-short as separate containers — if you're running bare-metal Supervisor, edit /etc/supervisor/conf.d/<bench>.conf to give the long group more processes (typical: 2–4 instead of 1). Overprovision long, not short.

What does bench purge-jobs actually delete?

Every pending job in every RQ queue for the site — default, short, and long. Jobs currently executing are not killed; they continue to completion. Scheduled events are not affected — the scheduler will re-enqueue them on the next tick. If a queue is truly wedged behind one bad job, purge-jobs --queue long scoped to that queue is the safer form.

Background jobs stopped on your ERPNext instance?

We deploy and operate ERPNext on Docker and bare-metal for multi-entity operators and complex organisations. If the scheduler or RQ workers have gone silent and the ladder above did not resolve it, we will help you work out where the queue is blocked and why.

Next step · ERPNext how-to & fixes

Now see support that keeps it fixed

Support and maintenance for ERPNext you already run: fixes, upgrades and the recurring issues your team should not have to chase.

M

Written by

Manoj

Founder of MithTech, an open-source ERP & automation engineering practice. Hands-on ERPNext/Frappe implementation across multi-branch, multi-warehouse Indian operations — GST/TDS/PT compliance, branch-level permissions, and custom Frappe apps that give management real-time visibility.

Free · By email

Get practical ERPNext & automation guides

New implementation guides, cost breakdowns and open-source tips for Indian businesses — occasionally, straight to your inbox. No spam.

Already a MithTech client?

Help the next operator choose.

Most teams evaluating ERPNext have no way to tell who actually delivers. If we’ve run an implementation for you, two lines on Google count for more than anything we can write about ourselves.

Leave a Google review

Only if we’ve actually worked together — Google filters reviews from non-customers, so an honest one is worth more than ten polite ones.

Keep reading

See what this looks like for your business

A 30-minute working session with a principal consultant. We pressure-test the architecture and outline the engagement model that fits your governance and procurement posture. You leave with a written brief.

0
Published on 9 August 2026

Manoj

Comments & ratings

No comments yet. Start a new discussion.