Skip to main content
All posts
· 5 min read

Laravel Queue Worker Died: How to Detect and Recover

Queue workers die quietly. OOM kills, segfaults, and restart loops leave no obvious trace while jobs pile up. Here's how to detect a dead worker before your queue backlog becomes a customer problem.

Your Laravel queue worker died. You don’t know yet. Jobs are piling up. In an hour, three hours, or tomorrow morning, someone will notice that emails aren’t sending or that a batch job didn’t complete.

This is what queue worker death looks like: nothing. No error page. No alert. Just silence — and an ever-growing backlog.

How queue workers die

Out of memory kills. PHP is memory-hungry under load. If a job leaks memory — or if a worker processes enough jobs that heap fragmentation accumulates — the OS kills the process. Horizon’s --max-jobs and --max-time flags are the right solution, but many apps don’t set them.

Segfaults. PHP extensions (Redis, image processing, PDF libraries) occasionally segfault. The process terminates immediately with no PHP exception and no log entry.

Restart loops. Horizon tries to restart dead supervisors. If the restart itself fails (missing env variable, failed migration, corrupt session), Horizon enters a restart loop and stops processing jobs. The Horizon dashboard may still show “running” because the master process is alive.

Server reboots and deploys. Queue workers don’t survive server restarts unless they’re managed by a process supervisor like Supervisor or systemd. After a rolling deploy, workers that aren’t restarted continue running old code — or they die and aren’t replaced.

What “running” doesn’t mean

Horizon’s status dashboard reads Redis state. If the master process is alive and writing heartbeats, Horizon reports “running” — even if every supervisor under it has terminated. You can have a green Horizon dashboard with zero workers actually processing jobs.

This is the fundamental trap: Horizon’s UI tells you the orchestration layer is up, not that jobs are being processed.

How to detect a dead worker

Queue depth as a signal. If jobs are being processed, queue depth fluctuates. If workers have died, depth climbs monotonically. A queue depth alert (e.g., alert when depth exceeds 1000 for more than 5 minutes) catches dead workers without requiring direct process monitoring.

Job age as a signal. The oldest job in the queue tells you how long a job has been waiting. In a healthy system, old jobs don’t accumulate. An “oldest job age” alert (e.g., alert when any job has been waiting more than 15 minutes) is a reliable signal that nothing is consuming the queue.

Supervisor heartbeat monitoring. Horizon writes supervisor heartbeats to Redis. Monitoring the time since last heartbeat per supervisor catches paused or stopped supervisors faster than queue depth (which takes time to accumulate).

You can read these signals from Redis directly:

$paused = Redis::connection('horizon')->smembers('horizon:paused_supervisors');

Or query Horizon’s internal state:

$supervisors = app(SupervisorRepository::class)->all();
foreach ($supervisors as $supervisor) {
    if (! $supervisor->status === 'running') {
        // alert
    }
}

Recovery

Restart Horizon:

php artisan horizon:terminate
# Process supervisor (Supervisor/systemd) restarts it automatically

Check for failed jobs:

php artisan queue:failed
php artisan queue:retry all

For OOM issues, reduce --max-jobs and --max-memory:

php artisan horizon --max-jobs=100 --max-memory=256

Or set these in your horizon.php config per supervisor.

The monitoring gap

Manual recovery fixes the current incident. It doesn’t tell you how long workers were dead, which jobs were affected, or whether this has happened before.

Crontinel monitors Horizon supervisor status directly from Redis — it detects a paused or stopped supervisor the moment it happens, not after the queue backlog reaches a threshold you notice manually.

composer require crontinel/laravel
php artisan crontinel:install

When a supervisor dies, you get an alert. When it recovers, you get a resolve. No more discovering dead workers from customer complaints.

See also

blog
Laravel Queue Worker Memory Leaks: Detect Them Before Workers Crash

Queue workers run for hours or days. Memory climbs slowly until the process is killed. Here's how memory leaks happen in Laravel workers, how to detect them, and how to keep workers healthy without masking the real problem.

blog
How to Detect Laravel Queue Worker Stalls: When Workers Go Silent

A Laravel queue worker that's alive but not processing jobs is worse than a crashed worker — it gives no alert, no error, and no warning. Here's how to detect stalled workers, common causes, and how Crontinel catches silent failures before they compound.

blog
Laravel Horizon Workers Idle While Jobs Pile Up? 6 Causes (And Fixes)

Horizon shows workers as running while 2,000 jobs pile up untouched. Here are the 6 real causes — and how to fix each one fast.

use cases
Detect When Laravel Horizon Workers Are Running But Not Processing Jobs

Your Horizon dashboard shows active supervisors and workers, but jobs sit in the queue for minutes. Here is how to catch worker starvation before it turns into a production incident.