Your Laravel queue worker died. You don’t know yet. Jobs are piling up. In an hour, three hours, or tomorrow morning, someone will notice that emails aren’t sending or that a batch job didn’t complete.
This is what queue worker death looks like: nothing. No error page. No alert. Just silence — and an ever-growing backlog.
How queue workers die
Out of memory kills. PHP is memory-hungry under load. If a job leaks memory — or if a worker processes enough jobs that heap fragmentation accumulates — the OS kills the process. Horizon’s --max-jobs and --max-time flags are the right solution, but many apps don’t set them.
Segfaults. PHP extensions (Redis, image processing, PDF libraries) occasionally segfault. The process terminates immediately with no PHP exception and no log entry.
Restart loops. Horizon tries to restart dead supervisors. If the restart itself fails (missing env variable, failed migration, corrupt session), Horizon enters a restart loop and stops processing jobs. The Horizon dashboard may still show “running” because the master process is alive.
Server reboots and deploys. Queue workers don’t survive server restarts unless they’re managed by a process supervisor like Supervisor or systemd. After a rolling deploy, workers that aren’t restarted continue running old code — or they die and aren’t replaced.
What “running” doesn’t mean
Horizon’s status dashboard reads Redis state. If the master process is alive and writing heartbeats, Horizon reports “running” — even if every supervisor under it has terminated. You can have a green Horizon dashboard with zero workers actually processing jobs.
This is the fundamental trap: Horizon’s UI tells you the orchestration layer is up, not that jobs are being processed.
How to detect a dead worker
Queue depth as a signal. If jobs are being processed, queue depth fluctuates. If workers have died, depth climbs monotonically. A queue depth alert (e.g., alert when depth exceeds 1000 for more than 5 minutes) catches dead workers without requiring direct process monitoring.
Job age as a signal. The oldest job in the queue tells you how long a job has been waiting. In a healthy system, old jobs don’t accumulate. An “oldest job age” alert (e.g., alert when any job has been waiting more than 15 minutes) is a reliable signal that nothing is consuming the queue.
Supervisor heartbeat monitoring. Horizon writes supervisor heartbeats to Redis. Monitoring the time since last heartbeat per supervisor catches paused or stopped supervisors faster than queue depth (which takes time to accumulate).
You can read these signals from Redis directly:
$paused = Redis::connection('horizon')->smembers('horizon:paused_supervisors');
Or query Horizon’s internal state:
$supervisors = app(SupervisorRepository::class)->all();
foreach ($supervisors as $supervisor) {
if (! $supervisor->status === 'running') {
// alert
}
}
Recovery
Restart Horizon:
php artisan horizon:terminate
# Process supervisor (Supervisor/systemd) restarts it automatically
Check for failed jobs:
php artisan queue:failed
php artisan queue:retry all
For OOM issues, reduce --max-jobs and --max-memory:
php artisan horizon --max-jobs=100 --max-memory=256
Or set these in your horizon.php config per supervisor.
The monitoring gap
Manual recovery fixes the current incident. It doesn’t tell you how long workers were dead, which jobs were affected, or whether this has happened before.
Crontinel monitors Horizon supervisor status directly from Redis — it detects a paused or stopped supervisor the moment it happens, not after the queue backlog reaches a threshold you notice manually.
composer require crontinel/laravel
php artisan crontinel:install
When a supervisor dies, you get an alert. When it recovers, you get a resolve. No more discovering dead workers from customer complaints.