Skip to main content
All posts
· 5 min read

How to Monitor Horizon Jobs in Production

Laravel Horizon gives you a beautiful dashboard - but if you're not watching it in production, you won't know when supervisors stop, queues pile up, or failed jobs start accumulating. Here's how to set up proper Horizon monitoring.

Laravel Horizon is the premium job queue dashboard for Laravel. It tracks your queues, shows you supervisor status, visualizes job throughput, and keeps a history of recent jobs. When it’s working, it’s excellent. When something goes wrong in production and nobody’s watching - you won’t know until your users start complaining.

What Horizon monitors for you

Horizon exposes a rich set of signals about your queue workers:

  • Supervisor status - each Horizon supervisor is a group of queue workers. If all supervisors for a queue stop, jobs accumulate with no processing.
  • Queue depth - how many jobs are waiting in each queue.
  • Failed jobs per minute - a spike in failures is a clear incident signal.
  • Paused state - Horizon queues can be paused manually. If this happens unintentionally, jobs queue up silently.
  • Job timing - Horizon tracks how long jobs take, which helps spot degradation.

Horizon stores all of this in Redis. The dashboard reads from Redis, so it’s real-time - but only if someone is looking at it.

The monitoring gap

The Horizon dashboard is for humans actively using it. It doesn’t alert you when:

  • All supervisors stop at 3am
  • Failed jobs start accumulating over the weekend
  • Queue depth climbs because a downstream service is slow
  • Horizon itself goes down (Redis is fine, but the dashboard is unreachable)

You need monitoring layered on top of Horizon to catch these cases.

What to monitor in production

Supervisor status

Horizon’s /api/horizon/status endpoint returns the status of all supervisors. The key field is status - it can be active, paused, or inactive. When all supervisors for a given queue are inactive, that queue is not being processed at all.

If you’re using Crontinel, it reads this automatically from your Horizon setup. You can also query it directly:

curl -H "Authorization: Bearer YOUR_API_KEY" \
  https://app.crontinel.com/api/horizon/status

A paused or inactive supervisor combined with a climbing queue depth is an immediate incident.

Queue depth per queue

Horizon tracks queueNames and their current depths. A queue that normally has 10-20 jobs and suddenly has 2,000 means workers aren’t keeping up. This could be because:

  • Workers died (OOM kill, segfault)
  • An external dependency is slow (database, Redis, HTTP calls)
  • A job is running in a tight loop and blocking the queue

Set a threshold based on your normal workload. When depth exceeds it, alert.

Failed jobs per minute

Horizon exposes failed_per_minute as a rolling rate. A sustained rate above 0 on a production queue is always worth investigating - even if it’s low, it means something is consistently failing.

Paused supervisors

If Horizon’s paused_at is set and the pause wasn’t intentional, this is an incident. Someone (or some code) called php artisan horizon:pause and didn’t call horizon:continue. Jobs will queue up behind it.

Setting up Horizon monitoring with Crontinel

composer require crontinel/laravel
php artisan crontinel:install

The Crontinel package reads Horizon’s /api/horizon/status endpoint and reports:

  • Supervisor health per queue
  • Failed jobs per minute
  • Paused state and when it was paused
  • Queue depths across all queues

From the Crontinel dashboard, you can set alerts on all of these signals - so you get notified at 3am instead of finding out from a user report.

Common Horizon failures in production

Memory limit kills

Horizon supervisors restart when they hit PHP memory limits. If your jobs process large datasets without releasing memory, each restart drops any in-flight job. Configure memory limits conservatively:

'environments' => [
    'production' => [
        'supervisor-1' => [
            'maxProcesses' => 10,
            'memory' => 1024, // megabytes - restart at 1GB
            'maxTime' => 3600, // restart after 1 hour
        ],
    ],
],

Monitoring catches this when you see a pattern of supervisors restarting frequently.

Long-running jobs blocking queues

A job that takes 10 minutes when most jobs take 10 milliseconds will consume a worker for that duration. If you have 2 workers and 3 long-running jobs queue up, your short jobs start timing out. Horizon shows job timing - use it to find the outliers.

Set timeout explicitly on long-running jobs:

ProcessReportJob::dispatch($report)
    ->onQueue('reports')
    ->timeout(600) // 10 minutes
    ->retryUntil(now()->addHours(1));

Redis connection drops

If Redis becomes unavailable, Horizon’s /api/horizon/status will return an error or timeout. If Crontinel can’t reach Horizon, it fires an alert: “Cannot reach Horizon status endpoint.” This is a distinct failure mode from slow queues - it means the entire queue system is unreachable, not just backed up.

Deployment race condition

During a deployment, old Horizon workers are killed and new ones start. If the old workers are killed before new ones are ready, there’s a brief window where no workers are processing jobs. Horizon supervisors show inactive during this window. A well-configured deployment should use horizon:pause before deploying and horizon:continue after new workers are ready.

Alerting thresholds

Set these based on your baseline:

SignalAlert threshold
Failed jobs/min> 0 on any production queue
Supervisor statusAny inactive or paused (unintentional)
Queue depth> 2x your p95 baseline
Oldest job age> 5 minutes on time-sensitive queues
Horizon unreachableAny timeout/error on /api/horizon/status

Dashboard vs monitoring

Horizon’s dashboard is excellent for investigating current queue state - you can see which jobs are running, how long they’ve been running, and what failed. Use it for debugging. But don’t rely on someone having it open to catch incidents. Layer alerting on top so you’re notified automatically.


Crontinel monitors Horizon supervisors, queue depth, failed jobs, and job timing across all your apps. Get started in two commands.

See also

use cases
Detect When Laravel Horizon Workers Are Running But Not Processing Jobs

Your Horizon dashboard shows active supervisors and workers, but jobs sit in the queue for minutes. Here is how to catch worker starvation before it turns into a production incident.

use cases
Monitor Laravel pulse:restart in Production

pulse:restart runs silently during deploys and fails silently too. If the new Pulse worker cannot start, you won't know until someone notices the dashboard froze. Here is how to catch restart failures before they cost you Pulse data.

blog
Laravel Horizon Shows Running But Jobs Aren't Processing

Horizon's status dashboard says everything is fine. Jobs are silently piling up. Here's why that happens and how to actually detect a dead supervisor.

blog
What Happens to Horizon When Redis Drops: Data Loss, Recovery, and Monitoring

Redis goes down and Horizon loses its connection. Jobs vanish, supervisors crash, and the dashboard goes blank. Here's exactly what breaks, what recovers automatically, and what you lose permanently.