Laravel Horizon is the premium job queue dashboard for Laravel. It tracks your queues, shows you supervisor status, visualizes job throughput, and keeps a history of recent jobs. When it’s working, it’s excellent. When something goes wrong in production and nobody’s watching - you won’t know until your users start complaining.
What Horizon monitors for you
Horizon exposes a rich set of signals about your queue workers:
- Supervisor status - each Horizon supervisor is a group of queue workers. If all supervisors for a queue stop, jobs accumulate with no processing.
- Queue depth - how many jobs are waiting in each queue.
- Failed jobs per minute - a spike in failures is a clear incident signal.
- Paused state - Horizon queues can be paused manually. If this happens unintentionally, jobs queue up silently.
- Job timing - Horizon tracks how long jobs take, which helps spot degradation.
Horizon stores all of this in Redis. The dashboard reads from Redis, so it’s real-time - but only if someone is looking at it.
The monitoring gap
The Horizon dashboard is for humans actively using it. It doesn’t alert you when:
- All supervisors stop at 3am
- Failed jobs start accumulating over the weekend
- Queue depth climbs because a downstream service is slow
- Horizon itself goes down (Redis is fine, but the dashboard is unreachable)
You need monitoring layered on top of Horizon to catch these cases.
What to monitor in production
Supervisor status
Horizon’s /api/horizon/status endpoint returns the status of all supervisors. The key field is status - it can be active, paused, or inactive. When all supervisors for a given queue are inactive, that queue is not being processed at all.
If you’re using Crontinel, it reads this automatically from your Horizon setup. You can also query it directly:
curl -H "Authorization: Bearer YOUR_API_KEY" \
https://app.crontinel.com/api/horizon/status
A paused or inactive supervisor combined with a climbing queue depth is an immediate incident.
Queue depth per queue
Horizon tracks queueNames and their current depths. A queue that normally has 10-20 jobs and suddenly has 2,000 means workers aren’t keeping up. This could be because:
- Workers died (OOM kill, segfault)
- An external dependency is slow (database, Redis, HTTP calls)
- A job is running in a tight loop and blocking the queue
Set a threshold based on your normal workload. When depth exceeds it, alert.
Failed jobs per minute
Horizon exposes failed_per_minute as a rolling rate. A sustained rate above 0 on a production queue is always worth investigating - even if it’s low, it means something is consistently failing.
Paused supervisors
If Horizon’s paused_at is set and the pause wasn’t intentional, this is an incident. Someone (or some code) called php artisan horizon:pause and didn’t call horizon:continue. Jobs will queue up behind it.
Setting up Horizon monitoring with Crontinel
composer require crontinel/laravel
php artisan crontinel:install
The Crontinel package reads Horizon’s /api/horizon/status endpoint and reports:
- Supervisor health per queue
- Failed jobs per minute
- Paused state and when it was paused
- Queue depths across all queues
From the Crontinel dashboard, you can set alerts on all of these signals - so you get notified at 3am instead of finding out from a user report.
Common Horizon failures in production
Memory limit kills
Horizon supervisors restart when they hit PHP memory limits. If your jobs process large datasets without releasing memory, each restart drops any in-flight job. Configure memory limits conservatively:
'environments' => [
'production' => [
'supervisor-1' => [
'maxProcesses' => 10,
'memory' => 1024, // megabytes - restart at 1GB
'maxTime' => 3600, // restart after 1 hour
],
],
],
Monitoring catches this when you see a pattern of supervisors restarting frequently.
Long-running jobs blocking queues
A job that takes 10 minutes when most jobs take 10 milliseconds will consume a worker for that duration. If you have 2 workers and 3 long-running jobs queue up, your short jobs start timing out. Horizon shows job timing - use it to find the outliers.
Set timeout explicitly on long-running jobs:
ProcessReportJob::dispatch($report)
->onQueue('reports')
->timeout(600) // 10 minutes
->retryUntil(now()->addHours(1));
Redis connection drops
If Redis becomes unavailable, Horizon’s /api/horizon/status will return an error or timeout. If Crontinel can’t reach Horizon, it fires an alert: “Cannot reach Horizon status endpoint.” This is a distinct failure mode from slow queues - it means the entire queue system is unreachable, not just backed up.
Deployment race condition
During a deployment, old Horizon workers are killed and new ones start. If the old workers are killed before new ones are ready, there’s a brief window where no workers are processing jobs. Horizon supervisors show inactive during this window. A well-configured deployment should use horizon:pause before deploying and horizon:continue after new workers are ready.
Alerting thresholds
Set these based on your baseline:
| Signal | Alert threshold |
|---|---|
| Failed jobs/min | > 0 on any production queue |
| Supervisor status | Any inactive or paused (unintentional) |
| Queue depth | > 2x your p95 baseline |
| Oldest job age | > 5 minutes on time-sensitive queues |
| Horizon unreachable | Any timeout/error on /api/horizon/status |
Dashboard vs monitoring
Horizon’s dashboard is excellent for investigating current queue state - you can see which jobs are running, how long they’ve been running, and what failed. Use it for debugging. But don’t rely on someone having it open to catch incidents. Layer alerting on top so you’re notified automatically.
Crontinel monitors Horizon supervisors, queue depth, failed jobs, and job timing across all your apps. Get started in two commands.