Skip to main content
All posts
· 5 min read

Python Scheduled Task Monitoring: Detect APScheduler, Celery Beat & Cron Failures

How to monitor Python scheduled tasks — APScheduler, Celery Beat, system cron, and custom schedulers. Detect missed runs, silent failures, and queue backpressure before users notice.

Python powers everything from data pipelines to background job processing. But unlike Laravel or Rails, Python doesn’t have a single dominant scheduling framework — teams use APScheduler, Celery Beat, system cron, or custom scheduling logic. Each has different failure modes, and most teams discover problems when a pipeline stops producing reports or a data sync falls behind.

This guide covers how to monitor Python scheduled tasks across every common setup, what to watch for beyond “did it run,” and how to catch failures before they become customer-facing issues.

The Python Scheduling Landscape

Unlike PHP frameworks with built-in schedulers, Python teams typically use one of these:

ToolTypical UseScheduling Model
APSchedulerIn-process scheduled jobsCron-like triggers, interval, date-based
Celery BeatDistributed task schedulingCrontab entries, periodic task registry
System cronShell scripts, management commandsTraditional crontab entries
Custom schedulersData pipelines, ETLDatabase-driven or in-memory queues
AsyncIO tasksFastAPI/async servicesasyncio.sleep loops, task queues

Each has different failure modes. A system cron job that silently fails produces no logs. A Celery Beat task that loses its schedule after a restart runs forever without anyone noticing. An APScheduler job stuck in a retry loop burns CPU without doing useful work.

What to Monitor

1. Heartbeat — Did It Run?

The foundation. Your task sends a signal when it starts or finishes. If the signal doesn’t arrive within the expected window, something is wrong.

Without monitoring: A Celery Beat task stops running at 2am. You discover it at 9am when the daily report doesn’t arrive.

With monitoring: Alert fires at 2:05am. You fix it before anyone notices.

# APScheduler heartbeat example
from apscheduler.schedulers.blocking import BlockingScheduler

scheduler = BlockingScheduler()

@scheduler.scheduled_job('cron', hour=2, minute=0)
def daily_report():
    # Signal that we're alive
    crontinel.ping("daily-report")
    
    # Do the work
    generate_report()
    
    # Signal completion
    crontinel.complete("daily-report")

2. Duration — Is It Slow?

A job that usually takes 30 seconds suddenly takes 10 minutes. It still “works,” but something is wrong — database lock, API rate limit, memory pressure.

Watch for:

  • Jobs taking 2x+ longer than their historical average
  • Gradual duration increase over days (resource leak)
  • Spike in duration correlating with traffic or data volume

3. Exit Code — Did It Succeed?

A job runs, finishes, but exits with code 1. In system cron, this produces no output unless you explicitly redirect stderr. In Celery, it raises an exception. In APScheduler, it might silently swallow the error.

# Bad: exception is caught and forgotten
try:
    process_data()
except Exception as e:
    logger.error(f"Failed: {e}")  # Logged, but who's watching?

# Good: signal failure to monitoring
try:
    process_data()
    crontinel.complete("process-data")
except Exception as e:
    crontinel.fail("process-data", error=str(e))
    raise

4. Queue Depth — Is Work Backing Up?

If your scheduled tasks push jobs to a queue (Redis, RabbitMQ, database), monitor the queue depth. A growing queue means workers can’t keep up.

Redis queue monitoring:

import redis

r = redis.Redis()

def check_queue_depth(queue_name="celery"):
    depth = r.llen(queue_name)
    if depth > 1000:
        alert(f"Queue {queue_name} has {depth} pending jobs")
    return depth

5. Schedule Drift — Is It Running on Time?

A task scheduled for 2:00am starts running at 2:15am because the previous run took too long. Or timezone changes cause jobs to fire at the wrong time.

With APScheduler:

scheduler = BlockingScheduler()

@scheduler.scheduled_job('cron', hour=2, minute=0, misfire_grace_time=300)
def daily_sync():
    # If this fires at 2:20am instead of 2:00am, something is wrong
    crontinel.track("daily-sync", expected="02:00")
    run_sync()

Monitoring APScheduler

APScheduler runs in-process, so if your application crashes, all scheduled jobs stop silently.

Setup

from apscheduler.schedulers.asyncio import AsyncIOScheduler
from apscheduler.events import EVENT_JOB_ERROR, EVENT_JOB_MISSED

scheduler = AsyncIOScheduler()

def on_job_error(event):
    crontinel.fail(f"job-{event.job_id}", error=str(event.exception))

def on_job_missed(event):
    crontinel.fail(f"job-{event.job_id}", error="Missed scheduled run")

scheduler.add_listener(on_job_error, EVENT_JOB_ERROR)
scheduler.add_listener(on_job_missed, EVENT_JOB_MISSED)

# Add jobs with heartbeat
@scheduler.scheduled_job('interval', minutes=5, id='health-check')
def health_check():
    crontinel.ping("health-check")
    check_services()
    crontinel.complete("health-check")

scheduler.start()

Common Failures

  1. Scheduler stops after app restart — No persistence configured. Jobs exist only in memory.
  2. Job throws exception — Without EVENT_JOB_ERROR listener, the error is logged but nobody sees it.
  3. Job misses its window — Without misfire_grace_time, APScheduler skips late jobs silently.

Monitoring Celery Beat

Celery Beat is a separate process that enqueues tasks on schedule. If Beat stops, no new tasks are created. If workers stop, tasks pile up in the queue.

Setup

# celery.py
from celery import Celery
from celery.signals import task_success, task_failure, task_prerun

app = Celery('myapp')

@task_prerun.connect
def task_prerun(sender, task_id, **kwargs):
    crontinel.ping(f"celery-{sender.name}")

@task_success.connect
def task_success(sender, result, **kwargs):
    crontinel.complete(f"celery-{sender.name}")

@task_failure.connect
def task_failure(sender, task_id, exception, **kwargs):
    crontinel.fail(f"celery-{sender.name}", error=str(exception))

Watch for Beat Process Health

# Check if Beat is running
ps aux | grep celery beat

# Check Celery flower dashboard
curl http://localhost:5555/api/workers

Common Failures

  1. Beat process dies — No new tasks are scheduled. Everything else looks fine.
  2. Worker pool exhausted — All workers busy, new tasks queue up.
  3. Task loses state after restart — Database-backed scheduler needed for persistence.

Monitoring System Cron

System cron is the simplest but least observable. Jobs run as shell commands. If they fail, you need explicit logging.

Setup

# crontab entry with logging
0 2 * * * /usr/bin/python3 /app/manage.py daily_report >> /var/log/cron/daily_report.log 2>&1; \
  [ $? -eq 0 ] && echo "$(date): SUCCESS" >> /var/log/cron/monitor.log || echo "$(date): FAILED" >> /var/log/cron/monitor.log

Better: Use a wrapper script

#!/bin/bash
# /usr/local/bin/cron-wrapper.sh
JOB_NAME=$1
shift

crontinel ping "$JOB_NAME"
"$@"
EXIT_CODE=$?

if [ $EXIT_CODE -eq 0 ]; then
    crontinel complete "$JOB_NAME"
else
    crontinel fail "$JOB_NAME" --error "Exit code: $EXIT_CODE"
fi

exit $EXIT_CODE
# crontab
0 2 * * * cron-wrapper.sh daily-report python3 /app/manage.py daily_report

Python SDK Installation

Crontinel provides a Python SDK that works with any scheduler:

pip install crontinel
from crontinel import Crontinel

crontinel = Crontinel(api_key="your-key")

# Simple heartbeat
crontinel.ping("my-task")
# ... do work ...
crontinel.complete("my-task")

# With context
crontinel.ping("my-task", metadata={"env": "production", "region": "us-east-1"})

Quick Comparison

FeatureAPSchedulerCelery BeatSystem CronCrontinel
Missed run detection✅ (with listener)❌❌✅
Duration tracking❌❌❌✅
Queue depthN/A✅ (with Flower)N/A✅
Auto-alerts❌❌❌✅
Works with all schedulers✅✅✅✅

See also

blog
Node.js Cron Monitoring: Detect Scheduled Task Failures Before Users Do

Complete guide to monitoring Node.js cron jobs, scheduled tasks, and background workers. Detect missed runs, silent failures, and queue backpressure with node-cron, Bull, and custom schedulers.

blog
Cron Monitoring Guide 2026: Detect Failures Before Users Do

Complete guide to cron monitoring for engineering teams. Learn how to detect missed schedule runs, silent failures, worker stalls, and queue backpressure — with setup examples for any framework.

blog
How to Set Up Laravel Cron Monitoring (The Right Way)

A step-by-step guide to setting up Laravel cron monitoring that actually works. Detect missed schedules, failed tasks, and silent cron failures before your users notice.

use cases
Laravel Cron Monitoring for Scheduled Tasks, Missed Runs, and Failures

Monitor every Laravel scheduled task in production so you catch missed runs, non-zero exits, and silent scheduler outages without instrumenting each task individually.