Bragi Docs Help

Scheduler Health

Bragi keeps an eye on the BragiScheduler process itself and surfaces a clear in-app signal if it stops responding. This is the difference between "everything looks fine in the UI but actually nothing is running" and "Bragi tells you the scheduler is down and what to do about it".

How it works

Every running BragiScheduler process writes a heartbeat row to the metadata database every 10 seconds. Each per-environment worker inside the process also records a "work loop tick" at the end of every scheduling cycle. The Bragi web app reads these two timestamps and decides on a per-environment health verdict.

The app uses two independent signals so it can distinguish a crashed process from one that is alive but stuck:

Status

Meaning

Healthy

Heartbeat is fresh and the work loop is progressing.

Stuck

Heartbeat is fresh, but the work loop has not advanced for several minutes. The process is alive but the per-environment worker has not made progress - typically a hung job or a deadlock.

Dead

No heartbeat has arrived for this environment in over a minute. The scheduler process is presumed to have crashed or become unreachable.

Never started

No scheduler has ever registered a heartbeat for this environment. Normally only seen on a brand-new deployment.

What changes in the UI

When the scheduler is anything other than Healthy:

  • A coloured icon appears in the app bar at the top of every page. Hover over it for a quick per-environment summary; click it to jump to the Scheduler page.

  • The Scheduler and History pages show a banner at the top explaining what's wrong, the machine name the scheduler was running on, its version, and when it last reported.

  • The elapsed "Running for" timer on jobs that were Running when the scheduler stopped freezes at the last moment we know the scheduler was healthy. A small snowflake icon next to the time indicates this. This prevents a misleading timer from counting upwards indefinitely after the scheduler is gone.

  • The Scheduler diagnostics expander on the Scheduler page shows the instance id, machine name, process id, version, started-at, last heartbeat and last work loop tick. Screenshot this expander straight into a support request.

Notifications

When the scheduler transitions from Healthy to Dead or Stuck, an email goes out to whichever addresses are configured for failed-job emails for that environment (the same list configured under Settings). A reminder is sent again every 24 hours while the issue persists. No "recovered" email is sent - the banner clearing is the signal.

What to do when you see Dead

You need someone with access to the scheduler host to act. Useful copy/paste from the banner:

Once the service is restarted, any jobs left as Running are automatically moved to Failed with a recovery message that identifies the machine that crashed. No manual cleanup is required.

What to do when you see Stuck

The process is responding but its work loop is jammed - usually a long-running job or a hung external call. The team responsible for the scheduler host should inspect the scheduler log on that machine before restarting, because killing it will leave any genuinely in-flight work to be cleaned up by FixUpEdgeCases on the next start.

Tuning thresholds

Defaults are sensible for most deployments:

Setting (in SchedulerHealth section of appsettings.json)

Default

HeartbeatInterval

10 seconds

HeartbeatDeadAfter

60 seconds (two missed heartbeats)

WorkLoopStuckAfter

5 minutes

MonitorInterval

30 seconds

UnhealthyReminderInterval

24 hours

PurgeAfter

30 days (how long to retain old heartbeat rows)

If your environment has individual scheduled jobs that legitimately run longer than five minutes, raise WorkLoopStuckAfter accordingly - otherwise the work-loop tick will trip the Stuck verdict while a normal slow job is in flight.

Multi-scheduler environments

The heartbeat schema supports multiple scheduler instances serving the same environment, with the per-environment verdict being the best across all instances. Today's deployments still run one scheduler per environment by convention - the data model is ready for a future change to that model.

02 October 2026