Job Callbacks
Job Callbacks let Bragi tell an external system when a Job run has finished. When the run ends, Bragi posts a small JSON payload to a URL you nominate, so the other system does not have to poll Bragi to find out what happened.
There are two kinds:
One shot - registered when a Job is triggered through the API, or when someone uses Run Now in the Bragi UI with Run Notifications turned on. It fires once, for that specific run, and is then closed off. Use this when an external process kicks off a Bragi Job and needs to know when that particular run has finished.
Persistent - registered against a Job and fires every time that Job runs. Use this for standing integrations, for example notifying a downstream system whenever the overnight load completes, or raising a ticket whenever a Job fails.
Persistent callbacks are managed on the Scheduler > Job Callbacks page or through the API. One shot callbacks are created by the API or by Run Now and cannot be edited, but they are listed on the same page so you can see what is pending and what has recently fired, and each can be deleted.
Managing persistent callbacks
Field | Description |
|---|---|
Name | The name of the Job Callback. |
Description | A summary explaining the purpose of the callback, useful for future reference and edits. |
Callback URL | The |
Environment | The Environment this callback is associated with. |
Job Config | The specific Job this callback watches. |
Fires On |
|
Max Attempts | Total delivery attempts, not retries on top of the first. The default of |
Enabled | Turn the callback off without deleting it. |
Caller Reference | Optional value echoed back in the payload so the receiver can tie the callback to something on its side. |
The Test button posts a sample payload to the URL straight away and reports the response, so you can check the endpoint is reachable before relying on it. It tries once regardless of Max Attempts, and does not run the job or record anything against the callback.
The detail view also shows when the callback last fired, how many times it has fired, the HTTP status code that came back and the last error, if any.
The payload
Bragi sends a POST with a Content-Type of application/json and a Content-Length header:
Field | Description |
|---|---|
| The Job that ran. |
| The Environment it ran in. |
| The run itself. Matches the instance shown in the Scheduler and History pages. |
| The run number for that Job. |
| The final status of the run, for example |
|
|
|
|
| UTC timestamps. |
| How long the run took. |
| Names of the tasks that failed. Empty on a successful run. |
| Whatever was supplied when the callback was registered. |
| The run mode the run executed with. |
| Reserved for a future release. Always null. |
Bragi does not send any authentication with the request. If your endpoint needs to be protected, use an unguessable URL or restrict it at the network level.
A callback that fails, whether the endpoint is unreachable, returns an error status or does not answer within 30 seconds, is retried until Max Attempts is used up. The failure is then recorded against the callback and in the Bragi log. It never affects the Job that triggered it.
A callback whose Job ran but whose outcome did not match Fires On is not sent, and nothing is recorded against a persistent callback. A one shot in that position is closed off instead, with Not sent - job outcome did not match trigger as its last error, because its run has been and gone.
Triggering a Job with a one shot callback
The Job must have Allow API Trigger enabled. Add the callback parameters to the existing RunJobNow call. See Bragi API for how to authenticate and call it:
Parameter | Description |
|---|---|
| Optional. Omit it and |
| Optional, one of |
| Optional. Echoed back as |
The callback is bound to the run this call deals with, so you do not need to correlate it yourself. It only ever fires for a run that started after your call, so anything you wrote to the Job's sources before calling is included in the run it reports on. If the run never reaches the scheduler, the one shot is closed off automatically after 24 hours.
When the Job is already running
Normally the call queues a new run and binds the callback to it. A Job only ever has one run going at a time, so if a run is already queued or in progress, what happens depends on whether it has started and whether it will produce what you asked for:
Situation | Result |
|---|---|
A run is queued but has not started, in the same run mode, and you supplied a | The callback is attached to the queued run and the call succeeds, reporting |
A run has already started, or is in a different run mode, and you supplied a | The call succeeds, reporting |
No | Fails with |
Callers that arrive while a run is in progress all share the next run rather than each queuing their own: if ten requests arrive during run 1, run 2 is queued once run 1 finishes and all ten callbacks fire when run 2 finishes, each with its own callerReference. Callers asking for a different run mode wait for a run of their own after that, oldest first.
This means a successful response does not always mean a new run was queued straight away, and the wait can be up to two runs long: the one in progress, then yours.
While a callback is waiting for its run to be queued, the One Shot Callbacks table shows its state as Next run.
Run Now in the Bragi UI works differently. A Run Notification just reports on whatever run the click produced, so when the Job is already running in the same run mode it is attached to the run in progress and no further run is queued.
Waiting for a Job: a worked example walks through building the receiving side, including the wait, the timeout and the pitfalls.
Retention
A one shot is registered for every triggered run, so an environment driven by an API integration creates them continuously. They are housekeeping, not configuration, and are removed automatically once their run has been dealt with and they were registered more than Job Callback Retention Days ago (Admin > Settings, default 7). The age is measured from when the callback was registered, not from when it fired. Setting it to 0 keeps them all, which will grow without limit.
Persistent callbacks are configuration and are never removed automatically.
The One Shot Callbacks table on the Job Callbacks page shows the most recent 100 for the current environment, so it stays readable however many have been registered.
Managing persistent callbacks through the API
Endpoint | Description |
|---|---|
| Registers a persistent callback. Takes |
| Removes a callback, either by |
| Lists the persistent callbacks for an |
RegisterCallback and UnregisterCallback require the Job to have Allow API Trigger enabled, the same as RunJobNow. Removing by callbackId skips that check, since it does not name a Job. GET /Jobs/Callbacks only reads, so it does not check it either.
A callback registered through the API always gets the default Max Attempts of 3. Change it on the Scheduler > Job Callbacks page if you need something else.
Job Callbacks and Job Monitors
Job Monitors email a person a readable summary of recent runs. Job Callbacks post a machine readable payload to another system. They are independent, so a Job can have both.