Job Engine — Administrator Guide
Manage scheduled jobs, monitor real-time execution, and administer the async message queue across your GoodBooks GB5 installation.
Quartz.NET 3.13 Cluster
Real-time SignalR monitoring
Multi-tenant
DLQ management
Overview
The Job Engine replaces the legacy GB4 Windows Service scheduler and MSMQ queue reader. It provides:
- Scheduled jobs — cron, interval, and one-time execution of any GB5 API endpoint or Dapr topic
- Async message queue — persistent SQL-backed queue (TJOBQUEUE) with retry and dead-letter management
- Execution history — full audit trail of every job run, with duration, errors, and correlation IDs
- Real-time monitoring — live SignalR dashboard with KPI cards that update as jobs run
Tenant Level
Manages your own tenant's jobs and queues
Job Definitions · Schedules · Live Monitor · Execution History · DLQ Manager
Sys Admin (AdminRights ≥ 1)
Cross-tenant view across one installation
Tenant Health · Cross-Tenant History
Super Admin (AdminRights = 2)
Cluster-wide visibility across all nodes
Cluster Topology · Server Comparison
Job Definitions (MJOBDEFINE)
Create and manage job definitions. Each job links a schedule (TSCHEDULER) to a target API endpoint (MWEBSERVICE).
📋
Select an existing job using the picklist on the left. The form loads the job's full details including its cron expression preview.
📅
CRON expression builder — embedded cron picker shows a human-readable description and the next 5 scheduled times as you type. Example: 0 0 8 1/1 * ? = "Every day at 08:00".
▶
Trigger Now — fires the job immediately without waiting for the next scheduled time. A dialog confirms the Execution ID assigned to the run.
⏸
Pause / Resume — toggles the job's active state. Paused jobs are removed from the Quartz scheduler in real time; they do not run until resumed. The status badge updates instantly.
🏷
Status badge — shows ACTIVE (green), PAUSED (amber), or IDLE (grey). Updated automatically when you load a job or after Pause/Resume actions.
Job fields reference
| Field | Description |
| Job Name | Descriptive name. Required. Used in execution history and SignalR events. |
| Scheduler | Links to a TSCHEDULER row (cron, interval, or one-time). Create schedules in the Schedule Manager screen first. |
| Web Service | The MWEBSERVICE entry that defines the HTTP endpoint to call. Supports GET and POST methods. |
| URI Parameter | Optional value substituted into the URL template's {Param} placeholder. |
| Max Retries | How many times to retry on failure before marking the execution as FAILED. Default: 3. |
| Timeout (seconds) | Maximum execution time before the job is killed. Default: 300s (5 minutes). |
| Allow Concurrent | If unchecked, Quartz will skip the next fire if the previous execution is still running. |
| Job Category | REPORT | DATASYNC | DBMAINTENANCE | CUSTOM. Used for filtering and dashboard grouping. |
| Dapr Topic | If set, the job publishes to this Dapr pub/sub topic instead of calling the web service URL. |
Schedule Manager (TSCHEDULER)
Create reusable schedules that can be linked to one or more job definitions. Split-pane: list on left, editor on right.
🗓
Three schedule types — select CRON, INTERVAL, or ONE_TIME using the radio buttons in the editor.
👁
CRON preview panel — as soon as you edit the cron expression, the panel below shows the human-readable description and a list of the next 5 run times.
🔘
Active toggle — disable a schedule without deleting it. Jobs linked to a disabled schedule will be paused automatically.
Schedule type details
| Type | Fields required | When to use |
| CRON | Cron Expression (6-part Quartz format) | Complex recurring schedules — daily at 8am, every weekday, first of month, etc. |
| INTERVAL | Interval Seconds | Simple repeating jobs — every 5 minutes, every hour |
| ONE_TIME | Scheduled Date/Time | A single future execution — end-of-quarter batch, migration run |
Common CRON expressions
0 0 8 1/1 * ?Every day at 08:00
0 0/30 * * * ?Every 30 minutes
0 0 6 ? * MON-FRIWeekdays at 06:00
0 0 0 1 1/1 ? *1st of each month midnight
0 0 18 ? * FRIEvery Friday at 18:00
0 0/5 * * * ?Every 5 minutes
Quartz cron format
GB5 uses Quartz.NET cron syntax: Seconds Minutes Hours DayOfMonth Month DayOfWeek [Year]. Use ? (not *) for DayOfMonth or DayOfWeek when specifying the other. All times are UTC.
Live Job Monitor
Real-time dashboard connected to SignalR. No refresh needed — jobs appear and update automatically as they execute.
🟢
Connection bar — green "Live" when connected. Amber "Connecting…" during reconnect. Reconnects automatically; no action needed.
📊
4 KPI cards — Running Now / Waiting / DLQ Count / Failed (today). Updated in real time as jobs start and complete.
🃏
Job status cards — one card per in-flight or recently finished execution. Blue pulsing dot = running; green = completed; red = failed. Click the kebab menu (⋮) to Trigger Again or View History.
🔔
Failure toast notifications — a warning toast appears for every job failure, including the error snippet. DLQ promotions are indicated with a special label.
DLQ count going up?
If the DLQ card shows a non-zero count, navigate to DLQ Manager to inspect and requeue affected items. Jobs in DLQ are not retried automatically.
Execution History
Searchable, paged log of every job execution for your tenant. New completed executions prepend to the top without a page refresh.
🔍
Filter bar — filter by Status, date range (From / To). Apply resets to page 1.
🖱
Row click → detail view — opens a lightbox with the full error body (for failures), result JSON, duration, host instance, and correlation ID.
Execution status codes
QUEUED→
RUNNING→
SUCCESS
RUNNING→
FAILED→ retries →
FAILED (no more retries)
| Status | Meaning |
| QUEUED | Scheduled but not yet started |
| RUNNING | Currently executing |
| SUCCESS | Completed normally |
| FAILED | Failed; will be retried according to backoff schedule |
| TIMEOUT | Execution exceeded the configured timeout seconds |
| CANCELLED | Manually cancelled before execution completed |
Understanding Trigger Types
| Trigger Type | Cause |
| SCHEDULED | Fired by Quartz at its configured time |
| MANUAL | "Trigger Now" button clicked in Job Define screen |
| API | Called via the TriggerJob REST endpoint programmatically |
| QUEUE | Initiated by an async queue message handler |
DLQ Manager (Dead Letter Queue)
Items enter the DLQ when they have exhausted all retry attempts (default: 3 retries with exponential backoff). The DLQ count KPI updates in real time via SignalR.
🔄
Requeue individual item — resets status to PENDING and retry count to 0. Item re-enters the normal processing queue immediately.
🔄
Requeue All — bulk requeue all DLQ items. Requires confirmation. Items are processed sequentially to avoid server overload.
🗑
Purge All — permanently deletes all DLQ items (hard delete). A confirmation dialog appears before this action. Cannot be undone.
ℹ
View Error — opens a dialog with the full error message and stack trace for each item, helping diagnose root causes.
Retry backoff schedule
| Attempt # | Wait before retry |
| 1st failure | 30 seconds |
| 2nd failure | 2 minutes |
| 3rd failure | 8 minutes |
| 4th failure | 30 minutes |
| 5th failure | 2 hours |
| Exhausted (> Max Retries) | → DLQ — no further retries |
Best practice
Before requeuing, always click "View Error" to understand the root cause. Requeuing a systemic failure (e.g. external service down) will just re-fail and return to DLQ.
Sys Admin Screens
Access requirement
These screens are visible only to users with AdminRights ≥ 1 in their session. The backend also enforces this check — unauthorized API calls are rejected.
Cross-tenant overview across all tenants on this installation. Auto-refreshes every 60 seconds.
📊
4 KPI cards — Total Tenants / Tenants With Failures / Total DLQ Items / Currently Running. Computed from the current refresh snapshot.
🔴
Row health colour coding — green = healthy, amber = DLQ items present, red = failures in last 24h. Scan at a glance.
🖱
Click a row — navigates to the Cross-Tenant Execution History filtered to that tenant's DatabaseName.
Same as the tenant-level Execution History, with an additional Tenant column and an optional Tenant filter field. Does not update in real time (manual refresh button provided).
Super Admin Screens
Access requirement
These screens require AdminRights = 2 (Super Admin). They query Quartz cluster state (QRTZ_SCHEDULER_STATE) directly and are intended for DevOps / infrastructure engineers.
Quartz cluster node map. Auto-refreshes every 30 seconds. Each card represents one scheduler instance (typically one per pod/container).
🟢
Active — the node sent a heartbeat within 2× its checkin interval. This is the normal state.
🔴
Stale — the node has not checked in within the expected window. This indicates the pod has crashed or been terminated. Quartz will automatically recover any jobs that were assigned to a stale node within the next checkin cycle.
📊
Jobs in flight — current number of executing jobs on this node. A continuously high number may indicate a slow or stuck job.
🖱
Click a node card — pre-selects it in Server Comparison for performance analysis.
Compare throughput, failure rates, and average durations across 2–4 cluster nodes side by side. On-demand — results load when you click Compare.
☑
Node selector — check 2 to 4 nodes from the available list, then click Compare. Results appear below.
📈
Per-node KPIs — Jobs/hour, Failure Rate %, Average Duration. Failure Rate highlighted red if above 5%.
🐛
Top Failing Jobs per node — identifies which specific jobs are causing failures on each server, helping pinpoint node-specific issues (e.g., missing shared drive, network partition).
Troubleshooting
Job not running at its scheduled time
1
Check job status
In Job Define, load the job. Status badge should show ACTIVE. If PAUSED, click Resume.
2
Check cluster topology
As Super Admin, open Cluster Topology. All nodes should show Active. A stale node will not fire jobs until it recovers (or Quartz assigns them to another node after 2× checkin interval).
3
Use Trigger Now
Click "Trigger Now" to force an immediate execution. If it succeeds, the schedule may be using an incorrect cron expression (check in CRON preview panel).
4
Check Execution History
Filter by this job. If you see FAILED rows, the job is running but erroring. Click the row to read the full error detail.
DLQ items not clearing after requeue
If requeued items keep returning to DLQ:
- Open the item's View Error dialog and read the full error. Look for connection errors, permission denials, or validation failures.
- Check whether the target web service is reachable from the Job Engine pod.
- If the error is a data validation issue, the payload itself may be invalid — contact the application team to correct the source record before requeuing.
- Use Purge to discard items that cannot be recovered and are blocking queue statistics.
Tenant Health shows a tenant as red
Click the tenant row to drill into Cross-Tenant History. Filter by FAILED status. Look for patterns — same job failing repeatedly usually means a service-level issue for that tenant (e.g., their database is offline or a third-party integration is unreachable).
Cluster node marked Stale
A stale node means the pod stopped sending heartbeats. Check Kubernetes pod logs for the JobEngineSL pod with that instance name. Quartz automatically recovers jobs from stale nodes within 2 × CheckinInterval (typically 30 seconds). No manual intervention is needed unless the node remains stale after recovery.
Purge is permanent
Purging the DLQ hard-deletes rows from TJOBQUEUE. This cannot be undone. Only purge items you are certain should not be reprocessed — for example, test messages or records from a decommissioned integration.
Quick reference: where to look for what
| Symptom | First screen to check |
| Jobs not running | Job Define → status badge; Cluster Topology → stale nodes |
| Job running but producing wrong results | Execution History → row detail → result JSON / error details |
| Notifications not arriving | DLQ Manager → items with type EMAIL/SMS/WEBHOOK |
| One tenant has problems, others are fine | Tenant Health (Sys Admin) → click tenant → Cross-Tenant History |
| High failure rate on one server only | Server Comparison (Super Admin) → compare affected node vs healthy node |
| Real-time dashboard not updating | Live Monitor connection bar — if Disconnected, refresh the page to reconnect |