Tips

Background Jobs in NestJS With BullMQ: A Practical Setup

Queues, retries with backoff, and idempotent processors for BullMQ in NestJS, so a slow job never blocks an HTTP response again.

Background Jobs in NestJS With BullMQ: A Practical Setup

Sending an email, generating a PDF invoice, or calling a slow third-party API inside an HTTP request handler ties up that request for as long as the slow work takes, and a timeout or a crash mid-request leaves the work half-done with no record it needs to be retried. A queue moves that work off the request path entirely.

Producer and processor: two different concerns

The HTTP handler only needs to enqueue a job and return — it does not wait for the job to run. A separate processor, potentially in a separate process entirely, picks jobs off the queue and does the actual work, which means the request returns in milliseconds regardless of how long email delivery actually takes.

// Producer — inside the checkout flow, returns immediately
@Injectable()
export class OrdersService {
  constructor(@InjectQueue('delivery') private deliveryQueue: Queue) {}

  async completeOrder(orderId: string) {
    await this.orderRepository.update({ id: orderId }, { status: 'paid' });
    await this.deliveryQueue.add(
      'deliver-order',
      { orderId },
      { attempts: 5, backoff: { type: 'exponential', delay: 2000 } },
    );
  }
}

// Processor — runs independently, retried automatically on throw
@Processor('delivery')
export class DeliveryProcessor extends WorkerHost {
  async process(job: Job<{ orderId: string }>): Promise<void> {
    await this.deliveryService.deliverOrder(job.data.orderId);
  }
}

Every processor must be safe to run twice

A crash between finishing the real work and BullMQ marking the job complete means the job gets retried and the processor runs again — this is not an edge case, it is the normal contract of an at-least-once queue. deliverOrder has to check whether delivery already happened before sending anything a second time, the same idempotency discipline a payment webhook needs.

  • Set attempts and backoff explicitly on every job — the BullMQ default is a single attempt with no retry, which silently drops a failed job.
  • Use exponential backoff, not a fixed delay, so a downstream outage does not turn every failed job into an immediate retry storm against a service that is already down.
  • Configure removeOnComplete and removeOnFail with sane limits — an unbounded job history grows Redis memory usage forever.
  • Give each job a deterministic jobId derived from its natural key (order-${orderId}-delivery) when duplicate enqueues are possible, so BullMQ deduplicates instead of running the same delivery twice.
  • A job that exhausts all retries needs to land somewhere visible — a dead-letter queue, an alert, a status column on the originating row — not silently disappear.

Separate queues by failure domain, not just by feature

A single default queue mixing password-reset emails with heavyweight report generation means a burst of slow reports delays every password-reset email behind them in line. Splitting into email, reports, and delivery queues, each with its own concurrency setting, keeps a slow class of job from starving a fast, latency-sensitive one.

BullModule.registerQueue(
  { name: 'email', defaultJobOptions: { attempts: 3 } },
  { name: 'reports', defaultJobOptions: { attempts: 1 } },
);

@Processor('email', { concurrency: 10 })
export class EmailProcessor extends WorkerHost { /* ... */ }

@Processor('reports', { concurrency: 2 })
export class ReportsProcessor extends WorkerHost { /* ... */ }

A queue is not a substitute for a database record of what happened. Persist the outcome of a job (delivered, failed, retried N times) on the entity it relates to, so support and admin tooling can answer "what happened to this order's delivery" without reading Redis directly.

Scheduling recurring work with repeatable jobs

A nightly report, a subscription renewal check, or a cleanup of expired sessions needs to run on a schedule rather than in response to a request — BullMQ's repeatable job option turns a cron expression into a self-scheduling job that re-enqueues itself, instead of relying on a separate cron daemon that has no visibility into the same retry and dead-letter machinery the rest of the queue uses.

await this.reportsQueue.add(
  'nightly-sales-report',
  {},
  {
    repeat: { pattern: '0 2 * * *' }, // 02:00 server time, every day
    jobId: 'nightly-sales-report',    // stable id prevents duplicate schedules
  },
);

@Processor('reports')
export class ReportsProcessor extends WorkerHost {
  async process(job: Job): Promise<void> {
    if (job.name === 'nightly-sales-report') {
      await this.reportsService.generateNightlySalesReport();
    }
  }
}

A stable jobId on the repeat registration is what keeps a redeploy from silently creating a second, overlapping schedule for the same report — BullMQ treats a repeat registration with the same id as an update to the existing schedule rather than a brand new one.

Monitoring queue health before it becomes an incident

A queue that is silently backing up — jobs arriving faster than the processor can handle them — looks fine from the outside right up until the backlog is large enough that a "background" email arrives hours late. Tracking queue depth and processing latency as first-class metrics turns that into an alert hours before a customer notices, not a support ticket after the fact.

async getQueueHealth(queueName: string) {
  const queue = this.queues.get(queueName);
  const counts = await queue.getJobCounts('active', 'waiting', 'failed', 'delayed');

  return {
    waiting: counts.waiting,
    active: counts.active,
    failed: counts.failed,
    // A growing "waiting" count over time, not its absolute value, is the
    // signal that processors are falling behind arrival rate.
  };
}

A rising failed count deserves its own alert independent of queue depth — a processor that has started throwing on every job (a bad deploy, an expired credential to a third-party service) will drain waiting down to zero just as fast as a healthy one, while quietly failing every single job, which queue-depth-only monitoring completely misses.

Deciding what genuinely belongs in a background job

Not every slow operation belongs in a queue — a user waiting on a password reset email genuinely wants confirmation the request was accepted within a second or two, and moving that specific email to a background queue with no visible feedback can make a fast operation feel broken if the queue is ever backed up. The right question is not "is this slow" but "does the user need to wait for this specific step to know their request succeeded."

A reasonable middle ground for latency-sensitive background work is a synchronous acknowledgment ("email queued") paired with an asynchronous delivery guarantee — the user gets fast feedback that something happened, and the queue still owns the actual reliability of getting the email out.

Worth deciding explicitly up front: what happens to a job whose payload references a row that has since been deleted — an order cancelled between being enqueued for delivery and the processor actually running. A processor that assumes the referenced data will always exist crashes with a confusing stack trace on every retry; one that checks for that case explicitly and marks the job as skipped, with a clear log line, turns an ordinary race condition into an unremarkable, self-explaining outcome instead of a paging alert at 3am.

Job payloads deserve the same size discipline as any other message passed between services: a large object embedded directly in the job data bloats Redis memory and slows every read of the queue, when a smaller reference — an entity id the processor looks up itself — usually serves just as well and keeps the payload the same tiny size regardless of how large the underlying entity grows over time.

Conclusion

BullMQ earns its place the moment any request handler does work slow enough, or unreliable enough, to risk the request itself. Enqueue and return fast, write every processor as if it will run twice, back off exponentially on failure, and split queues along failure domains so one slow job class cannot starve a fast one.

Member discussion

Share your thoughts with the ToshStack community.

Join the discussion

Become a member of ToshStack to start commenting.

Already a member? Sign in