Tips

Designing Idempotent Payment Webhooks That Never Double-Charge

Signature verification, idempotency keys, and safe retries for payment gateway webhooks in a NestJS checkout flow.

Designing Idempotent Payment Webhooks That Never Double-Charge

Every payment gateway — Stripe, VNPay, PayPal — retries a webhook that did not get a fast 2xx response, and network blips mean "did not get a response" and "the handler already succeeded but the ack was lost" are indistinguishable from the gateway's side. An endpoint that is not built for duplicates will eventually process the same successful payment twice.

Verify the signature before touching anything else

A webhook URL is public by necessity, which makes it a target for forged requests claiming a payment succeeded. Every gateway signs its payload with a secret only your server knows; verifying that signature — using the raw request body, not the parsed JSON, because re-serializing can change byte-for-byte formatting and break the signature — must happen before any business logic runs.

@Post('webhooks/vnpay')
async handleVnpayWebhook(
  @Body() body: VnpayWebhookDto,
  @Headers('x-vnpay-signature') signature: string,
) {
  const isValid = this.vnpayService.verifySignature(body, signature);
  if (!isValid) {
    throw new UnauthorizedException('Invalid webhook signature');
  }

  await this.paymentsService.handleWebhookEvent(body);
  return { received: true }; // fast 2xx, always
}

An idempotency key turns retries into no-ops

The gateway's own transaction id is a natural idempotency key: before applying any effect, check whether that id has already been recorded as processed, inside the same transaction that records it now. If it has, return success immediately without touching the order again — the gateway sees the same 2xx it would have gotten the first time, and nothing downstream fires twice.

  • Store the gateway transaction id with a unique constraint, not just an application-level check — a race between two near-simultaneous retries needs the database, not if statements, to be the tie-breaker.
  • Update the order status with a guarded transition (status = paid WHERE status = pending), so a duplicate webhook that slips past the idempotency check still cannot move a refunded order back to paid.
  • Always return the fast acknowledgment even on a webhook you choose not to act on (an unrecognized event type) — the gateway does not need your business logic to succeed, only your endpoint to respond.
  • Log every webhook payload before processing it, signature failures included — that history is the only way to reconstruct what happened when a customer disputes a charge weeks later.
  • Never trust the webhook payload for the order amount; look up the order server-side by its id and compare against what the gateway reports, rejecting a mismatch instead of trusting the number in the request.

Process asynchronously, acknowledge synchronously

If handling a webhook involves sending an email, updating a search index, and notifying a fulfillment service, doing all of that inline before responding risks a gateway timeout that triggers an unnecessary retry. Acknowledge as soon as the payment state itself is durably recorded, and push the side effects onto a queue that can retry independently of the webhook delivery.

async handleWebhookEvent(event: VnpayWebhookDto): Promise<void> {
  await this.dataSource.transaction(async (manager) => {
    const alreadyProcessed = await manager.exists(PaymentEntity, {
      where: { gatewayTransactionId: event.transactionId },
    });
    if (alreadyProcessed) return;

    await manager.save(PaymentEntity, {
      gatewayTransactionId: event.transactionId,
      orderId: event.orderId,
      status: 'succeeded',
    });

    await manager.update(
      OrderEntity,
      { id: event.orderId, status: 'pending' },
      { status: 'paid' },
    );
  });

  await this.deliveryQueue.add('deliver-order', { orderId: event.orderId });
}

A webhook handler that is not safe to call twice is not safe to call once — you cannot know, from inside the handler, whether this is the first delivery or a retry of a call that already succeeded.

Testing webhooks locally without a live gateway account

A webhook handler cannot be trusted until it has been tested against a request shaped exactly like the real gateway would send — headers, signature, and all — not just a hand-built JSON body sent through Postman with the signature check commented out. Most gateways ship a CLI (Stripe CLI, VNPay's sandbox tools) that forwards real signed test events to a local server through a tunnel.

# Forward Stripe's real test-mode webhook events to a local server,
# signed exactly the way production events are.
stripe listen --forward-to localhost:3000/webhooks/stripe

# Trigger a specific event to test the handler's reaction to it.
stripe trigger payment_intent.succeeded

A local integration test suite should also replay a captured real payload with a correct signature computed against the same secret the test environment uses, so the signature verification path itself is under test, not just skipped in favor of a shortcut that leaves it silently untested until production traffic finds the bug.

  • Keep a small fixture file of real (test-mode) captured payloads per event type — gateways occasionally change field shapes, and a stale hand-written fixture drifts from reality.
  • Test the duplicate-delivery path explicitly: send the same event twice and assert the second call is a no-op, not just that the first call succeeds.
  • Test an invalid signature explicitly and assert a 401 — a webhook endpoint with no negative test for this is untested in the one place it matters most.

Refunds and chargebacks are webhooks too

A refund initiated from the gateway's own dashboard — not through your API — still needs to reach your system through the same webhook channel, and it needs the same idempotency treatment: a refund.succeeded event keyed on the gateway's refund id, checked against what has already been recorded before any order status changes. A chargeback (the customer's bank forcibly reversing the charge) is the same shape again, just initiated by a different party.

async handleRefundEvent(event: RefundWebhookDto): Promise<void> {
  await this.dataSource.transaction(async (manager) => {
    const alreadyProcessed = await manager.exists(RefundEntity, {
      where: { gatewayRefundId: event.refundId },
    });
    if (alreadyProcessed) return;

    await manager.save(RefundEntity, {
      gatewayRefundId: event.refundId,
      orderId: event.orderId,
      amount: event.amount,
    });

    // A partial refund does not necessarily flip the whole order to refunded.
    const order = await manager.findOneOrFail(OrderEntity, { where: { id: event.orderId } });
    const isFullRefund = BigInt(event.amount) >= BigInt(order.totalAmount);
    await manager.update(OrderEntity, { id: order.id }, {
      status: isFullRefund ? 'refunded' : 'partially_refunded',
    });
  });
}

The full-versus-partial distinction is easy to skip in a first implementation and expensive to retrofit later, because "was this order ever fully refunded" is exactly the question support and finance ask most often when a customer disputes a charge.

What the customer-facing order status should say during processing

A payment that succeeded at the gateway but has not yet been reflected in the order (because the webhook has not arrived, or is queued behind other work) is a real, if brief, window most checkout flows have to design for explicitly. Showing a generic "processing" status rather than optimistically flipping to "paid" on the client before the server confirms it avoids a worse failure mode: a customer who sees "paid" and starts using a product that a failed or reversed webhook later un-pays.

A short client-side poll (or a WebSocket push once the webhook actually lands) that upgrades "processing" to "paid" within a few seconds keeps the perceived wait short without ever showing a status the server has not actually confirmed.

A webhook endpoint is also a reasonable place to record latency metrics specifically, separate from the rest of the API — a payment gateway that starts taking noticeably longer to deliver its webhooks is often an early signal of an incident on the gateway's own side, visible in your own metrics before the gateway's status page catches up. Treating webhook delivery latency as a signal worth alerting on, not just webhook failures, catches degradation earlier than waiting for an outright error.

Conclusion

Signature verification keeps forged events out; an idempotency key on the gateway's transaction id keeps duplicates from double-processing; a guarded status transition keeps out-of-order delivery from corrupting state; and pushing side effects to a queue keeps the acknowledgment fast enough that the gateway never needs to retry a request that actually succeeded.

Member discussion

Share your thoughts with the ToshStack community.

Join the discussion

Become a member of ToshStack to start commenting.

Already a member? Sign in