Skip to content

The payment row is the outbox

Moving payment processing out of web requests and onto GoodJob and Postgres, with the payment record as the durable intent, jobs that can be replayed by payment ID, and the numbers from a load test of the result.

13 min read

A payment platform built on Rails sent money to a bank API from inside web requests. Approving a batch of transfers held a web thread for the whole batch, and the size of a batch was limited by the ingress timeout rather than by anything in the design. The processing was moved to background jobs in the same Postgres database, with GoodJob as the job runner.

This post describes the resulting design: the approved payment row as the record of intent, the job as a dispatch that can always be recreated from it, the four checks that make a replay safe, the ledger that lets the bank call run outside any transaction, and how every payment reaches a final state. It closes with production and load test numbers for a single worker.

Payments inside the web request

Approving a batch was one HTTP request. The controller wrote the approval records, then looped over the payments and processed each one inline:

  1. Check the balance and the customer’s consent.
  2. Lock the sender’s balance row and the external account’s balance row with SELECT ... FOR UPDATE.
  3. Write the ledger transfer inside that lock.
  4. Call the bank API, still inside the lock and the open transaction.
  5. Store the bank’s transaction ID and schedule a status check.

Each web process ran five threads, with two processes per replica: ten requests at a time. The HTTP client had no configured timeout, so a slow bank call could wait for the Ruby defaults of 60 seconds to open and 60 seconds to read. In front of all of it sat an ingress timeout of 240 seconds.

One request thread, one payment after another Ingress timeout, 240 s The balance row stays locked for the whole payment, bank call included 504 to the browser; the thread keeps sending Ledger write, balance row locked Bank API call
One request approves a batch. Every payment holds the balance row lock across its bank call. Once the batch outlasts the ingress timeout, the browser receives a 504 while the same thread keeps sending payments.

Three consequences followed from that shape.

The batch size limit was the ingress timeout. At half a second per bank call, 240 seconds holds fewer than 500 payments before any database time is counted. Past that point the gateway answered the browser with a 504 while the web server kept the thread running, so the user saw a failure for a batch that was still being sent.

A failure in the middle left the batch split. All approvals were written first and all transfers were sent afterwards, with no transaction around either loop. The first exception stopped the loop: payments before it had been sent, the failing one was marked failed, and the payments after it stayed approved without anything that would send them later.

One account’s payments could not run in parallel. The balance row lock was held for the length of the bank call, so two requests paying from the same account queued behind each other on that row, each holding a web thread and a database connection while it waited.

A separate batch upload flow had already been moved to a background job, but that job processed the whole batch in one execution, one payment after another. It avoided the timeout and kept every other property above.

The payment row as the outbox

The transactional outbox pattern writes a business change and a message about it in one database transaction, then lets a relay deliver the message. The guarantee is that the change and the intent to act on it are never separated: either both commit or neither does.

A payment system already stores its intent. An approved payment row, keyed by the payment’s UUID, is a durable statement that money should move. Everything needed to send it is on that row or reachable from it. The job that sends it does not need to carry any information of its own; it only needs to name the payment.

GoodJob keeps its jobs in a Postgres table in the application’s own database. Enqueueing a job is an INSERT into good_jobs on the same connection as the rest of the request, followed by a NOTIFY that Postgres delivers only when the transaction commits. With GoodJob’s default of enqueue_after_transaction_commit = false, perform_later inside an open transaction joins that transaction. There is no broker to keep consistent with the database and no dual write.

Web request approve batch Postgres, one database payments id 7f3c… approved sent to bank good_jobs ProcessPaymentJob(7f3c…) Both rows commit together, or neither does missing job: re-enqueue by payment ID Worker Bank NOTIFY
The payment row and the job row live in the same database. On commit, Postgres notifies the worker, which calls the bank and moves the payment from approved to sent. A payment without a job can always be given a new one by ID.

The job carries only the payment ID, and the approval step changed from sending a payment to enqueueing it:

on approve(payment):
    enqueue ProcessPayment(payment.id)          # INSERT into good_jobs, same transaction

ProcessPayment(payment_id):
    one running job per payment ID
    payment = load(payment_id) or stop
    process(payment)

Because the job is derived from the payment row, the job is not the thing that must never be lost. If a payment ends up without a job anyway, because an enqueue ran outside the approving transaction or a job was discarded, the payment is still approved and still carries its ID. Enqueueing it again is always correct, provided a second run of the job for the same payment cannot send money twice. The next section lists the checks that enforce that.

Four checks keyed on the payment ID

A replay can happen for ordinary reasons: a worker killed during a deploy, a retry after a timeout, a manual re-enqueue of approved payments that have no job. Four checks, all keyed on the payment ID, make each of these safe.

Job lock per payment ID Ledger line unique per payment Bank ID stored on payment Idempotency key = payment ID Bank First run Replay, ID stored Replay, ID lost created skipped, already sent same transfer Ledger check on a replay: the existing debit is reused, no second debit is written
The first run passes all four checks. A replay after the bank ID was stored stops at the third check. A replay after a crash that lost the ID reaches the bank, which returns the original transfer for the same idempotency key.

1. A job lock per payment. A concurrency key built from the payment ID allows one running job per payment. perform_limit: 1 makes a second job for the same payment wait and retry until the first has finished, and enqueue_limit: 1 drops a duplicate enqueue while one is still queued. Two details are worth knowing. In GoodJob 3.30, total_limit is ignored when both enqueue_limit and perform_limit are set, so a configuration that sets all three is only enforcing the latter two. And the enqueue limit counts queued jobs, not running ones: a new enqueue while a job for the same payment is running is accepted, and then waits on the perform limit.

2. One debit per payment. The debit is written only if the payment has none yet, and a unique constraint in the database allows one debit per payment. A replay reuses the existing debit instead of writing a second one.

3. The bank’s transaction ID on the payment. Once the bank has accepted a transfer, its ID is stored on the payment, and the bank client returns without sending anything when the ID is already present.

4. The payment ID as the idempotency key. Every request to the bank carries Idempotency-Key: <payment id>, and retries of a request reuse it. If a crash happens after the bank accepted the transfer but before its ID was stored, the next attempt sends the same key and receives the original transfer back. Reusing the key on retry was adopted after the bank confirmed that repeating a request with the same key is safe for timeouts and authentication errors.

The checks cover replays of one payment. They do not cover a user who creates the same transfer again as a new payment, which gets a new ID and a new key. That case belongs to duplicate detection at creation time, not to the processing pipeline.

Ledger first, bank call second

The first version of the job kept the original processor, which called the bank inside the ledger lock. Moving the work into a job removed the request timeout, but a payment still held the account’s balance row for the whole bank call. Payments from one account still ran one at a time, and each waiting job still held a worker thread and a database connection.

The way out depended on the ledger. Money is recorded in a double-entry ledger. Every movement is a transfer that writes two lines, one on each account, with amounts that sum to zero, and each account’s balance is maintained from its lines. A payment writes a debit transfer from the customer’s account to an external account that stands for money leaving the platform, and each line has a status: pending, settled or failed. Lines are appended and never edited or deleted. A line’s status is kept apart from the line itself, as its own append-only records, so a status change adds a record and leaves the line untouched. Money is corrected only by appending new lines.

Two properties of that ledger allow the debit to commit before the bank call. The committed debit transfer reserves the money: the customer’s balance drops at commit, so the next payment from the same account sees the reduced balance without waiting for this payment’s bank call. And a committed transfer can be undone later by appending the opposite transfer, so nothing needs a transaction to stay open in case the bank says no.

The processor was split into two phases:

phase 1, under the row locks, one short transaction:
    stop if a debit for this payment exists
    write the debit, mark the payment pending
    commit

phase 2, no transaction, no locks:
    call the bank with idempotency key = payment ID
    record the result

Phase one takes the row locks, writes the debit, marks the payment pending and commits. Phase two calls the bank with no transaction open and no rows locked, then records the result. Lock waits and deadlocks in phase one are raised so that the job retries them; any other failure there marks the payment failed before any money has left.

Bank call inside the ledger lock Payment 1Payment 2Payment 3 1.8 s Ledger committed first, bank call after Payment 1Payment 2Payment 3 0.65 s 0 0.5 s 1 s 1.5 s Row lock Bank call Waiting
Three payments from the same account. With the bank call inside the lock, each payment waits for the previous one's bank call. With the ledger committed first, the lock covers only the ledger write and the bank calls overlap. Durations are illustrative: a 500 ms bank call and a 50 ms ledger write.

Reversal instead of rollback

After the split, the outcome of the bank call is written as new ledger state rather than decided by a commit or a rollback. If the bank accepts the transfer, a settled status is recorded for the two debit lines and no new lines are written. If the bank rejects it, or its retries run out, a reversal transfer is appended in the opposite direction, from the external account back to the customer, and a failed status is recorded for all four lines. The customer’s balance returns to where it was, and the history shows the attempt and its reversal.

bank call in progress, no locks held Bank accepts code account amount status debit customer −100.00 pending settled debit external +100.00 pending settled Customer balance 1,000.00 900.00 Bank rejects code account amount status debit customer −100.00 pending failed debit external +100.00 pending failed reversal external −100.00 failed reversal customer +100.00 failed Customer balance 1,000.00 900.00 1,000.00
One payment of 100.00 under both outcomes. Phase one commits the debit transfer and the customer's balance drops at once. An accepted transfer only records a new status for the two lines. A rejected one appends the opposite transfer, and the balance returns.

A rollback can only undo work inside a transaction that is still open, which is why the first design kept the transaction open for the length of the bank call. A compensating transfer undoes committed work at any later time and from any process: the job that sent the payment, a retry an hour later, or the recovery job. A reversal that runs twice finds the first one already written and writes nothing. That is what keeps the row lock to the tens of milliseconds of the ledger write instead of the length of an HTTP call.

Every payment reaches a final state

Each bank response is classified by its response code into one of two outcomes.

ResponseMeaningHandling
Timeout, connection failure, authentication error, temporary error at the bankThe transfer may or may not have happenedRetry with growing delays, same idempotency key, ledger untouched
Rejection such as an invalid or closed accountThe transfer did not and will not happenReverse at once
Retries exhaustedThe outcome is still unknownReverse, and rely on the nightly reconciliation to catch a transfer the bank did execute

Database lock timeouts and deadlocks are retried on their own short schedule, since they say nothing about the bank. The reversal path is a single recovery job: it writes whatever part of the ledger pair is missing, appends the reversal, marks the payment failed, and writes nothing on a second run.

The final status reaches the ledger through two independent streams. The primary stream is the bank’s webhook. The bank acknowledges each instant transfer at once and reports settled or rejected later, and the webhook enqueues a status job that settles or reverses the payment. A check scheduled thirty minutes after sending picks up any payment whose webhook never arrived. The second stream is a nightly reconciliation that compares every ledger balance with the bank’s balance for the same account and reports any difference.

Together, the two streams keep the ledger in line with the bank. During major outages at the payment processor, the reconciliation report showed which accounts differed, and those were corrected by hand.

Capacity

In production, a single worker replica has processed up to 100,000 payments per day. A load test measured how far one replica can go.

Load test setup
Worker1 process, 1 vCPU, 10 threads
DatabasePostgres 16, 4 CPUs, commit latency about 6.7 ms, matching a zone-redundant managed server
BankSimulated, 500 ms median response, asynchronous acknowledgement
Data500 source accounts
Result
Throughput10.3 payments per second
Job time, p50 / p95930 ms / 1,200 ms
Worker CPU, p950.9 of 1 vCPU
Postgres CPU0.21 of 4 cores
Lock waits, errorsnone

One process can do roughly as many payments per second as its threads divided by the time a job holds a thread. The load test recorded the median job time, not the mean, so this is an approximation rather than a bound:

payments per second  ≈  threads / median seconds per job  =  10 / 0.93  ≈  10.8

A thread sweep shows where the thread count stops being the limit and the CPU takes over. The 5- and 20-thread runs used the synchronous bank path.

ThreadsPayments per secondJob p50Worker CPULimited by
56.4751 ms0.53Threads
1010.3930 ms0.78vCPU
209.81,934 ms0.93vCPU

Payments from a single source account ran at the same rate as payments spread over 500 accounts: the balance row is locked only for the ledger write.

Daily volume from one workerPayments per secondPer day
Measured rate, even over 24 hours10.3890,000
Measured rate, 8-hour window10.3300,000
Estimate with production round trips and webhook status jobs, 24 hoursabout 8about 690,000
Payments per second, one worker process per vCPU 0 10 20 30 40 1 process, 5 threads limited by threads 6.4/s 555k per day 1 process, 10 threads limited by the vCPU 10.3/s 888k per day 1 process, 20 threads vCPU saturated 9.8/s 843k per day 2 processes, 10 threads projected, 2 vCPU 20.6/s 1.8M per day 4 processes, 10 threads projected, 4 vCPU 41.1/s 3.6M per day 500k per day Measured Projected, linear per process
One worker process per vCPU. The 10-thread run used the asynchronous bank path that production runs; the 5- and 20-thread runs used the synchronous path. The hatched bars multiply the one-process result by two and four and assume the bank's rate limit allows it. The dashed line is 500,000 payments per day at an even rate.

Past ten payments per second, more processes help and more threads do not. Each process adds its threads plus two to the database’s connection count. The other levers are persistent HTTP connections to the bank and a bank rate limit that permits the result.

Why this design

The design was chosen under two constraints. It had to ship quickly, inside the existing Rails application and its Postgres database, with no new infrastructure to run. And it had to be exact, because every job moves money. The first version moved payments off the web request without adding any infrastructure, and the later changes built on the same foundation.

It is not the design a larger system would end with. Throughput is capped per worker process, the job table shares the database with everything else, and retries, timers and waiting on webhooks live in hand-written job code. A payment whose retries run out while its outcome is unknown is reversed, and a transfer the bank did execute is found by reconciliation, not prevented; holding such payments in an unknown state until the bank confirms the outcome would be stricter.

AlternativeWhat it addsWhat it costs
A log broker, fed from an outbox table by change data captureDispatch separated from the database, replay, fan-out to other consumersA broker and a relay to run; consumers still need the same idempotency keys
Event sourcingPayment state as a replayable log; the append-only ledger is already halfway thereA rewrite of how payment state is stored and queried
A durable workflow engineRetries, timers and webhook waits as workflow code instead of job stateA new runtime and a new programming model
Locks per account in an external lock serviceAccount locking outside the databaseRejected: an external lock can expire or be lost in a failover while its holder still acts on it, and making it safe needs fencing tokens checked by the database, which puts the database back in the path

Checklist

  • Treat the domain row as the outbox. Store the intent on the row the business already owns, key the job on that row’s ID, and make re-enqueueing by ID always correct.
  • Enqueue inside the transaction that changes the state, and check the adapter’s enqueue_after_transaction_commit setting before relying on it.
  • Key every layer on the same ID: job concurrency, the ledger debit, the stored provider ID and the idempotency key.
  • Commit the ledger before calling an external API, hold row locks for the database work only, and undo with reversal transfers instead of rollbacks.
  • Size the database pool from the thread count plus GoodJob’s utility connections plus any threads started inside jobs.
  • Set explicit timeouts on every HTTP client a job calls.
  • Classify every provider response as retry or final, and reconcile balances with the provider on a schedule.
  • Connect GoodJob workers to Postgres directly. Its advisory locks are held per session, so it does not work through a connection pooler in transaction mode.

Further reading