The payment row is the outbox
Moving payment processing out of web requests and onto GoodJob and Postgres, with the payment record as the durable intent, jobs that can be replayed by payment ID, and the numbers from a load test of the result.
13 min read
A payment platform built on Rails sent money to a bank API from inside web requests. Approving a batch of transfers held a web thread for the whole batch, and the size of a batch was limited by the ingress timeout rather than by anything in the design. The processing was moved to background jobs in the same Postgres database, with GoodJob as the job runner.
This post describes the resulting design: the approved payment row as the record of intent, the job as a dispatch that can always be recreated from it, the four checks that make a replay safe, the ledger that lets the bank call run outside any transaction, and how every payment reaches a final state. It closes with production and load test numbers for a single worker.
Payments inside the web request
Approving a batch was one HTTP request. The controller wrote the approval records, then looped over the payments and processed each one inline:
- Check the balance and the customer’s consent.
- Lock the sender’s balance row and the external account’s balance row with
SELECT ... FOR UPDATE. - Write the ledger transfer inside that lock.
- Call the bank API, still inside the lock and the open transaction.
- Store the bank’s transaction ID and schedule a status check.
Each web process ran five threads, with two processes per replica: ten requests at a time. The HTTP client had no configured timeout, so a slow bank call could wait for the Ruby defaults of 60 seconds to open and 60 seconds to read. In front of all of it sat an ingress timeout of 240 seconds.
Three consequences followed from that shape.
The batch size limit was the ingress timeout. At half a second per bank call, 240 seconds holds fewer than 500 payments before any database time is counted. Past that point the gateway answered the browser with a 504 while the web server kept the thread running, so the user saw a failure for a batch that was still being sent.
A failure in the middle left the batch split. All approvals were written first and all transfers were sent afterwards, with no transaction around either loop. The first exception stopped the loop: payments before it had been sent, the failing one was marked failed, and the payments after it stayed approved without anything that would send them later.
One account’s payments could not run in parallel. The balance row lock was held for the length of the bank call, so two requests paying from the same account queued behind each other on that row, each holding a web thread and a database connection while it waited.
A separate batch upload flow had already been moved to a background job, but that job processed the whole batch in one execution, one payment after another. It avoided the timeout and kept every other property above.
The payment row as the outbox
The transactional outbox pattern writes a business change and a message about it in one database transaction, then lets a relay deliver the message. The guarantee is that the change and the intent to act on it are never separated: either both commit or neither does.
A payment system already stores its intent. An approved payment row, keyed by the payment’s UUID, is a durable statement that money should move. Everything needed to send it is on that row or reachable from it. The job that sends it does not need to carry any information of its own; it only needs to name the payment.
GoodJob keeps its jobs in a Postgres table in the application’s own database. Enqueueing a job is an INSERT into good_jobs on the same connection as the rest of the request, followed by a NOTIFY that Postgres delivers only when the transaction commits. With GoodJob’s default of enqueue_after_transaction_commit = false, perform_later inside an open transaction joins that transaction. There is no broker to keep consistent with the database and no dual write.
The job carries only the payment ID, and the approval step changed from sending a payment to enqueueing it:
on approve(payment):
enqueue ProcessPayment(payment.id) # INSERT into good_jobs, same transaction
ProcessPayment(payment_id):
one running job per payment ID
payment = load(payment_id) or stop
process(payment)
Because the job is derived from the payment row, the job is not the thing that must never be lost. If a payment ends up without a job anyway, because an enqueue ran outside the approving transaction or a job was discarded, the payment is still approved and still carries its ID. Enqueueing it again is always correct, provided a second run of the job for the same payment cannot send money twice. The next section lists the checks that enforce that.
Four checks keyed on the payment ID
A replay can happen for ordinary reasons: a worker killed during a deploy, a retry after a timeout, a manual re-enqueue of approved payments that have no job. Four checks, all keyed on the payment ID, make each of these safe.
1. A job lock per payment. A concurrency key built from the payment ID allows one running job per payment. perform_limit: 1 makes a second job for the same payment wait and retry until the first has finished, and enqueue_limit: 1 drops a duplicate enqueue while one is still queued. Two details are worth knowing. In GoodJob 3.30, total_limit is ignored when both enqueue_limit and perform_limit are set, so a configuration that sets all three is only enforcing the latter two. And the enqueue limit counts queued jobs, not running ones: a new enqueue while a job for the same payment is running is accepted, and then waits on the perform limit.
2. One debit per payment. The debit is written only if the payment has none yet, and a unique constraint in the database allows one debit per payment. A replay reuses the existing debit instead of writing a second one.
3. The bank’s transaction ID on the payment. Once the bank has accepted a transfer, its ID is stored on the payment, and the bank client returns without sending anything when the ID is already present.
4. The payment ID as the idempotency key. Every request to the bank carries Idempotency-Key: <payment id>, and retries of a request reuse it. If a crash happens after the bank accepted the transfer but before its ID was stored, the next attempt sends the same key and receives the original transfer back. Reusing the key on retry was adopted after the bank confirmed that repeating a request with the same key is safe for timeouts and authentication errors.
The checks cover replays of one payment. They do not cover a user who creates the same transfer again as a new payment, which gets a new ID and a new key. That case belongs to duplicate detection at creation time, not to the processing pipeline.
Ledger first, bank call second
The first version of the job kept the original processor, which called the bank inside the ledger lock. Moving the work into a job removed the request timeout, but a payment still held the account’s balance row for the whole bank call. Payments from one account still ran one at a time, and each waiting job still held a worker thread and a database connection.
The way out depended on the ledger. Money is recorded in a double-entry ledger. Every movement is a transfer that writes two lines, one on each account, with amounts that sum to zero, and each account’s balance is maintained from its lines. A payment writes a debit transfer from the customer’s account to an external account that stands for money leaving the platform, and each line has a status: pending, settled or failed. Lines are appended and never edited or deleted. A line’s status is kept apart from the line itself, as its own append-only records, so a status change adds a record and leaves the line untouched. Money is corrected only by appending new lines.
Two properties of that ledger allow the debit to commit before the bank call. The committed debit transfer reserves the money: the customer’s balance drops at commit, so the next payment from the same account sees the reduced balance without waiting for this payment’s bank call. And a committed transfer can be undone later by appending the opposite transfer, so nothing needs a transaction to stay open in case the bank says no.
The processor was split into two phases:
phase 1, under the row locks, one short transaction:
stop if a debit for this payment exists
write the debit, mark the payment pending
commit
phase 2, no transaction, no locks:
call the bank with idempotency key = payment ID
record the result
Phase one takes the row locks, writes the debit, marks the payment pending and commits. Phase two calls the bank with no transaction open and no rows locked, then records the result. Lock waits and deadlocks in phase one are raised so that the job retries them; any other failure there marks the payment failed before any money has left.
Reversal instead of rollback
After the split, the outcome of the bank call is written as new ledger state rather than decided by a commit or a rollback. If the bank accepts the transfer, a settled status is recorded for the two debit lines and no new lines are written. If the bank rejects it, or its retries run out, a reversal transfer is appended in the opposite direction, from the external account back to the customer, and a failed status is recorded for all four lines. The customer’s balance returns to where it was, and the history shows the attempt and its reversal.
A rollback can only undo work inside a transaction that is still open, which is why the first design kept the transaction open for the length of the bank call. A compensating transfer undoes committed work at any later time and from any process: the job that sent the payment, a retry an hour later, or the recovery job. A reversal that runs twice finds the first one already written and writes nothing. That is what keeps the row lock to the tens of milliseconds of the ledger write instead of the length of an HTTP call.
Every payment reaches a final state
Each bank response is classified by its response code into one of two outcomes.
| Response | Meaning | Handling |
|---|---|---|
| Timeout, connection failure, authentication error, temporary error at the bank | The transfer may or may not have happened | Retry with growing delays, same idempotency key, ledger untouched |
| Rejection such as an invalid or closed account | The transfer did not and will not happen | Reverse at once |
| Retries exhausted | The outcome is still unknown | Reverse, and rely on the nightly reconciliation to catch a transfer the bank did execute |
Database lock timeouts and deadlocks are retried on their own short schedule, since they say nothing about the bank. The reversal path is a single recovery job: it writes whatever part of the ledger pair is missing, appends the reversal, marks the payment failed, and writes nothing on a second run.
The final status reaches the ledger through two independent streams. The primary stream is the bank’s webhook. The bank acknowledges each instant transfer at once and reports settled or rejected later, and the webhook enqueues a status job that settles or reverses the payment. A check scheduled thirty minutes after sending picks up any payment whose webhook never arrived. The second stream is a nightly reconciliation that compares every ledger balance with the bank’s balance for the same account and reports any difference.
Together, the two streams keep the ledger in line with the bank. During major outages at the payment processor, the reconciliation report showed which accounts differed, and those were corrected by hand.
Capacity
In production, a single worker replica has processed up to 100,000 payments per day. A load test measured how far one replica can go.
| Load test setup | |
|---|---|
| Worker | 1 process, 1 vCPU, 10 threads |
| Database | Postgres 16, 4 CPUs, commit latency about 6.7 ms, matching a zone-redundant managed server |
| Bank | Simulated, 500 ms median response, asynchronous acknowledgement |
| Data | 500 source accounts |
| Result | |
|---|---|
| Throughput | 10.3 payments per second |
| Job time, p50 / p95 | 930 ms / 1,200 ms |
| Worker CPU, p95 | 0.9 of 1 vCPU |
| Postgres CPU | 0.21 of 4 cores |
| Lock waits, errors | none |
One process can do roughly as many payments per second as its threads divided by the time a job holds a thread. The load test recorded the median job time, not the mean, so this is an approximation rather than a bound:
payments per second ≈ threads / median seconds per job = 10 / 0.93 ≈ 10.8
A thread sweep shows where the thread count stops being the limit and the CPU takes over. The 5- and 20-thread runs used the synchronous bank path.
| Threads | Payments per second | Job p50 | Worker CPU | Limited by |
|---|---|---|---|---|
| 5 | 6.4 | 751 ms | 0.53 | Threads |
| 10 | 10.3 | 930 ms | 0.78 | vCPU |
| 20 | 9.8 | 1,934 ms | 0.93 | vCPU |
Payments from a single source account ran at the same rate as payments spread over 500 accounts: the balance row is locked only for the ledger write.
| Daily volume from one worker | Payments per second | Per day |
|---|---|---|
| Measured rate, even over 24 hours | 10.3 | 890,000 |
| Measured rate, 8-hour window | 10.3 | 300,000 |
| Estimate with production round trips and webhook status jobs, 24 hours | about 8 | about 690,000 |
Past ten payments per second, more processes help and more threads do not. Each process adds its threads plus two to the database’s connection count. The other levers are persistent HTTP connections to the bank and a bank rate limit that permits the result.
Why this design
The design was chosen under two constraints. It had to ship quickly, inside the existing Rails application and its Postgres database, with no new infrastructure to run. And it had to be exact, because every job moves money. The first version moved payments off the web request without adding any infrastructure, and the later changes built on the same foundation.
It is not the design a larger system would end with. Throughput is capped per worker process, the job table shares the database with everything else, and retries, timers and waiting on webhooks live in hand-written job code. A payment whose retries run out while its outcome is unknown is reversed, and a transfer the bank did execute is found by reconciliation, not prevented; holding such payments in an unknown state until the bank confirms the outcome would be stricter.
| Alternative | What it adds | What it costs |
|---|---|---|
| A log broker, fed from an outbox table by change data capture | Dispatch separated from the database, replay, fan-out to other consumers | A broker and a relay to run; consumers still need the same idempotency keys |
| Event sourcing | Payment state as a replayable log; the append-only ledger is already halfway there | A rewrite of how payment state is stored and queried |
| A durable workflow engine | Retries, timers and webhook waits as workflow code instead of job state | A new runtime and a new programming model |
| Locks per account in an external lock service | Account locking outside the database | Rejected: an external lock can expire or be lost in a failover while its holder still acts on it, and making it safe needs fencing tokens checked by the database, which puts the database back in the path |
Checklist
- Treat the domain row as the outbox. Store the intent on the row the business already owns, key the job on that row’s ID, and make re-enqueueing by ID always correct.
- Enqueue inside the transaction that changes the state, and check the adapter’s
enqueue_after_transaction_commitsetting before relying on it. - Key every layer on the same ID: job concurrency, the ledger debit, the stored provider ID and the idempotency key.
- Commit the ledger before calling an external API, hold row locks for the database work only, and undo with reversal transfers instead of rollbacks.
- Size the database pool from the thread count plus GoodJob’s utility connections plus any threads started inside jobs.
- Set explicit timeouts on every HTTP client a job calls.
- Classify every provider response as retry or final, and reconcile balances with the provider on a schedule.
- Connect GoodJob workers to Postgres directly. Its advisory locks are held per session, so it does not work through a connection pooler in transaction mode.
Further reading
- Pattern: Transactional outbox, microservices.io. The pattern this design adapts.
- Compensating Transaction pattern, Azure Architecture Center. Undoing committed work with new work, as the ledger reversal does.
- 100X Faster: How We Supercharged Netflix Maestro’s Workflow Engine, Netflix Technology Blog. A queue table written in the same transaction as the state it describes, at a much larger scale.