- Kind
- Personal capstone
- Year
- 2026
- Role
- Design, build, tests, operations
- Stack
- Java 17Spring Boot 3.5PostgreSQL 15RabbitMQResilience4jFlywayMicrometerTestcontainers
- Source
- Private · walkthrough on request
Credit-card onboarding pipeline
A five-stage card onboarding service, built so that a slow vendor, a dead broker or a duplicate message ends in a defined state instead of a second credit enquiry.
The hard partA duplicate message here is not wasted work. It is a second hard enquiry on a real person's credit file.
- 7 ADRs, including the ones that record what is not built yet
- 138 tests; one runs the whole pipeline against real Postgres and RabbitMQ
- 8 Prometheus alert rules, 4 Grafana dashboards
The problem
Digital credit-card onboarding runs through stages that each depend on a system the bank does not control: a KYC check, a credit-bureau enquiry, a decision, account provisioning in core banking, and a copy back to the legacy platform. I built this on my own time to work through how that flow fails, and to make every failure end somewhere visible.
Run it synchronously and any slow vendor fails the whole application, and the customer starts again. In this domain "start again" is not neutral. A bureau pull costs money and leaves a hard enquiry on the applicant's credit file, and enquiry velocity is itself a signal the decision engine reads. A retry that runs twice damages the input to the next decision.
The source is private. The design, the trade-offs and the failure behaviour are all below, and a code walkthrough is available on request.
Shape of the system
A modular monolith: one Spring Boot deployable, one PostgreSQL, strict package boundaries. Intake is synchronous and does no external calls. Everything slow happens in RabbitMQ consumers, one queue per stage.
202. Each hop between stages goes through RabbitMQ, published only after the stage’s transaction commits.Intake answers before anything slow happens. POST /api/v1/applications
validates, checks the idempotency key, checks for a live duplicate
applicant, and then writes four rows in one transaction: the applicant,
the application, an outbox_events row, and an audit event. It returns
202 Accepted. No vendor has been called yet, so web threads are never
parked on someone else's latency.
The outbox closes the dual-write gap. Saving to Postgres and publishing to RabbitMQ are two systems; either can fail after the other succeeded. Writing the "message to send" as a row in the same transaction means an accepted application always has a durable event, even if the broker is down. A poller publishes pending rows every two seconds and marks them published, or counts the failure and tries again, up to five times before the row is marked failed for inspection.
Each stage does one external call outside any transaction, then commits a guarded state transition, then publishes the next message after the commit. The KYC check runs before the bureau call on purpose: pay for and record a credit enquiry only for an identity that verified.
Break it
Every number in this simulation is read from the service's configuration: three attempts with exponential backoff, a breaker that opens when half of the last calls failed, a 30-second open window, a two-second outbox poll. Pick a failure and watch where it ends.
Failure lab· simulated from the service’s config
In request
Client
mobile app
Intake API
filters, validation
Postgres
app + outbox, 1 tx
After 202
Outbox poller
every 2s
RabbitMQ
queue per stage
Stages
KYC
before any spend
Bureau
retry + breaker
circuit closed
Decision
pure evaluator
Core banking
claim, then call
Legacy mirror
vendor row
On nack
Dead letters
audited failure
Idempotency, in three places
At-least-once delivery means duplicates are a certainty. One guard is not enough, because duplicates arrive from three different directions.
The client. A phone that lost the 202 will post again. Requests
carry an Idempotency-Key; the service stores a SHA-256 of the
canonical request body next to it. Same key and same body replays the
stored response with 200. Same key with a different body is a 409,
because silently accepting it would mean the customer and the bank
disagree about what was applied for.
The applicant. A new key with the same PAN, while an earlier application is still in flight, is also refused. PAN is stored AES-256-GCM-encrypted with a random IV, so the ciphertext cannot be searched; a separate SHA-256 column carries the lookup, and Postgres enforces it with a partial unique index on live rows only, so erased applicants (India's DPDP Act) do not block a new application.
The broker. Each consumer re-reads the application and skips work
whose state has already moved on. A redelivered application.created
finds the application past RECEIVED and is acknowledged without a
second bureau call.
Decisions
The trade-offs are written up as architecture decision records. The short version of each:
RabbitMQ, not Kafka. This is a low-volume, stage-shaped work queue. It needs per-message acks, dead-lettering and independent consumers, not retained replay or partition throughput. RabbitMQ gives those with less to operate. The cost is no replay, and retry progression has to be designed explicitly.
Transactional outbox, not publish-after-commit. The dual-write problem is real and the outbox is the cheapest correct fix. The cost: delivery is at-least-once, so every consumer has to be idempotent.
Choreography, not an orchestrator. The flow is linear. Each stage consumes, transitions, publishes. No workflow engine to run. The cost: the flow is spread across bindings and consumers, so the application's status column is the one place that shows progress.
A modular monolith first. One deployable makes the outbox trivial, because it shares a database with the writer, and keeps one log stream for debugging. Package boundaries are strict so a stage can be extracted later if it needs to scale on its own; the bureau and decision services exist as extraction scaffolds, not yet live.
Encrypted column plus a hash. AES-GCM must never use a fixed IV, so deterministic lookup needs a second column. The ADR says plainly that a plain SHA-256 is weaker than a keyed HMAC, which is the next item below.
What it does not handle yet
Each of these is written down in the repository, with the fix.
Concurrent duplicates at the bureau stage. The status guard protects
against a duplicate that arrives after the first finished. Two copies
processed at the same moment both see RECEIVED and both call the
bureau. The core-banking stage already avoids this by claiming the work
(DECISION_MADE → CORE_BANKING_PENDING, flushed) before its external
call; the bureau stage should do the same with a conditional update, and
pass the application id to the vendor as an idempotency key.
A crash between the vendor call and the commit. The failure lab's
last scenario. Core banking creates the account, the pod dies, the
redelivered message hits the guard and is skipped, and the application
stays in CORE_BANKING_PENDING with an account that exists. It needs a
reconciliation sweep that asks core banking about stale pending rows,
and a durable consumer inbox so message deduplication stops being
coupled to business state (ADR-007, proposed).
Outbox row claiming. With more than one replica, two pollers can
publish the same pending row. Consumers tolerate the duplicate, but the
fix is SELECT … FOR UPDATE SKIP LOCKED.
Key management. The PAN hash should be an HMAC with a key held in a KMS, and field encryption should move to envelope encryption with a versioned key id. Both are specified in ADRs, neither is built.
Operating it
Micrometer metrics feed Prometheus. Eight alert rules cover what an on-call person would need to act on: intake error rate and latency, each circuit breaker opening, queue depth, anything in a dead-letter queue, a spike in duplicate-applicant refusals, and PII appearing in logs. Logs mask PAN and date of birth at the layout level, and every request carries a correlation id from the HTTP filter through the outbox row into the message headers. Four Grafana dashboards are provisioned from the repository.
CI runs the 138 tests with a 70% coverage gate. One integration test starts real PostgreSQL and RabbitMQ in Testcontainers and drives an application through the entire pipeline. A nightly job runs OWASP Dependency-Check, Trivy and CodeQL.
Outcome
Every failure path ends in a named, audited state. Retries and the circuit breaker are configured per vendor, poisoned work dead-letters into a terminal failed state, and a repeated submission returns the original response instead of opening a second application.