Skip to content
Rohit Behera
← All work
Kind
Personal capstone
Year
2026
Role
Design, build, tests, operations
Stack
Java 17Spring Boot 3.5PostgreSQL 15RabbitMQResilience4jFlywayMicrometerTestcontainers
Source
Private · walkthrough on request

Credit-card onboarding pipeline

A five-stage card onboarding service, built so that a slow vendor, a dead broker or a duplicate message ends in a defined state instead of a second credit enquiry.

The hard partA duplicate message here is not wasted work. It is a second hard enquiry on a real person's credit file.

  • 7 ADRs, including the ones that record what is not built yet
  • 138 tests; one runs the whole pipeline against real Postgres and RabbitMQ
  • 8 Prometheus alert rules, 4 Grafana dashboards

The problem

Digital credit-card onboarding runs through stages that each depend on a system the bank does not control: a KYC check, a credit-bureau enquiry, a decision, account provisioning in core banking, and a copy back to the legacy platform. I built this on my own time to work through how that flow fails, and to make every failure end somewhere visible.

Run it synchronously and any slow vendor fails the whole application, and the customer starts again. In this domain "start again" is not neutral. A bureau pull costs money and leaves a hard enquiry on the applicant's credit file, and enquiry velocity is itself a signal the decision engine reads. A retry that runs twice damages the input to the next decision.

The source is private. The design, the trade-offs and the failure behaviour are all below, and a code walkthrough is available on request.

Shape of the system

A modular monolith: one Spring Boot deployable, one PostgreSQL, strict package boundaries. Intake is synchronous and does no external calls. Everything slow happens in RabbitMQ consumers, one queue per stage.

Onboarding pipeline architectureInside the request, the intake API validates, checks the idempotency key and duplicate applicants, then writes the applicant, application, outbox event and audit event to PostgreSQL in one transaction and returns 202. An outbox poller publishes pending events to RabbitMQ every two seconds. Four consumers run in sequence: bureau (KYC first, then the credit bureau), decision, core banking, and legacy mirror. Calls to vendors retry three times behind a circuit breaker; messages that still fail are dead-lettered and moved to an audited failed state.INSIDE THE REQUESTno vendor is called hereClientmobile appIntake APIauth · rate limit · validationidempotency key · PAN-hash dedupPostgreSQL — one transactionapplicants PAN AES-256-GCM, pan_hashapplications RECEIVED, idempotency_keyoutbox_events PENDING ← the message to sendaudit_events append-only202 Accepted, in millisecondsRELAYat-least-onceOutbox pollerevery 2 s · ≤5 tries, then FAILEDRabbitMQdirect exchangequeue per stagePENDING rowsSTAGESeach: one vendor call, then a guarded transitionBureau stageKYC first, thenthe credit bureauDecisionpure evaluatorrecord + status, 1 txCore bankingclaim the work,then provisionLegacy mirrorvendor-shaped rowDoneCOMPLETEDVENDORS & FAILUREKYC + bureauexternalCore banking systemexternalretry ×3 · breakerretry ×3 · breakerDead-letter queuesconsumer → *_FAILED, auditedmetric + DLQMessagesPresent alertnack once retries are spent
Solid arrows happen inside the customer’s request; everything below the first lane happens after the 202. Each hop between stages goes through RabbitMQ, published only after the stage’s transaction commits.

Intake answers before anything slow happens. POST /api/v1/applications validates, checks the idempotency key, checks for a live duplicate applicant, and then writes four rows in one transaction: the applicant, the application, an outbox_events row, and an audit event. It returns 202 Accepted. No vendor has been called yet, so web threads are never parked on someone else's latency.

The outbox closes the dual-write gap. Saving to Postgres and publishing to RabbitMQ are two systems; either can fail after the other succeeded. Writing the "message to send" as a row in the same transaction means an accepted application always has a durable event, even if the broker is down. A poller publishes pending rows every two seconds and marks them published, or counts the failure and tries again, up to five times before the row is marked failed for inspection.

Each stage does one external call outside any transaction, then commits a guarded state transition, then publishes the next message after the commit. The KYC check runs before the bureau call on purpose: pay for and record a credit enquiry only for an identity that verified.

Break it

Every number in this simulation is read from the service's configuration: three attempts with exponential backoff, a breaker that opens when half of the last calls failed, a 30-second open window, a two-second outbox poll. Pick a failure and watch where it ends.

Failure lab· simulated from the service’s config

Inject a failure

The bureau stops answering while applications keep arriving.

After 5 failed calls the circuit opens and later applications fail fast instead of piling onto a vendor that is already struggling. Each lands in a visible, audited failed state, not limbo.

In request

  1. Client

    mobile app

  2. Intake API

    filters, validation

  3. Postgres

    app + outbox, 1 tx

After 202

  1. Outbox poller

    every 2s

  2. RabbitMQ

    queue per stage

Stages

  1. KYC

    before any spend

  2. Bureau

    retry + breaker

    circuit closed

  3. Decision

    pure evaluator

  4. Core banking

    claim, then call

  5. Legacy mirror

    vendor row

On nack

  1. Dead letters

    audited failure

Status—

    Idempotency, in three places

    At-least-once delivery means duplicates are a certainty. One guard is not enough, because duplicates arrive from three different directions.

    The client. A phone that lost the 202 will post again. Requests carry an Idempotency-Key; the service stores a SHA-256 of the canonical request body next to it. Same key and same body replays the stored response with 200. Same key with a different body is a 409, because silently accepting it would mean the customer and the bank disagree about what was applied for.

    The applicant. A new key with the same PAN, while an earlier application is still in flight, is also refused. PAN is stored AES-256-GCM-encrypted with a random IV, so the ciphertext cannot be searched; a separate SHA-256 column carries the lookup, and Postgres enforces it with a partial unique index on live rows only, so erased applicants (India's DPDP Act) do not block a new application.

    The broker. Each consumer re-reads the application and skips work whose state has already moved on. A redelivered application.created finds the application past RECEIVED and is acknowledged without a second bureau call.

    Decisions

    The trade-offs are written up as architecture decision records. The short version of each:

    RabbitMQ, not Kafka. This is a low-volume, stage-shaped work queue. It needs per-message acks, dead-lettering and independent consumers, not retained replay or partition throughput. RabbitMQ gives those with less to operate. The cost is no replay, and retry progression has to be designed explicitly.

    Transactional outbox, not publish-after-commit. The dual-write problem is real and the outbox is the cheapest correct fix. The cost: delivery is at-least-once, so every consumer has to be idempotent.

    Choreography, not an orchestrator. The flow is linear. Each stage consumes, transitions, publishes. No workflow engine to run. The cost: the flow is spread across bindings and consumers, so the application's status column is the one place that shows progress.

    A modular monolith first. One deployable makes the outbox trivial, because it shares a database with the writer, and keeps one log stream for debugging. Package boundaries are strict so a stage can be extracted later if it needs to scale on its own; the bureau and decision services exist as extraction scaffolds, not yet live.

    Encrypted column plus a hash. AES-GCM must never use a fixed IV, so deterministic lookup needs a second column. The ADR says plainly that a plain SHA-256 is weaker than a keyed HMAC, which is the next item below.

    What it does not handle yet

    Each of these is written down in the repository, with the fix.

    Concurrent duplicates at the bureau stage. The status guard protects against a duplicate that arrives after the first finished. Two copies processed at the same moment both see RECEIVED and both call the bureau. The core-banking stage already avoids this by claiming the work (DECISION_MADE → CORE_BANKING_PENDING, flushed) before its external call; the bureau stage should do the same with a conditional update, and pass the application id to the vendor as an idempotency key.

    A crash between the vendor call and the commit. The failure lab's last scenario. Core banking creates the account, the pod dies, the redelivered message hits the guard and is skipped, and the application stays in CORE_BANKING_PENDING with an account that exists. It needs a reconciliation sweep that asks core banking about stale pending rows, and a durable consumer inbox so message deduplication stops being coupled to business state (ADR-007, proposed).

    Outbox row claiming. With more than one replica, two pollers can publish the same pending row. Consumers tolerate the duplicate, but the fix is SELECT … FOR UPDATE SKIP LOCKED.

    Key management. The PAN hash should be an HMAC with a key held in a KMS, and field encryption should move to envelope encryption with a versioned key id. Both are specified in ADRs, neither is built.

    Operating it

    Micrometer metrics feed Prometheus. Eight alert rules cover what an on-call person would need to act on: intake error rate and latency, each circuit breaker opening, queue depth, anything in a dead-letter queue, a spike in duplicate-applicant refusals, and PII appearing in logs. Logs mask PAN and date of birth at the layout level, and every request carries a correlation id from the HTTP filter through the outbox row into the message headers. Four Grafana dashboards are provisioned from the repository.

    CI runs the 138 tests with a 70% coverage gate. One integration test starts real PostgreSQL and RabbitMQ in Testcontainers and drives an application through the entire pipeline. A nightly job runs OWASP Dependency-Check, Trivy and CodeQL.

    Outcome

    Every failure path ends in a named, audited state. Retries and the circuit breaker are configured per vendor, poisoned work dead-letters into a terminal failed state, and a repeated submission returns the original response instead of opening a second application.

    loading index…Full retrieval trace →