Track Each Record in an Import
Run a bounded Batch Job with implemented success and rejection adapters, then reason about bulk outcomes and replay.
A nightly order export contains 80,000 rows. Most are valid, a few carry malformed amounts, and halfway through the run a request times out. Restarting the whole file might duplicate work; abandoning it might leave orders missing. A loop is only a small part of that problem.
The fixture in this chapter is three records, small enough to inspect by hand, and it stays that way while the job of tracking each record moves out of our own code. Batch adds a record lifecycle and completion report — a million rows first would only make those semantics harder to see.
Batch is an Enterprise Edition capability. Checkpoint 21 retains the local service but adds a separate import ledger. Its status processed means the quantity check passed and the local adapter recorded that result. It does not submit the record to the order-acceptance API.

Open this checkpoint in ACB
Stop the previous local application, then use File → Open Folder to open book/checkpoints/21-batch-import. Open src/main/mule/app.xml and use Flow List to select the named flow for each example. Start the controlled dependency in a separate terminal from the companion root with python3 book/stubs/server.py. Choose Run and Debug → Run Mule Application and wait for deployment. Save canvas edits and use Save and Hot-deploy to Local Runtime before repeating a request.
Run python3 book/run.py verify 21 from the companion root to exercise this running checkpoint with its synthetic fixtures. The verifier supplies requests and checks results; it does not start the ACB application. Keep the editor on this checkpoint while reading a failure so an old deployment cannot supply a misleading answer.
Initialize the destination before launching
The ledger is the only place this exercise records what happened to each row, so it has to start empty; otherwise a previous run’s outcomes are indistinguishable from this one’s. POST /lab/import/setup creates and resets import_outcomes(order_id, status). Then POST book/fixtures/import-records.json to /lab/batch. The launch returns 202 with started: true; read GET /lab/import/outcomes until all three identities have an outcome.
Example 056 — Classify a three-record Batch import
The complete source is in checkpoints/21-batch-import/src/main/mule/app.xml.
Choose Flow List → start-batch. The Listener accepts POST /lab/batch. Select Batch Job: Job Name SmallImport, Block Size 2, Max Failed Records -1. Expand its processing steps in order:
| Step | Acceptance and processors |
|---|---|
check-quantity | Choice condition payload.qty <= 0; matching route raises APP:BAD_QUANTITY |
save-valid-records | Accept Policy NO_FAILURES; Flow Reference record-import-success |
save-rejected-records | Accept Policy ONLY_FAILURES; Flow Reference record-import-rejection |
Expand On Complete and inspect its Logger message expression:
write({loaded: payload.loadedRecords, failed: payload.failedRecords,
successful: payload.successfulRecords}, 'application/json')
After the Batch Job, Set Variable assigns numeric 202 to httpStatus, and Transform Message returns JSON {started: true}. These components acknowledge launch; the On Complete section reports the job outcome.
The first step raises an error for quantity zero. The successful-record step accepts records without failures. The rejection step accepts records with failures. IMPORT-1 and IMPORT-3 therefore reach the success adapter; IMPORT-2 reaches the rejection adapter.
maxFailedRecords="-1" allows this small job to continue through record failures — it is not a sensible universal response to every outage. blockSize="2" controls internal blocks, not a business group that must commit together. On Complete receives a report with record counts, not the transformed source array. The logger reads that report after processing.
Example 057 — Record an accepted import row
The complete source is in checkpoints/21-batch-import/src/main/mule/app.xml.
Choose Flow List → record-import-success. Select Database Update, connection Orders_DB. In Advanced, its Target Variable is written, preserving the record payload. Its SQL is:
MERGE INTO import_outcomes(order_id,status) KEY(order_id) VALUES (:id,'processed')
Set Input Parameters to expression {id: payload.orderId}.
Example 058 — Record a rejected import row
The complete source is in checkpoints/21-batch-import/src/main/mule/app.xml.
Choose Flow List → record-import-rejection. Its Database Update uses Orders_DB, Target Variable written and Input Parameters {id: payload.orderId}. Its SQL records the other disposition:
MERGE INTO import_outcomes(order_id,status) KEY(order_id) VALUES (:id,'rejected')
Both adapters are supplied subflows. Their H2 MERGE uses the order ID as a key, so repeating this fixed fixture updates the same local outcome row. That is a narrow demonstration ledger. A filename is no substitute for that key: tomorrow’s export may reuse it with different contents, and “retry the same export” has to stay distinguishable from “process a newer export with the same name”. A real versioned import needs an import/source identity and ordering rule so a later replay cannot overwrite newer results.
The rejection adapter can fail too. If its database write fails, the system has not preserved a usable rejection merely because the original quantity error was recognized. Retain the immutable source and identify the import attempt so unresolved rows can be investigated.

Batch Job block size and maximum failed-record settings.
Read business outcomes beside runtime counts
The expected ledger is two processed rows and one rejected row. A handled rejection still originated as a failed record in the Batch lifecycle; do not interpret a runtime success counter as the destination’s business acceptance count without understanding the adapter. At 80,000 rows that gap becomes the question an operations team actually asks — why did the source contain 80,000 rows and the destination receive 79,700 — and a counter that cannot separate the two kinds of outcome cannot answer it.
An acceptance expression can skip a step without failing the record. If a delivery step selects only USD, an EUR row may bypass it without becoming a recorded rejection. When unsupported currency must be reported, implement that classification explicitly. Each source row needs a terminal business classification: delivered, deliberately excluded with a reason, rejected, or unresolved pending reconciliation.
Group delivery only when the destination supports it
A fixed Batch Aggregator collects eligible records into arrays for a bulk operation. Its size is separate from internal blockSize; the former must respect the destination’s request limit — that bound belongs to the destination’s contract, and an internal block size cannot change it. A final group can be smaller than the configured maximum. A streaming aggregator instead imposes sequential-access constraints and cannot be combined with a fixed size setting. Batch phases.
Suppose a destination accepts 100 items and returns HTTP 200 with 98 successes and two item errors. The adapter must correlate and store all 100 results. HTTP success establishes the request outcome, not acceptance of every record. Raising one blanket error after partial success loses the distinction unless item evidence was retained first.
Conversely, if the destination commits all 100 and the response disappears, the batch owns an uncertain group. Stable item identities and a destination idempotency/version contract permit reconciliation or safe replay. The Batch Job does not add a transaction around those remote effects.
Failures and recovery at scale
A failed-record limit can stop further processing; concurrent work can make the observed count exceed that threshold. It does not roll back earlier external commits. Whether continuing is useful depends on the kind of failure: working through malformed rows is often the point of a dirty historical file, while for a rejected credential every remaining block achieves nothing but noise. Distinguish malformed records from systemic failures such as a rejected credential before deciding whether continued attempts are useful. Batch failure handling.
Recovery evidence has to outlive the thing that failed, which means keeping the immutable input and the outcome ledger outside storage that can disappear with a replica. The local H2 ledger makes this exercise observable — it does not establish CloudHub replacement durability. Runtime restart recovery and replay from a retained source are separate recovery paths; an operating design should know which remains available after local state is lost.
More concurrency consumes connections, memory and destination quota as well as CPU, so the measures worth watching are committed item outcomes and unresolved work rather than throughput alone — more records submitted per second are no use if more of them stay unresolved. Making the launch return sooner or loading records faster does not prove that the destination completed more business work. A finite nightly import with per-record recovery and bulk delivery earns this machinery; a continuous event feed usually needs a message-driven design with explicit acknowledgement and backpressure, rather than an endlessly growing “nightly” input.
Try it
1. Reconcile the small import. What should the three ledger rows say, and why is the initial 202 insufficient?
Show answer
IMPORT-1 and IMPORT-3 are processed; IMPORT-2 is rejected. The 202 acknowledges launch, before the asynchronous job has established those per-record results.
2. Interpret a skipped currency. Is a row rejected merely because a step acceptance expression evaluates false?
Show answer
No. Skipping that step is selection, not a recorded business rejection. Add an explicit classification path when the contract requires one.
3. Inspect a mixed bulk response. A 100-item HTTP request returns 98 successes and two item errors. What must survive before retrying?
Show answer
Persist results correlated to all source identities, separating confirmed effects, rejected items and uncertain outcomes. Replay only under a destination identity/version rule that tolerates repeated or older work.
Batch tracks work that must be explained. A cache holds answers that may be recomputed. Confusing those two kinds of state makes recovery much harder.
Comments