Appendix D: Synchronization and Platform Variants
Keep SFTP completion, ordered watermarks, token refresh, managed state and Connected Mode governance as independent workshops.
A scheduled job needs a durable answer to “where did we finish?” A token cache needs a trustworthy answer to “is this credential still usable?” Both can store small values, but their loss and concurrency rules differ. This appendix separates those workflows instead of presenting them as one generic Object Store example.
Use chapters 16, 18 and 22 before synchronization; chapters 23–24 before token work; and chapter 29 before the Connected Mode gateway variant. Each workshop below states the additional environment it needs. The remote procedures are not prerequisites for the local capstone.
Prepare the workshops in ACB
Use Explorer to open book/workshops/synchronization/page.sql beside checkpoint 16’s Scheduler and checkpoint 22’s Object Store components. The SQL editor holds the source query; the canvas shows where a source adapter, destination handoff and cursor update would have to occur. The supplied query alone is not a runnable synchronization application.
For the SFTP variant, use the component palette to locate the SFTP source and inspect its connection, matching, polling and post-action settings in a separate sandbox project. Populate them only from the agreed server contract. For token refresh, keep credentials out of inline expressions and use a provider-specific configuration. For Connected Mode, edit the configuration artifacts in ACB and perform registration through the documented protected terminal procedure. None of these external services is created by opening the editor.
Remote file handoff: prerequisites and completion
A remote SFTP workshop requires a controlled server, agreed host key, synthetic input directory, producer completion protocol, client credential and archive/quarantine permissions. Keep immutable source bytes and a delivery identity/checksum. The local File round trip established serialization and attributes only.
Polling requires an eligibility rule
A polling source starts work when it discovers files that meet an eligibility rule agreed with the producer. For the partner’s midnight file, that rule decides whether your flow starts on a half-written delivery. A common design uses a temporary filename while writing and a final name after completion; another uses a separate ready marker carrying the expected count and checksum — either way the producer, not the consumer, declares completion.
Those are design choices, not guarantees supplied by a filename suffix. A rename must have the required behavior on the actual filesystem or SFTP server. If the producer writes directly to the final name, the consumer needs another way to know that all bytes are present.
The File and SFTP sources offer controls such as polling frequency, filename matching and checks for changes in size. A delay between size checks can reduce the chance of reading a file still being written, but a producer can pause longer than that interval. Stable size is a heuristic unless the delivery contract makes it authoritative. The File source configuration describes these controls; the late-writer scenario was not simulated in the local round trip.
When two replicas can see the same inbox, polling also needs a claim mechanism. An in-process variable cannot arbitrate between separate Mule processes — a source setting, server-side move, lease table or external scheduler may participate in that design instead. Your test has to start two consumers and show what prevents both from performing the business effect.
Watermarks do not identify business completion
Replay is where the difference between a watermark and a business record shows. A watermark helps a polling source remember a discovery position, commonly using file metadata such as modification time to decide what to consider next. It is not the same record as “order A-1001 was committed successfully.”
A corrected file can keep its name; a copied file can receive a new modification time; a failure can occur after some rows have committed. Replaying solely from the discovery position can either miss work or repeat effects. A delivery ledger tied to stable delivery and record IDs expresses the business state more directly.
For example, an orders feed can record a delivery identifier, checksum, received time, total row count, accepted count, rejected count and final disposition. If the same identifier arrives with different bytes, the consumer can flag a conflict. If a retry carries the same bytes after an interrupted run, row-level idempotency can prevent repeating accepted orders. This ledger is a proposed application design; it was not implemented by the small file test.
Moving that delivery into archive/ gives operators a useful location for its bytes; the ledger explains which orders were accepted and which still need recovery.
Success actions and failure actions
Where you handle an error affects what the source sees after processing. Convert a failed import into successful completion, and the source can run its success action — moving or deleting the file according to its configuration — even though the business objective was not met.
File and SFTP expose post-action settings, including whether actions apply after failure. Their exact configuration belongs in a deployment review because deletion can remove the only replay input. The SFTP reference describes move, deletion and failure-related controls. No remote post-action was run for this chapter.
A practical design keeps immutable received bytes until its retention policy permits deletion. Invalid deliveries move to a quarantine location with a reason and an operator-visible status. Retryable infrastructure failures remain distinguishable from permanently invalid data. The absence of a file from the inbox should never be the only explanation of what happened to it.
Consider a ten-thousand-row file whose final row contains an invalid date. If earlier rows were already committed independently, moving the whole file back to the inbox repeats them. Recovery needs either an atomic unit that truly includes all rows or a replay design that recognizes prior successes. The transactions and reliability chapters explain why an HTTP call, a file move and a database write do not automatically share one transaction.
The additional SFTP boundary
SFTP introduces a remote host whose identity must be checked, authentication material that must be protected, and a network that can fail between operations. A connection configuration contains host, port, user and authentication settings; production values belong outside the flow XML. The remote account needs the permissions required by the agreed inbox and archive operations.
Suppose a first connection fails and someone disables host verification to make it succeed. That removes a different check from the one a corrected password would fix. Checking the server’s host key authenticates the server to the client; presenting the client’s password or private key authenticates the client to the server.
Your deployment test should include an unexpected host key, an expired or rotated client credential, a denied archive move and an interrupted transfer. Expected outcomes include a clear failure category and retention of enough evidence to retry deliberately. The connector’s ability to connect once does not establish those operational behaviors.
The local example establishes File Write serialization, File Read parsing and file attributes. It is a foundation for a remote test, not a substitute for one.
A synchronization cursor needs a source contract
The next workshop requires a source database with stable ordering and precision, an authoritative retained cursor, one declared scheduler owner and an idempotent destination. Its reference query supplies the ordering decision; the adapters depend on those real services and are not silently supplied by a flow name.
Watermarking, or picking up where you left off
A watermark records a completed position in a source. You read it before querying and advance it only after the selected work is durably accepted by the destination. Object Store supplies storage for the cursor; it does not make source selection and destination writes one transaction. Mule’s watermark example illustrates the basic pattern.
A timestamp alone is often insufficient for paging, because two orders can share updated_at and advancing after processing only one can skip the other. Use a stable tiebreaker such as order ID and store both fields. The query and cursor comparison must use the same ordering — including timestamp precision and the ID tiebreaker.
This SQL fragment describes a proposed orders-source contract:
Example 085 — Select a page with a stable tiebreaker
Source: workshops/synchronization/page.sql.
SELECT id, updated_at, total
FROM orders
WHERE updated_at > :last_time
OR (updated_at = :last_time AND id > :last_id)
ORDER BY updated_at, id
LIMIT :page_size
The fragment assumes a source with a stable ordering and consistent timestamp precision, and its parameter and limit syntax must be checked for your actual database driver. Bind values through Database connector input parameters, and retain the selected rows or final cursor in a variable before destination operations replace payload.
The algorithm is: read the initialized cursor, select a bounded ordered page, retain its final tuple, durably complete or hand off the destination work, then advance the cursor. An empty page leaves it unchanged. Unexpected absence of an established cursor requires a recovery decision; it must not silently replay from an arbitrary old date.
No readNextOrderPage or writePageIdempotently implementation is hidden behind this description. Implement those adapters against the selected source and destination, then interrupt the run after each boundary before claiming a synchronization guarantee.
Write-last still permits duplicates
If all destination writes succeed and the cursor update fails, the next run reads that page again. That is safer than skipping it, provided the destination recognizes the same order/version. If the cursor advances before destination completion, a crash can lose work permanently. The trade is deliberate: replay with idempotency instead of silent omission.
Partial page completion needs the same decision. Either make the destination operation atomic for the page or tolerate reprocessing successful rows. A Batch job that starts asynchronously does not satisfy the “page finished” boundary merely because the launching flow returns. Wait for the appropriate completion signal or store a separate durable handoff that now owns the work.
Concurrent schedulers introduce a second race. Two runs can read the same cursor and process overlapping pages, then an older run can overwrite a newer cursor. A single scheduler setting on one process is not a guarantee across all replicas. Use a supported single-owner scheduler topology or an external lease/checkpoint mechanism whose atomicity is established.
A source row can also appear late with an earlier timestamp. The tuple cursor does not solve that. The source contract may need a lookback window with deduplication, a commit sequence or change-data capture. Deletions also need an explicit representation. Watermarking is correct only relative to the source’s change semantics; ordering a query does not invent those semantics.
Token refresh needs a validity and ownership rule
This workshop requires an actual provider contract: expiry fields, supported refresh/grant behavior, whether issuing a token invalidates earlier tokens, and a permitted way to coordinate refresh across deployed callers. Test only with a disposable provider/client configuration.
Store an envelope containing the token’s expiry and required context. Check validity with a margin for clock uncertainty and request duration before use. A cache TTL merely controls storage — it cannot make an expired or revoked token valid. If a token expires at 15:00 and the policy refreshes two minutes early, a 14:59 cache hit must not authorize its use.
Two replicas can both read a miss, both obtain a token and both store it. If a fresh token revokes the prior one, one replica may now hold an invalid value. A single refresh owner or supported distributed lease can coordinate the operation, but its acquisition, expiry and ownership checks must be designed together. A local lock does not span independent replicas.
Treat backend unavailability separately from absent data so an outage does not cause every replica to refresh simultaneously. Never log a token to prove refresh worked; record safe identities, expiry metadata and outcome instead.
Object Store backends are not interchangeable
A local store, a Mule cluster’s shared state and managed Object Store v2 have different storage and coordination boundaries. persistent="true" does not turn a standalone process into the managed service. Individual connector operations do not create a transaction spanning a provider call and several stores.
Managed v2 retention, value-size limits and account request allowances need an explicit capacity/retention review. Static and rolling TTL semantics differ; do not carry a local maxEntries assumption into managed sizing. A cursor whose required recovery window exceeds the backend’s retention belongs in a more appropriate authoritative store. Object Store v2 guide.
A shared dictionary also does not establish a shared lock. Verify coordination under the actual deployment topology, including a crashed owner and a lease expiring during slow work. Payment and order idempotency still belong around the authoritative business effect, not an expiring token marker.
Connected Mode gateway workshop
This variant requires an authorized Omni registration identity, organization/environment, gateway entitlement and access to API Manager. Use a separate registration directory from the Local Mode example. The registration command selects Connected Mode; startup then consumes the generated identity, and the route/policies come from the managed API instance.
Create an API instance associated with the registered gateway, set the consumer listener and upstream path to the same /api/orders contract used in chapter 29, and apply one policy at a time. Do not drop a Local Mode ApiInstance file beside it and assume that file overrides the managed route.
The technical registration command uses flexctl registration create with --connected=true and the selected organization, environment and connected-app credentials. Follow the version’s control-plane-specific procedure in a protected execution environment because command arguments can expose substituted secrets. Store generated registration material privately. Connected Mode registration.
After registration, test startup, a valid request, a policy rejection and an unavailable upstream. Then disconnect management on a warm instance and attempt a separate cold start. Retain mode, gateway image digest, API instance, route and policy versions with the observations. No registration or gateway traffic was run here.
Try it
1. A file stops growing briefly. Does unchanged size prove the producer finished?
Show answer
Only if the agreed producer protocol makes it authoritative. A paused writer can resume. A tested final-name handoff or ready marker with checksum/count provides a clearer completion boundary.
2. Complete a page, lose the cursor update. What must the destination permit on the next run?
Show answer
Safe replay of already completed items using stable identity and source version. Advancing the cursor first risks omission; advancing last deliberately permits replay. Timestamp plus ID handles ties but not late older commits or deletions.
3. Refresh on two replicas. What does atomic Store fail to protect?
Show answer
The workflow across reading, provider invocation and result publication. Both replicas can call the provider before either stores. Coordinate the full refresh protocol under the actual topology.
4. Change gateway mode. Why keep registration and route ownership explicit?
Show answer
Connected and Local modes have different authorities for API/policy configuration. A second file is not an automatic override. Use separate identities/configuration and establish the intended source before testing request behavior.
Comments