SDK-674 Add switchProject for runtime project switching - #1091
joaodordio wants to merge 7 commits into
Conversation
Multi-region apps currently have no way to move a running install from one Iterable project to another without an app restart, which forces either a relaunch or shipping separate builds per region. switchProject tears down the outgoing project in place and stands up a replacement instance against the new key and config. It disables the push token on the outgoing project, resets its in-app and embedded state, purges its persisted offline queue while preserving queued disableDevice tasks, and clears its identity and project-scoped storage so nothing replays into the incoming project. Calls made during the switch window are queued and replayed in order, but only where replaying against the new project is the desired outcome. Anything carrying the outgoing project's campaign, template or message IDs runs against the project the SDK is on when it is made, since replaying it would attribute it to a project those IDs do not exist in. Push handling is deliberately outside the gate for that reason, and because holding the OS completion handler for the length of a switch risks the watchdog. The teardown waits, with a bound, for the outgoing disableDevice to reach the request layer before releasing that instance. A request built with the outgoing API key cannot be sent once its instance is gone, and the offline processor holds its auth provider weakly, so a release that won that race dropped the disable without reporting a failure. It does not wait for the network response, which in offline mode may never arrive. The callback reports true only when every cleanup step was clean and a device disable was confirmed. false means the SDK is on the new project but a step was noisy, never that the switch failed or rolled back, and it is the expected result for an app that does not use push. Mirrors the Android implementation in SDK-673, including the gated call set and the disable dispatch bound.
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #1091 +/- ##
==========================================
+ Coverage 73.67% 75.02% +1.35%
==========================================
Files 114 118 +4
Lines 10202 10618 +416
==========================================
+ Hits 7516 7966 +450
+ Misses 2686 2652 -34 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
The note claimed the previous project's exhausted JWT retry budget could suppress the new project's first token request until pauseAuthRetries(false) was called. That is not what the code does. The retry state is private instance state on AuthManager, and switchProject builds a fresh dependency container, so the new project gets a new AuthManager starting at retryCount 0 with pauseAuthRetry false. Comment and CHANGELOG only, no behaviour change. Android carried the same inaccurate note and is corrected alongside this in SDK-673.
| config: config, | ||
| apiEndPointOverride: apiEndPointOverride, | ||
| dependencyContainer: dependencyContainer, | ||
| callback: { started in deliverOnMainThread(callback, started) }) |
There was a problem hiding this comment.
When called before initialization, iOS normally reports true if initialization succeeds. Android reports false, since no previous project was cleaned up and no device disable was confirmed. The iOS documentation also describes that case as false.
This is a parity issue - should this result be aligned across both platforms?
There was a problem hiding this comment.
The successful cold-start case is aligned now, but the initialization-failure case still diverges.
Swift forwards initialize2's started value, so a failed start() reports .switchedWithWarnings; Android's pre-initialization branch unconditionally reports SWITCHED_CLEANLY.
Since result parity is part of this contract, suggest aligning this edge case and add a failing-start pre-initialization test.
switchProject took (apiKey, config) as two independently supplied halves of one thing. IterableConfig carries the dataRegion and the authDelegate, so nothing stopped one project's key being paired with another project's region or auth delegate. That mismatch produces an auth failure, or a request sent to the wrong region, that looks unrelated to the project it came from, and the adoption guide had to warn about it at length. Pairing the two makes it unrepresentable. The public signature is now switchProject(project:callback:). The apiKey/config form is internal, used by the tests that inject a dependency container. Replacing rather than overloading because this is unreleased: keeping both public would leave the loose form available forever and reintroduce exactly what the pair removes. IterableProject's initializer is failable and returns nil for an empty or whitespace-only key, so an unusable project is not constructible and a blank key can no longer reach a teardown. Android throws IllegalArgumentException for the same case; the guarantee is identical, the mechanism follows each platform's convention, since a blank key is a normal outcome of a failed region lookup and trapping on it would be wrong here. IterableConfig holds its delegates weakly and this does not change that, so the doc comment warns that a long-lived project needs its delegates retained elsewhere.
The callback handed back a bare boolean, and the meaning of false had to be explained everywhere it appeared: it does not mean the switch failed, it does not mean it was rolled back, it is expected in normal operation for any app without push, and the correct response to it is identical to the correct response to true. A return value that needs that much prose to avoid being misread is carrying the wrong shape. The callback now receives an IterableProjectSwitchResult of .switchedCleanly or .switchedWithWarnings. Both names lead with "switched" so the thing that is true in both cases, that the SDK is on the new project, is unmissable at the call site rather than something the reader has to remember from the docs. Declared @objc with an Int raw value so Objective-C callers get IterableProjectSwitchResultSwitchedCleanly and friends. That rules out associated values, which is the main reason there is no warning detail attached to the warnings case. Internal plumbing still carries the teardown boolean set by the individual steps, including the gate's queued callbacks. It converts once in the public wrapper through IterableProjectSwitchResult.from, so the tests that inject a dependency container keep working against the internal overload unchanged. Note for a future change: adding a case to a public Swift enum is source-breaking for exhaustive switches in customer code, so if a warning reason is ever wanted it is cheaper to decide before release than after.
…n review Three defects and one documentation correction, all from PR review. A switchProject for a second project while a switch was still running was silently coalesced with the first. The SDK finished on the first project while the second caller's callback reported a completed switch, so an app whose region or brand picker was tapped twice sent every later event to the project the user had already moved away from, with nothing in the API to detect it. The gate now tells the two kinds of second request apart by where they are headed: one asking for the project already in flight joins it, so a picker tapped twice on the same destination still runs one teardown, and one asking for a different project is queued and run when the switch in flight lands. The gate is handed to that queued request rather than lowered and raised again. Lowering it in between leaves a window in which a brand new switchProject takes the gate first and is then overtaken by the older queued request, so the SDK settles on the project the app asked for second-to-last while every callback still reports success. The switch callback is exactly where an app makes that next call from, which puts it in that window. A queued request therefore runs with resumingGate set, and owns releasing the gate it inherited if it bails out before the teardown starts, since a gate left raised would queue every later SDK call and never replay it. The gate was also lowered before the queue was replayed, so a setEmail made as the switch landed could run ahead of the calls already waiting behind the gate and then be overwritten by the older one replayed behind it, which is the FIFO order the gate exists to provide. endSwitch now decides the queue is empty while holding the same lock queueOrExecute enqueues under, and lowers the gate only then. Attribution was guarded by a flag that was read and then written outside a lock. That is check-then-act: a universal link redirect resolving over the network could pass the check, have the switch clear attribution, and then put the previous project's campaign back into the storage the new project reads, so the new project's first attributed event would carry a campaignId it has never heard of. The attributionInfo setter and clearProjectScopedStorage now take one lock, so a redirect either writes before the clear and is cleared by it, or sees the flag and stands down. Documented that a switch called before initialize reports .switchedCleanly. There is no previous project, so no teardown step could have been noisy. Android was reporting warnings there and has been aligned.
Master had moved one commit ahead, the 6.7.5 release prep, which cut a new [6.7.5] heading and an empty [Unreleased] above it. Everything auto-merged, but the CHANGELOG merged wrong in a way git could not see. This branch's switchProject entries were written under [Unreleased], and that heading is now [6.7.5], so the merge filed all of them under a version that shipped without switchProject. Moved them back under [Unreleased]; [6.7.5] is now byte for byte what master says it contains.
…r last Review found four ordering defects in the switch gate. The first three are one mistake: the gate asked what project the SDK was on, when the question is what project the app most recently asked to be on. Those two differ for the whole length of a teardown, which includes the device disable wait. - A request for the project being left looked like "already there", so it reported a clean switch, tore nothing down, and let the switch in flight land anyway. The app was then told it was on A while every event went to B. - A repeat of the destination in flight joined it even with another destination queued behind it, so the chain settled past the project asked for last. - A repeat was merged into any earlier queued entry, so C, D, C ran C then D and reported the last request when the first one landed. One routing rule replaces all three. The requested project is the tail of the queue, or the switch in flight when the queue is empty, or the live key when nothing is running. A request for it is absorbed by whatever is going to deliver it, anything else goes on the tail, so adjacent repeats still share a single teardown and a double tapped picker behaves as before. The fourth was introduced by the previous handover fix. One flag was both queueing calls and serializing switch requests, so holding it across a handover queued the setEmail the contract tells apps to make from the callback and then replayed it into the next project in the chain. isRaised now covers only call queueing and drops as each switch lands, while inFlightApiKey marks the chain and keeps new requests ordered behind it. The next switch is dispatched through the main queue so it runs after the callbacks endSwitch posted there. Each defect has a test that fails without the fix.
📝 Summary
🎟️ Jira Ticket: SDK-674
📖 Description
Multi-region apps have no way today to move a running install from one Iterable project to another without an app restart, which forces either a relaunch or separate builds per region. This adds a single entry point that tears the outgoing project down in place and stands up a replacement instance against the new key and config.
The sequence disables the push token on the outgoing project, resets its in-app and embedded state, purges its persisted offline queue while preserving queued
disableDevicetasks, clears its identity from memory and storage along with the rest of its project-scoped storage (activation criteria, unsent unknown-user events and sessions, stored attribution), releases the outgoing instance so itsdeinitruns, and re-initializes. The method returns immediately and does the work off the calling thread.Three decisions worth reviewer attention.
Not everything is queued. Calls made during the switch window are held and replayed, but only where replaying against the new project is the desired outcome. Anything carrying the outgoing project's campaign, template or message IDs runs against the project the SDK is on when it is made, because replaying it would attribute it to a project those IDs do not exist in. Push handling is deliberately outside the gate for that reason, and because holding the OS completion handler for the length of a switch risks the watchdog.
IterableAppIntegration's behaviour is unchanged for apps that never callswitchProject. The gated set matches Android:setEmail,setUserId,updateEmail,updateUser,updateCart,trackPurchase,track,registerToken,logoutUserand theemail/userIdsetters.The teardown waits for the disable handoff. A request built with the outgoing API key cannot be sent once its instance is gone, and
OfflineRequestProcessorholds its auth provider weakly, so releasing the instance before the disable reached the request layer dropped it silently without reporting a failure, leaving the device registered on the project the app had left. The teardown now awaits the handoff, meaning persisted to the queue or in flight online, with a bound matching Android'sDISABLE_DISPATCH_TIMEOUT_MS. It deliberately does not await the network response, which in offline mode may never arrive.attributionInforefuses late writes. The setter drops writes once the project has been switched away from. Work the outgoing instance owns can resolve after the clear, a universal link redirect returning being the obvious case, and writing it back would attribute the new project's first event to a campaign it has never heard of.The callback is delivered on the main thread.
truemeans every cleanup step completed and a device disable was confirmed.falsemeans the SDK is on the new project but a step was noisy; it never means the switch failed or rolled back, and it is the expected result for an app that does not use push.switchProjectdoes not re-identify the user; callsetEmail/setUserIdfrom the callback.Paired with SDK-673 on Android, which implements the same contract. Divergence between the two is a bug, so the gated call set, the
falseconditions and the dispatch bound were all aligned deliberately.🧪 How to test?
Full suite green on
iPhone 17 Pro Max: 837 tests, 0 failures (719 unit-tests, 15 notification-extension-tests, 103 offline-events-tests).New coverage:
SwitchProjectTests(17): identity and storage clearing, manager rebuild, instance deallocation, config swap, FIFO drain, rapid and repeated switches, no-op on unchanged key, callback thread,falseconditions.SwitchProjectGateTests(8): a push tapped mid-switch is not handed to the new project, the silent-push completion handler is not held, the gate holds identity while releasing project-scoped calls, disable handoff signalling and its bound, attribution not written back after teardown, a failed re-init reportingfalse, empty and blank API key rejection.SwitchProjectOfflineTests: the teardown waits for the disable to reach the queue, and a queued disable survives the switch still targeting the outgoing project.RequestHandlerTests/testDeleteAllTasksOnLogoutwas stale (assertingcount == 0against the preserve-disableDevicebehaviour from SDK-297) and sat inskippedTests. Fixed and re-enabled.Manually: initialize against project A, identify a user, register for push, then
switchProjectto project B. Confirm the device is disabled on A, that B has no trace of A's identity, criteria, unknown-user events or attribution, and that calls made during the window land on B in order.🧾 Changelog
Added, see
CHANGELOG.md. Documents the callback contract, the conditions producingfalse, which calls are queued versus run inline, and the disable handoff wait.📹 Loom recording if applicable
N/A.
🐞 Github Issues solved
None.
📚 Docs PR if applicable
Required, tracked by SDK-676 for both platforms. Not yet opened.
Known, not addressed here
Pendingnever resolves, sosetEmail's completion handler is stranded when a queued register is discarded by a logout purge. Android fixes this in SDK-675; iOS has the same gap and it predates this change.unit-tests/testCriteriaUserIdTokenCheckis load-sensitive rather than flaky in the usual sense, and unrelated to this change. It waits 5 seconds inside a 6 second timeout, so it has under a second of headroom. On an unloaded machine it passes at 5.17s; on a run where the unit target took 1242s instead of 296s it exceeded the bound and failed. Worth widening that bound or removing the fixed wait, separately from this PR.The project pair (added after review feedback)
switchProjectnow takes anIterableProject, which pairs a project's API key with theIterableConfigto run it with:IterableConfigcarriesdataRegionandauthDelegate, so a loose key and config allowed one project's key to be combined with another project's region or auth delegate. That mismatch surfaces as an auth failure, or a request sent to the wrong region, that looks unrelated to the project it came from. Pairing them makes it unrepresentable.The initializer is failable and returns
nilfor an empty or whitespace-only key, so an unusable project is not constructible. Android throwsIllegalArgumentExceptionfor the same case: the guarantee is identical, the mechanism follows each platform's convention, since a blank key is a normal outcome of a failed region lookup and trapping on it would be wrong.The
apiKey/configform is now internal, used by the tests that inject a dependency container. Replaced rather than overloaded because this is unreleased; keeping both public would leave the loose form available forever.One caveat documented on the type:
IterableConfigholds its delegatesweak, and this does not change that. A paired value object invites being held as astatic let, so if the auth delegate is only retained by the config it deallocates and auth stops working silently. Android holds its handlers strongly and does not have this problem. This is the one way the abstraction makes something more dangerous rather than less, so it is called out in the doc comment and in the customer guide.Also corrected in this branch: the JWT retry budget note previously claimed an exhausted budget could suppress the new project's first token request. It cannot. The retry state is private instance state on
AuthManagerand the switch builds a fresh dependency container, so the new project starts clean. Comment and CHANGELOG only.The result type (added after review feedback)
The callback reports an
IterableProjectSwitchResult, either.switchedCleanlyor.switchedWithWarnings, rather than a boolean.The boolean was the problem the customer guide kept having to write around:
falsedid not mean the switch failed, did not mean it was rolled back, was the expected outcome for any app without push, and called for exactly the same handling astrue. A value needing that much explanation to avoid being misread is the wrong shape. Both case names lead with "switched" so the part that holds in either case, that the SDK is on the new project, is visible at the call site instead of being something the reader has to recall.Declared
@objcwith anIntraw value, so Objective-C callers getIterableProjectSwitchResultSwitchedCleanlyand friends. That rules out associated values, which is the main reason there is no warning detail attached to the warnings case.Internal plumbing still carries the teardown boolean the individual steps set, including the gate's queued callbacks, and converts once in the public wrapper. That is why the tests injecting a dependency container still drive the internal overload unchanged.
One thing worth a reviewer's opinion: adding a case to a public Swift enum is source-breaking for exhaustive switches in customer code. If we ever want the noisy step named in the result rather than only in the logs, it is much cheaper to decide that before release than after. I left it at two cases rather than guessing at a taxonomy.
Test totals after this change: 841 (723 unit, 15 notification-extension, 103 offline-events), 0 failures.