Skip to content

SDK-674 Add switchProject for runtime project switching - #1091

Open
joaodordio wants to merge 7 commits into
feature/switch-project-apifrom
feature/SDK-674-switch-project
Open

joaodordio wants to merge 7 commits into
feature/switch-project-apifrom
feature/SDK-674-switch-project

Conversation

@joaodordio

@joaodordio joaodordio commented Aug 25, 2026 •

Copy link
Copy Markdown
Member

📝 Summary

Adds IterableAPI.switchProject(project:callback:), moving a running install between Iterable projects in place with no restart and no state leaking across.

🎟️ Jira Ticket: SDK-674

📖 Description

Multi-region apps have no way today to move a running install from one Iterable project to another without an app restart, which forces either a relaunch or separate builds per region. This adds a single entry point that tears the outgoing project down in place and stands up a replacement instance against the new key and config.

The sequence disables the push token on the outgoing project, resets its in-app and embedded state, purges its persisted offline queue while preserving queued disableDevice tasks, clears its identity from memory and storage along with the rest of its project-scoped storage (activation criteria, unsent unknown-user events and sessions, stored attribution), releases the outgoing instance so its deinit runs, and re-initializes. The method returns immediately and does the work off the calling thread.

Three decisions worth reviewer attention.

Not everything is queued. Calls made during the switch window are held and replayed, but only where replaying against the new project is the desired outcome. Anything carrying the outgoing project's campaign, template or message IDs runs against the project the SDK is on when it is made, because replaying it would attribute it to a project those IDs do not exist in. Push handling is deliberately outside the gate for that reason, and because holding the OS completion handler for the length of a switch risks the watchdog. IterableAppIntegration's behaviour is unchanged for apps that never call switchProject. The gated set matches Android: setEmail, setUserId, updateEmail, updateUser, updateCart, trackPurchase, track, registerToken, logoutUser and the email/userId setters.

The teardown waits for the disable handoff. A request built with the outgoing API key cannot be sent once its instance is gone, and OfflineRequestProcessor holds its auth provider weakly, so releasing the instance before the disable reached the request layer dropped it silently without reporting a failure, leaving the device registered on the project the app had left. The teardown now awaits the handoff, meaning persisted to the queue or in flight online, with a bound matching Android's DISABLE_DISPATCH_TIMEOUT_MS. It deliberately does not await the network response, which in offline mode may never arrive.

attributionInfo refuses late writes. The setter drops writes once the project has been switched away from. Work the outgoing instance owns can resolve after the clear, a universal link redirect returning being the obvious case, and writing it back would attribute the new project's first event to a campaign it has never heard of.

The callback is delivered on the main thread. true means every cleanup step completed and a device disable was confirmed. false means the SDK is on the new project but a step was noisy; it never means the switch failed or rolled back, and it is the expected result for an app that does not use push. switchProject does not re-identify the user; call setEmail/setUserId from the callback.

Paired with SDK-673 on Android, which implements the same contract. Divergence between the two is a bug, so the gated call set, the false conditions and the dispatch bound were all aligned deliberately.

🧪 How to test?

Full suite green on iPhone 17 Pro Max: 837 tests, 0 failures (719 unit-tests, 15 notification-extension-tests, 103 offline-events-tests).

New coverage:

  • SwitchProjectTests (17): identity and storage clearing, manager rebuild, instance deallocation, config swap, FIFO drain, rapid and repeated switches, no-op on unchanged key, callback thread, false conditions.
  • SwitchProjectGateTests (8): a push tapped mid-switch is not handed to the new project, the silent-push completion handler is not held, the gate holds identity while releasing project-scoped calls, disable handoff signalling and its bound, attribution not written back after teardown, a failed re-init reporting false, empty and blank API key rejection.
  • SwitchProjectOfflineTests: the teardown waits for the disable to reach the queue, and a queued disable survives the switch still targeting the outgoing project.
  • RequestHandlerTests/testDeleteAllTasksOnLogout was stale (asserting count == 0 against the preserve-disableDevice behaviour from SDK-297) and sat in skippedTests. Fixed and re-enabled.

Manually: initialize against project A, identify a user, register for push, then switchProject to project B. Confirm the device is disabled on A, that B has no trace of A's identity, criteria, unknown-user events or attribution, and that calls made during the window land on B in order.

🧾 Changelog

Added, see CHANGELOG.md. Documents the callback contract, the conditions producing false, which calls are queued versus run inline, and the disable handoff wait.

📹 Loom recording if applicable

N/A.

🐞 Github Issues solved

None.

📚 Docs PR if applicable

Required, tracked by SDK-676 for both platforms. Not yet opened.


Known, not addressed here

  • A purged offline task's Pending never resolves, so setEmail's completion handler is stranded when a queued register is discarded by a logout purge. Android fixes this in SDK-675; iOS has the same gap and it predates this change.
  • unit-tests/testCriteriaUserIdTokenCheck is load-sensitive rather than flaky in the usual sense, and unrelated to this change. It waits 5 seconds inside a 6 second timeout, so it has under a second of headroom. On an unloaded machine it passes at 5.17s; on a run where the unit target took 1242s instead of 296s it exceeded the bound and failed. Worth widening that bound or removing the fixed wait, separately from this PR.

The project pair (added after review feedback)

switchProject now takes an IterableProject, which pairs a project's API key with the IterableConfig to run it with:

guard let euProject = IterableProject(apiKey: euApiKey, config: config) else { return }
IterableAPI.switchProject(project: euProject) { cleanTeardown in ... }

IterableConfig carries dataRegion and authDelegate, so a loose key and config allowed one project's key to be combined with another project's region or auth delegate. That mismatch surfaces as an auth failure, or a request sent to the wrong region, that looks unrelated to the project it came from. Pairing them makes it unrepresentable.

The initializer is failable and returns nil for an empty or whitespace-only key, so an unusable project is not constructible. Android throws IllegalArgumentException for the same case: the guarantee is identical, the mechanism follows each platform's convention, since a blank key is a normal outcome of a failed region lookup and trapping on it would be wrong.

The apiKey/config form is now internal, used by the tests that inject a dependency container. Replaced rather than overloaded because this is unreleased; keeping both public would leave the loose form available forever.

One caveat documented on the type: IterableConfig holds its delegates weak, and this does not change that. A paired value object invites being held as a static let, so if the auth delegate is only retained by the config it deallocates and auth stops working silently. Android holds its handlers strongly and does not have this problem. This is the one way the abstraction makes something more dangerous rather than less, so it is called out in the doc comment and in the customer guide.

Also corrected in this branch: the JWT retry budget note previously claimed an exhausted budget could suppress the new project's first token request. It cannot. The retry state is private instance state on AuthManager and the switch builds a fresh dependency container, so the new project starts clean. Comment and CHANGELOG only.


The result type (added after review feedback)

The callback reports an IterableProjectSwitchResult, either .switchedCleanly or .switchedWithWarnings, rather than a boolean.

The boolean was the problem the customer guide kept having to write around: false did not mean the switch failed, did not mean it was rolled back, was the expected outcome for any app without push, and called for exactly the same handling as true. A value needing that much explanation to avoid being misread is the wrong shape. Both case names lead with "switched" so the part that holds in either case, that the SDK is on the new project, is visible at the call site instead of being something the reader has to recall.

Declared @objc with an Int raw value, so Objective-C callers get IterableProjectSwitchResultSwitchedCleanly and friends. That rules out associated values, which is the main reason there is no warning detail attached to the warnings case.

Internal plumbing still carries the teardown boolean the individual steps set, including the gate's queued callbacks, and converts once in the public wrapper. That is why the tests injecting a dependency container still drive the internal overload unchanged.

One thing worth a reviewer's opinion: adding a case to a public Swift enum is source-breaking for exhaustive switches in customer code. If we ever want the noisy step named in the result rather than only in the logs, it is much cheaper to decide that before release than after. I left it at two cases rather than guessing at a taxonomy.

Test totals after this change: 841 (723 unit, 15 notification-extension, 103 offline-events), 0 failures.

Multi-region apps currently have no way to move a running install from
one Iterable project to another without an app restart, which forces
either a relaunch or shipping separate builds per region.

switchProject tears down the outgoing project in place and stands up a
replacement instance against the new key and config. It disables the
push token on the outgoing project, resets its in-app and embedded
state, purges its persisted offline queue while preserving queued
disableDevice tasks, and clears its identity and project-scoped storage
so nothing replays into the incoming project.

Calls made during the switch window are queued and replayed in order,
but only where replaying against the new project is the desired
outcome. Anything carrying the outgoing project's campaign, template or
message IDs runs against the project the SDK is on when it is made,
since replaying it would attribute it to a project those IDs do not
exist in. Push handling is deliberately outside the gate for that
reason, and because holding the OS completion handler for the length of
a switch risks the watchdog.

The teardown waits, with a bound, for the outgoing disableDevice to
reach the request layer before releasing that instance. A request built
with the outgoing API key cannot be sent once its instance is gone, and
the offline processor holds its auth provider weakly, so a release that
won that race dropped the disable without reporting a failure. It does
not wait for the network response, which in offline mode may never
arrive.

The callback reports true only when every cleanup step was clean and a
device disable was confirmed. false means the SDK is on the new project
but a step was noisy, never that the switch failed or rolled back, and
it is the expected result for an app that does not use push.

Mirrors the Android implementation in SDK-673, including the gated call
set and the disable dispatch bound.
@joaodordio joaodordio self-assigned this Aug 25, 2026
@joaodordio joaodordio added the BCIT Run Business Critical Integration Tests label Aug 25, 2026
@codecov

codecov Bot commented Aug 25, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.66667% with 40 lines in your changes missing coverage. Please review.
✅ Project coverage is 75.02%. Comparing base (d97e76e) to head (62412c6).

Files with missing lines Patch % Lines
swift-sdk/SDK/IterableAPI.swift 56.86% 22 Missing ⚠️
swift-sdk/Internal/InternalIterableAPI.swift 91.91% 11 Missing ⚠️
swift-sdk/SDK/IterableAPI+SwitchProject.swift 98.23% 4 Missing ⚠️
...l/api-client/Request/OfflineRequestProcessor.swift 84.61% 2 Missing ⚠️
swift-sdk/SDK/IterableAppIntegration.swift 80.00% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1091      +/-   ##
==========================================
+ Coverage   73.67%   75.02%   +1.35%     
==========================================
  Files         114      118       +4     
  Lines       10202    10618     +416     
==========================================
+ Hits         7516     7966     +450     
+ Misses       2686     2652      -34     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

The note claimed the previous project's exhausted JWT retry budget could suppress
the new project's first token request until pauseAuthRetries(false) was called.
That is not what the code does. The retry state is private instance state on
AuthManager, and switchProject builds a fresh dependency container, so the new
project gets a new AuthManager starting at retryCount 0 with pauseAuthRetry false.

Comment and CHANGELOG only, no behaviour change. Android carried the same
inaccurate note and is corrected alongside this in SDK-673.
Comment thread swift-sdk/SDK/IterableAPI+SwitchProject.swift Outdated
Comment thread swift-sdk/SDK/IterableAPI+SwitchProject.swift Outdated
Comment thread swift-sdk/Internal/InternalIterableAPI.swift
config: config,
apiEndPointOverride: apiEndPointOverride,
dependencyContainer: dependencyContainer,
callback: { started in deliverOnMainThread(callback, started) })

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When called before initialization, iOS normally reports true if initialization succeeds. Android reports false, since no previous project was cleaned up and no device disable was confirmed. The iOS documentation also describes that case as false.

This is a parity issue - should this result be aligned across both platforms?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The successful cold-start case is aligned now, but the initialization-failure case still diverges.
Swift forwards initialize2's started value, so a failed start() reports .switchedWithWarnings; Android's pre-initialization branch unconditionally reports SWITCHED_CLEANLY.

Since result parity is part of this contract, suggest aligning this edge case and add a failing-start pre-initialization test.

switchProject took (apiKey, config) as two independently supplied halves of one
thing. IterableConfig carries the dataRegion and the authDelegate, so nothing
stopped one project's key being paired with another project's region or auth
delegate. That mismatch produces an auth failure, or a request sent to the wrong
region, that looks unrelated to the project it came from, and the adoption guide had
to warn about it at length. Pairing the two makes it unrepresentable.

The public signature is now switchProject(project:callback:). The apiKey/config form
is internal, used by the tests that inject a dependency container. Replacing rather
than overloading because this is unreleased: keeping both public would leave the
loose form available forever and reintroduce exactly what the pair removes.

IterableProject's initializer is failable and returns nil for an empty or
whitespace-only key, so an unusable project is not constructible and a blank key can
no longer reach a teardown. Android throws IllegalArgumentException for the same
case; the guarantee is identical, the mechanism follows each platform's convention,
since a blank key is a normal outcome of a failed region lookup and trapping on it
would be wrong here.

IterableConfig holds its delegates weakly and this does not change that, so the
doc comment warns that a long-lived project needs its delegates retained elsewhere.
The callback handed back a bare boolean, and the meaning of false had to be
explained everywhere it appeared: it does not mean the switch failed, it does not
mean it was rolled back, it is expected in normal operation for any app without
push, and the correct response to it is identical to the correct response to true.
A return value that needs that much prose to avoid being misread is carrying the
wrong shape.

The callback now receives an IterableProjectSwitchResult of .switchedCleanly or
.switchedWithWarnings. Both names lead with "switched" so the thing that is true in
both cases, that the SDK is on the new project, is unmissable at the call site
rather than something the reader has to remember from the docs.

Declared @objc with an Int raw value so Objective-C callers get
IterableProjectSwitchResultSwitchedCleanly and friends. That rules out associated
values, which is the main reason there is no warning detail attached to the
warnings case.

Internal plumbing still carries the teardown boolean set by the individual steps,
including the gate's queued callbacks. It converts once in the public wrapper
through IterableProjectSwitchResult.from, so the tests that inject a dependency
container keep working against the internal overload unchanged.

Note for a future change: adding a case to a public Swift enum is source-breaking
for exhaustive switches in customer code, so if a warning reason is ever wanted it
is cheaper to decide before release than after.
…n review

Three defects and one documentation correction, all from PR review.

A switchProject for a second project while a switch was still running was
silently coalesced with the first. The SDK finished on the first project while
the second caller's callback reported a completed switch, so an app whose region
or brand picker was tapped twice sent every later event to the project the user
had already moved away from, with nothing in the API to detect it. The gate now
tells the two kinds of second request apart by where they are headed: one asking
for the project already in flight joins it, so a picker tapped twice on the same
destination still runs one teardown, and one asking for a different project is
queued and run when the switch in flight lands.

The gate is handed to that queued request rather than lowered and raised again.
Lowering it in between leaves a window in which a brand new switchProject takes
the gate first and is then overtaken by the older queued request, so the SDK
settles on the project the app asked for second-to-last while every callback
still reports success. The switch callback is exactly where an app makes that
next call from, which puts it in that window. A queued request therefore runs
with resumingGate set, and owns releasing the gate it inherited if it bails out
before the teardown starts, since a gate left raised would queue every later SDK
call and never replay it.

The gate was also lowered before the queue was replayed, so a setEmail made as
the switch landed could run ahead of the calls already waiting behind the gate
and then be overwritten by the older one replayed behind it, which is the FIFO
order the gate exists to provide. endSwitch now decides the queue is empty while
holding the same lock queueOrExecute enqueues under, and lowers the gate only
then.

Attribution was guarded by a flag that was read and then written outside a lock.
That is check-then-act: a universal link redirect resolving over the network
could pass the check, have the switch clear attribution, and then put the
previous project's campaign back into the storage the new project reads, so the
new project's first attributed event would carry a campaignId it has never heard
of. The attributionInfo setter and clearProjectScopedStorage now take one lock,
so a redirect either writes before the clear and is cleared by it, or sees the
flag and stands down.

Documented that a switch called before initialize reports .switchedCleanly.
There is no previous project, so no teardown step could have been noisy. Android
was reporting warnings there and has been aligned.
Master had moved one commit ahead, the 6.7.5 release prep, which cut a new
[6.7.5] heading and an empty [Unreleased] above it.

Everything auto-merged, but the CHANGELOG merged wrong in a way git could not
see. This branch's switchProject entries were written under [Unreleased], and
that heading is now [6.7.5], so the merge filed all of them under a version that
shipped without switchProject. Moved them back under [Unreleased]; [6.7.5] is
now byte for byte what master says it contains.
@joaodordio
joaodordio marked this pull request as ready for review September 8, 2026 02:43
@joaodordio
joaodordio requested a review from a team as a code owner September 8, 2026 02:43
Comment thread swift-sdk/SDK/IterableAPI+SwitchProject.swift Outdated
…r last

Review found four ordering defects in the switch gate. The first three are one
mistake: the gate asked what project the SDK was on, when the question is what
project the app most recently asked to be on. Those two differ for the whole
length of a teardown, which includes the device disable wait.

- A request for the project being left looked like "already there", so it
  reported a clean switch, tore nothing down, and let the switch in flight land
  anyway. The app was then told it was on A while every event went to B.
- A repeat of the destination in flight joined it even with another destination
  queued behind it, so the chain settled past the project asked for last.
- A repeat was merged into any earlier queued entry, so C, D, C ran C then D and
  reported the last request when the first one landed.

One routing rule replaces all three. The requested project is the tail of the
queue, or the switch in flight when the queue is empty, or the live key when
nothing is running. A request for it is absorbed by whatever is going to deliver
it, anything else goes on the tail, so adjacent repeats still share a single
teardown and a double tapped picker behaves as before.

The fourth was introduced by the previous handover fix. One flag was both
queueing calls and serializing switch requests, so holding it across a handover
queued the setEmail the contract tells apps to make from the callback and then
replayed it into the next project in the chain. isRaised now covers only call
queueing and drops as each switch lands, while inFlightApiKey marks the chain
and keeps new requests ordered behind it. The next switch is dispatched through
the main queue so it runs after the callbacks endSwitch posted there.

Each defect has a test that fails without the fix.
@joaodordio
joaodordio changed the base branch from master to feature/switch-project-api September 23, 2026 15:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

BCIT Run Business Critical Integration Tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants