osiris_writer: Track diff between writer and replica offsets - #197
the-mikedavis wants to merge 2 commits into
Conversation
8272b66 to
0fa0df1
Compare
|
In addition to this (or maybe instead of this) we could track replica freshness. The stream coordinator is calculates "freshness" as a requirement in |
0fa0df1 to
680b7a9
Compare
|
I updated this to perform the same calculation as the stream coordinator does: https://github.com/rabbitmq/rabbitmq-server/blob/765d2c5d748f1a3227b97e966a31a73f4b561867/deps/rabbit/src/rabbit_stream_coordinator.erl#L209-L221 (Also see discussion in rabbitmq/rabbitmq-server#15098) So we could use this metric instead of querying the replication state with a call. |
680b7a9 to
82b2ac6
Compare
|
Tick the box to add this pull request to the merge queue (same as
|
82b2ac6 to
36a9e40
Compare
|
Tick the box to add this pull request to the merge queue (same as
|
|
I fixed some edge-cases where the gauge would spike during initialization, and I added a sum of offsets that need replication, the "replication backlog". The total value is not useful probably, but seeing it grow or shrink would let you know how replication is doing. |
|
Tick the box to add this pull request to the merge queue (same as
|
This introduces a metric calculated at every batch which records the difference between the timestamps of the last chunks in the logs of the writer and its replicas, and the sum of offsets that need to be replicated. This can be used to watch replicas catch up on replication, or to diagnose situations when replication is being starved out (network-wise) by high-throughput publishing. The replication diff should usually be low but, for streams seeing traffic, non-zero. This calculation is also used in the stream coordinator when adding a member, via `osiris_writer:query_replication_state/1`. With this change the stream coordinator could be updated to use the counter instead of calling the writer.
Exposes the replica_staleness gauge directly through seshat/ETS. Callers that just need the aggregate ms lag can skip the gen_batch_server:call round-trip and pending-batch queueing that query_replication_state/1 has.
eac1202 to
e08f25e
Compare
|
Tick the box to add this pull request to the merge queue (same as
|
This introduces a metric calculated at every batch which records the difference between the timestamps of the last chunks in the logs of the writer and its replicas. This can be used to watch replicas catch up on replication, or to diagnose situations when replication is being starved out (network-wise) by high-throughput publishing. The replication diff should usually be low but, for streams seeing traffic, non-zero.
This calculation is also used in the stream coordinator when adding a member, via
osiris_writer:query_replication_state/1. With this change the stream coordinator could be updated to use the counter instead of calling the writer.