Problem
MetricBatchWriter uses an unbounded ConcurrentLinkedQueue with no size cap:
// MetricBatchWriter.scala
private val queue = new ConcurrentLinkedQueue[MetricRow]()
def enqueue(row: MetricRow): Unit = { queue.add(row) }
The background flush thread drains this queue every 5 s (configurable). If the database is unavailable or slow, the flush stalls while incoming requests continue to enqueue MetricRow objects. Because the queue is unbounded, this can fill the heap and trigger an OutOfMemoryError.
Failure scenario
- Database goes down (or latency spikes).
flush() blocks / retries while requests keep calling queue.add().
- Queue grows without bound → heap exhaustion → OOM.
MetricRow objects hold references to response body strings, further accelerating memory pressure.
Proposed fix
Add a bounded capacity and a drop-oldest eviction policy:
private val MAX_QUEUE_SIZE = 50_000
private val queue = new java.util.concurrent.ArrayBlockingQueue[MetricRow](MAX_QUEUE_SIZE)
def enqueue(row: MetricRow): Unit = {
if (!queue.offer(row)) {
// drop oldest to make room, then retry once
queue.poll()
queue.offer(row)
logger.warn(s"MetricBatchWriter: queue full — oldest metric dropped (capacity=)")
}
}
Why drop-oldest: metrics are time-series data. Keeping recent traffic visible is more valuable than preserving stale backlog that will never catch up. The warn log makes the drop observable.
Additional notes
- The capacity limit (50 000) is a starting point; tune based on observed flush rate × acceptable latency.
- Alternatively, expose it as a prop (
metrics.queue.max_size).
- A bounded queue also means
OOM from this path now appears as a log warn rather than a JVM crash — a meaningful operability improvement.
Related: introduced in the same PR that fixed the MetricTest flaky race (fire-and-forget Future → sync enqueue).
Problem
MetricBatchWriteruses an unboundedConcurrentLinkedQueuewith no size cap:The background flush thread drains this queue every 5 s (configurable). If the database is unavailable or slow, the flush stalls while incoming requests continue to enqueue
MetricRowobjects. Because the queue is unbounded, this can fill the heap and trigger anOutOfMemoryError.Failure scenario
flush()blocks / retries while requests keep callingqueue.add().MetricRowobjects hold references to response body strings, further accelerating memory pressure.Proposed fix
Add a bounded capacity and a drop-oldest eviction policy:
Why drop-oldest: metrics are time-series data. Keeping recent traffic visible is more valuable than preserving stale backlog that will never catch up. The warn log makes the drop observable.
Additional notes
metrics.queue.max_size).OOMfrom this path now appears as a log warn rather than a JVM crash — a meaningful operability improvement.Related: introduced in the same PR that fixed the
MetricTestflaky race (fire-and-forget Future → sync enqueue).