Skip to Main Content

MongoByte MongoDB Logo

Welcome to the new MongoDB Feedback Portal!

{Improvement: "Your idea"}
We’ve upgraded our system to better capture and act on your feedback.
Your feedback is meaningful and helps us build better products.

Status Submitted
Created by ayla c
Created on Sep 28, 2026

Add an opt-in change stream mode to emit shard key–changing updates as delete + insert

What problem are you trying to solve?

Focus on the what and why of the need you have, not the how you'd like it solved.

CDC pipelines that read change streams on sharded collections often key target rows by documentKey (shard key + _id), because _id alone is not guaranteed to be globally unique across a sharded cluster. When an update changes a shard key value, the key used to identify the downstream row changes. A pipeline keyed by documentKey must therefore remove the row under the old key and materialize it under the new key.

Currently, how this is reported depends on whether the new shard key belongs to the same shard:

  • Different shard. The shard raises WouldChangeOwningShard, and mongos runs a delete + insert in a transaction. Change streams emit delete (old key) + insert (new key).

  • Same shard. The write remains an update in the oplog, logged as a diff with the old key in o2. Change streams emit a single update event whose documentKey is the old key.

In the same-shard case, an upsert-based pipeline either leaves a stale row under the old key or writes new values under the wrong key. updateLookup can return null after the shard key changes, because the lookup uses the old documentKey. Obtaining the exact event-time post-image through the change stream requires changeStreamPreAndPostImages.

The same logical operation produces different event shapes depending on chunk placement. Placement changes over time with balancing and migrations, so applications reading the change stream must handle both shapes.

What would you like to see happen?

Describe the desired outcome or enhancement.

Provide an opt-in change stream mode that emits shard key–changing updates as delete with the old documentKey, followed by insert with the new documentKey and the full post-update document, regardless of whether the document moves to another shard.

This would give applications reading change streams a consistent representation of downstream key changes, independent of chunk placement. The delete must be emitted before the insert, and each event must have its own resume token, so that resuming from the delete's token still delivers the insert.

An implementation that relies on post-images would already remove the custom split logic from every pipeline. Ideally, this would also work without requiring collection-wide pre-image capture solely to handle rare shard key changes. For example, same-shard shard key updates could be written as delete + insert, the way cross-shard updates already are.

Why is this important to you or your team?

Explain how the request adds value or solves a business need.

We operate a CDC platform that replicates many sharded collections for internal teams. If a shard key–changing update is not handled correctly, the target silently keeps stale or mis-keyed rows. Nothing fails, so the drift goes unnoticed until someone finds inconsistent data.

To prevent this today, we either restrict shard key updates on CDC-enabled collections, which constrains application teams, or maintain custom handling in the pipeline for an event shape that depends on chunk placement. Emitting shard key–changing updates consistently as delete + insert would let us support shard key updates without restricting them or relying on custom handling.

What steps, if any, are you taking today to manage this problem?

We enable changeStreamPreAndPostImages, open the stream with fullDocument: "required", and split each shard key–changing update into delete + upsert in our CDC pipeline by comparing documentKey with the shard key fields in fullDocument. This works, but:

  • every pipeline that keys by documentKey has to implement this split itself, and must handle two different event shapes for the same logical operation;

  • a split done in the pipeline produces two records that share a single resume token. If the token is committed after the delete is applied but before the insert, resuming skips the insert, so the pipeline must also guarantee that both records are committed together;