A Delhi-based D2C brand doing ~₹5L/day in GMV was losing an average ₹18 lakh/year to incident downtime. Twelve weeks of alerting redesign, on-call restructuring, and AI-driven anomaly detection later. Steady-state 22-minute MTTR, 3x earlier incident catches, and support tickets dropped 34%.
Engagement: Q2 2025 · Ongoing operations · 10 min read · Sector: D2C / E-commerce
A Delhi-based D2C brand. Anonymized per our client NDA. Running a mid-market beauty and personal care storefront on Shopify Plus with custom checkout, ~₹5 lakh/day in GMV, three warehouses on ShipRocket, three payment gateways, and a WhatsApp order support team of six.
When they signed with us, their incident response looked like this: something breaks. A customer complains on Instagram DM. Support team pings the developer on WhatsApp. Developer opens laptop from wherever they are, tries to figure out what's broken by tailing logs across three servers, eventually finds the issue, deploys a fix, comes back to update the team. Average time from customer report to full resolution: 4.2 hours.
At their GMV, that's roughly ₹87,000 in lost orders per incident. They were having 20-25 incidents per year that qualified as "user-visible". A mix of checkout failures, payment gateway timeouts, ShipRocket sync issues, inventory desync bugs, and image CDN outages. Real total downtime cost: ~₹18 lakh/year, not counting the reputational cost of Instagram customers watching their orders fail in public.
⚠ The real problem wasn't the incidents
Every e-commerce business has incidents. The problem here was that the business was learning about them from customers, not from monitoring. Alerts existed but they were noisy. Slack channel with 200+ alerts/day, most false positives, all ignored. That's alert fatigue, and it's the root cause of most long-MTTR shops.
Before changing anything, we established a baseline nobody wanted to see. We pulled 90 days of incident data from their Slack alert channel, PagerDuty (barely used), Shopify webhook failures, payment gateway retry logs, and cross-referenced against customer support tickets that mentioned errors.
Findings:
Presented this data to the client's founder. Their reaction: "This is worse than I thought. Fix it."
Before adding AI or anything sophisticated, we did unsexy fundamentals. This alone dropped MTTR to ~2 hours by end of Week 4.
We killed 60% of the existing alerts. Threshold-based CPU/memory alerts that never mattered. Duplicate alerts firing from three different sources for the same underlying issue. "Warning" level alerts that had never once corresponded to an actual problem.
What we kept + what we added:
PagerDuty was already licensed but nobody was on rotation. We set up a proper primary + secondary rotation across three of their dev team + our Ops Pod on-call. Escalation: 5-minute ack window, escalates to secondary at 10 min, escalates to Ops Pod lead at 15 min.
Result: acknowledgment time dropped from 41 minutes to under 4 minutes within one rotation cycle.
Faster alerts + faster ack are worth nothing if the person who acks doesn't know what to do. So we spent 4 weeks writing runbooks for the top 15 recurring incident types.
Each runbook followed a strict template:
Runbooks were tested via gameday drills. We simulated three of the top incidents on staging and timed the response. First drill: 45 minutes to resolution. By the fourth drill: 18 minutes.
✓ The runbook insight
Every incident that happens twice deserves a runbook. Every incident that happens three times deserves automation. Every incident that happens five times deserves a fundamental architectural fix. This progression alone eliminated 6 of the top-15 recurring incidents within 3 months.
By Week 8, MTTR was already down to ~45 minutes. This is where AI enters the picture. The remaining gains came from catching incidents before they became customer-visible.
We deployed a simple but effective ML pipeline that watched three streams:
The order-rate anomaly detection alone caught 4 incidents in the first month that would previously have taken 30-60 minutes to detect. In one case, a Shopify script edit caused checkout to fail silently for one specific product SKU. The anomaly detection caught the drop in that SKU's order rate within 4 minutes. Traditional threshold-based monitoring would never have seen it.
The most valuable component turned out to be the simplest. A Python job that embeds new error log messages and compares them against a corpus of known errors. When a completely new error pattern appears, someone investigates within minutes.
This caught: two new third-party API breaking changes, one CDN provider deprecation notice buried in error responses, and a slow-brewing memory leak in a background worker. All three would have grown into customer-visible incidents; all three were fixed before that happened.
✓ The compounding second-order effect
Customer support ticket volume dropped 34% within 3 months. Not because customers were happier per se, but because they weren't opening tickets about broken orders any more. Because the orders weren't breaking, and when they did, the incident was resolved before customers noticed. Fewer angry Instagram DMs. Higher NPS. Better repeat purchase rate.
The pattern here. Audit, alert quality, on-call discipline, runbook-driven response, AI anomaly detection layered on top. Works across most e-commerce businesses at this scale. The specific tools change based on the client's stack, but the sequence doesn't.
This is what we mean by managed application operations. Building AI on top of broken monitoring gets you nothing. Fix the operational foundation first, then use AI as a multiplier. That's the order.
Every mid-market e-commerce business we've worked with had a similar pattern. Decent tooling, no operational discipline, MTTR way too long, teams burned out. The transformation is more about process than technology, though good technology accelerates it.
If your incident response looks anything like the "before" picture above, reach us via the contact form. We'll do a 90-day audit + transformation plan the same way we did here, with concrete milestones and measured outcomes.
RAG on 50,000+ Indian court judgments. Sub-500ms retrieval, near-zero hallucinations.
Read case study CS-02 Scale · Messaging Infra10,000 SMS/second, 99.9% uptime, DLT compliance gate, WhatsApp Business API at scale.
Read case study CS-04 Compliance · BFSIRegulatory change detection and DPDP compliance for Indian BFSI. RBI, SEBI, IRDAI circulars to tickets within hours.
Read case studyChat with Deci or use our contact form. We'll do a quick audit call, tell you what your MTTR probably is right now, and whether it's worth engaging further.