The queue architecture, multi-carrier failover, DLT compliance layer, and WhatsApp Business API orchestration that lets BookMySMS handle millions of daily messages across SMS, WhatsApp, RCS, and Voice. Without dropping a beat when a carrier goes down or a template gets rejected.
Built 2018 · Continuously operated · 10 min read · Sector: Messaging Infrastructure
Anyone can send a hundred SMS a second. Get to 10,000 and the entire system changes shape. Carrier throttling that never mattered at low volume becomes your bottleneck. TRAI DLT rules that were background noise become gate-keepers on every request. WhatsApp Business API rate limits force you into per-phone-number budgeting. And when one of your seven carriers has a bad hour, "just retry" isn't a strategy. That's how you send the same OTP six times to one user.
BookMySMS started as Decipher's internal messaging tool for our own client projects. When traffic crossed a certain threshold, we spun it out as a standalone product. Today it powers messaging for banks, insurance companies, EdTech platforms, e-commerce brands, and government projects. Sending millions of messages daily across SMS, WhatsApp Business API, RCS, and Voice.
This case study walks through the architectural decisions that got us to sustained 10,000 SMS/second throughput with 99.9% uptime SLA. And what breaks when you scale a messaging platform.
Every SMS in India must resolve to a TRAI-registered DLT template. Get it wrong and the carrier rejects the message. You're charged for the send anyway, and worse, repeated rejections harm your sender reputation across all operators.
Our approach: enforce DLT at the API layer, not at the send layer. When a client submits a message, we resolve:
If any of these fail, the API returns a 400 with a specific error code (e.g., DLT_TEMPLATE_NOT_FOUND, DLT_CONTENT_MISMATCH). The message never enters the queue. This alone cut our carrier-side rejections by ~85% when we moved from post-hoc DLT checking to pre-send validation.
We're integrated with 7+ Indian SMS carriers plus WhatsApp Business API, Twilio for international, and multiple Voice providers. Each has different pricing, throughput limits, operator coverage, and. Critically. Different failure modes.
The routing engine maintains a health score for each carrier, updated every 30 seconds. Score inputs:
When a carrier's score drops below threshold, new messages route to alternatives within one second. When the carrier recovers, we ramp traffic back gradually (not all at once. That just triggers rate limits again).
In early 2024, our largest SMS carrier had a 45-minute regional outage during OTP peak hours (3 AM IST. Banking systems doing overnight batch operations). Their API returned 200 OK for accepted messages but silently dropped 60% of deliveries.
Traditional "retry on failure" would have doubled our send volume trying to recover. Our health scoring caught the drop within 90 seconds (delivery rate crashed from 97% to 38%), reweighted routing to secondary carriers, and drained the queue through them at 70% of original throughput.
The client's monitoring didn't page anyone. They noticed the incident on our weekly delivery report the following Monday. That's what failover done right looks like.
WhatsApp is not SMS. Different rules, different constraints:
BookMySMS abstracts all of this. Client submits a WhatsApp send via the same unified API used for SMS. Internally:
Carrier degradation response. Before health-based routing it took 2-4 hours to detect + mitigate a carrier degradation. With health-based routing: 90 seconds, automatic detection + re-routing.
Carrier-side rejection rate. Before pre-send DLT validation: ~12%. After pre-send DLT validation: under 2%.
Building this system was 20% of the work. Running it is the other 80%. Every day, someone at Decipher is:
This is what managed application operations means in practice. It's not just monitoring a dashboard. It's owning the SLA end-to-end while the client team focuses on their actual business.
Not everyone needs 10,000 SMS/second. But the same architectural patterns. Pre-validation gates, health-based routing, unified API over messy underlying providers, columnar analytics. Apply whenever you're building a platform that sits between real users and multiple upstream providers. Payment gateways, video streaming, IoT ingestion, ad-tech bidding. Different domain, same shape of problem.
If you're building or scaling a system like this and want to skip the 18-month learning curve, reach us via the contact form. We'll tell you honestly whether the problem is worth building from scratch or whether an existing platform will get you 80% there.
RAG on 50,000+ Indian court judgments. Sub-500ms retrieval, near-zero hallucinations.
Read case study CS-03 Managed Ops · E-commerceMTTR from 4.2h to 22min for a D2C brand. The alerting redesign + AI anomaly detection.
Read case study CS-04 Compliance PlatformDPDP + GDPR + ISO 27001 + SOC 2 on one modular platform.
View all case studiesWhether it's messaging, payments, IoT, or another domain where multiple upstream providers meet real user demand. We've done this. Reach us via the contact form to talk architecture.