How Data 360 Segmentation Processes a Quadrillion Records Across Arbitrary Customer Data Models
Data 360’s segmentation engine tackles the massive challenge of processing a quadrillion records monthly across highly customized and varied customer data models. It dynamically interprets arbitrary schemas and relationships at runtime to ensure reliable audience segmentation, powering key marketing and personalization workflows on Salesforce. The platform balances extreme scale and diverse workloads through intelligent resource allocation, observability, and phased query planning to maintain high reliability and cost efficiency. Salesforce teams can learn from its approach to handling metadata scalability, skewed data distributions, and multi-system execution while remaining accountable for SLAs. This provides valuable architectural insights for building robust, scalable data-driven segmentation and activation workflows within complex Salesforce ecosystems.
- Design segmentation engines to dynamically interpret arbitrary customer data models at runtime.
- Implement intelligent workload size estimation and adaptive compute allocation for cost-effective scaling.
- Use phased query planning to handle complex metadata and relationship graphs efficiently.
- Invest in observability, alerting, and automated troubleshooting to manage operational scale.
- Optimize execution strategies to mitigate data skew, duplication, and partitioning issues early.
In our Engineering Energizers Q&A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Deepak Pushpakar, Software Engineering Architect for Segmentation and Activation within Data 360 . His team processes a quadrillion records per month on customer data sets spread across disparate storage systems. The complexity is significant: thousands of tables, thousands of relationships, highly variable data quality, completely custom data models, and data volumes that range from thousands to hundreds of billions of records per job. Explore how his team maintained reliable audience segmentation despite arbitrary customer schemas and relationship graphs across Data 360, and how his team overcame the metadata scalability constraints that threatened query planning, workload execution, and platform usability at extreme scale.