Executive summary
I created scorecards and “when to use” guidance so three streaming platforms addressed distinct tenant problems instead of developing overlapping capabilities. I also aligned engineering, product, program management, infrastructure, security, finance, observability, and tenants around shared decisions and accountability.
My role and scope
- Defined business needs, requirements, scorecards, and platform-selection guidance.
- Supplied product leadership before a formal product owner joined in 2021, then served as accountable engineering leader and product partner.
- Resolved overlap and priority conflicts with product, program, technical leads, tenants, and executives.
- Connected objectives and feedback to quarterly commitments and engineering milestones.
- Influenced shared infrastructure roadmaps and developed durable technical ownership.
- Secured alignment across senior vice president, assistant vice president, director, domain product, Security, Kubernetes, and Observability stakeholders.
The business problem
When decisions begin with technology, streaming products can duplicate functionality, fragment expertise, confuse tenants, and spread limited capacity across competing roadmaps. The portfolio needed a repeatable method to identify each tenant problem, choose the appropriate platform, and share common capabilities rather than rebuild them.
Strategy and product decisions
Spark Streaming served mature Kubernetes-based microbatch processing, ML pipelines, and Medallion and Lambda architectures where minutes of latency were acceptable. General batch processing was outside the product scope.
Kafka Connect served low-latency, lower-cost, lower-complexity integration and CDC. Complex transformations through extensive Single Message Transforms were out of scope. Walmart Global Tech’s Messaging Proxy Service article documents one Kafka Connect sink pattern.
Flink served complex, stateful, low-latency streaming, but carried greater operational and staffing risk and therefore required selective adoption.
Illustrative public scorecard
This reconstructs the decision model for public use; it is not a verbatim copy of Walmart’s internal document.
| Dimension | Spark Streaming | Kafka Connect | Flink |
|---|---|---|---|
| Preferred workload | Complex microbatch and ML/data architectures | Focused integration and CDC | Complex, stateful, low-latency streaming |
| Latency | Minutes acceptable | Low | Lowest-latency needs |
| Transformation | High complexity | Low; avoid complex SMT chains | High complexity |
| Portfolio constraint | Streaming, not general batch | Not a general processing engine | Selective use when staffing and maturity support it |
Organizational and cross-functional leadership
Weekly newsletters and recurring reviews exposed adoption, cost, commitments, pages, tickets, vulnerabilities, dependencies, and delivery risk. RFCs and design documents scaled review to the risk of each change; early demos and working groups invited Product, Security, Kubernetes, Observability, UX, and tenants to shape decisions before implementation.
During measured adoption, defects decreased approximately 30%, rework approximately 50%, and delivery success increased approximately 40%. Engineers using the process also received stronger evaluations overall; this is correlation, not proof that documents alone caused the difference.
Execution and operating model
Shared capabilities
The products shared Kubernetes, managed ArgoCD, GitOps, secrets and certificates, observability, vulnerability remediation, disaster recovery, storage, workload tools, and infrastructure upgrades. Managed ArgoCD reached 100% streaming adoption, later expanded to three other Data Platform organizations, and helped Dev Tools establish a supported offering.
Decision rights and accountability
When a proposed control-plane reorganization was declined, I created RACI documentation and required control-plane participation in paging. Control-plane pages initially increased approximately 50% as ownership shifted; data-plane pages declined approximately 20%. Systemic fixes then reduced the control-plane increase to about 10%.
Controlled Flink launch
I challenged the product requirement and the risk of broad availability without enough expertise. The organization delivered production MVP capabilities, launched first pilots in FYE25, onboarded two production pipelines, and limited rollout after an operational incident showed the need for more automation and expertise.
Measured results and people outcomes
- Reduced overlap and made workload placement, onboarding, roadmap, and support expectations more consistent; no single cost or headcount measure was recorded for this qualitative impact.
- Four of seven direct reports received above-average evaluations, including one exceptional rating.
- Every team member used structured 30/60/90-day plans; onboarding included end-user platform use, on-call shadowing, and progress toward PagerDuty participation.
Leadership lessons
Technology choice is not the strategy. Begin with tenant and business outcomes, give each product a distinct purpose, make prioritization transparent, and create operating mechanisms that preserve accountability even when the preferred organization design is not adopted.
Continue the conversation
Review my LinkedIn profile or contact contact@bengesoftwarellc.com.