Director-level initiativePortfolio strategyOperating modelProduct strategy

Initiative period: 2020–2025 · measured design-process adoption period: 2023–2024

Enterprise Streaming Portfolio Strategy and Cross-Functional Operating Model

I defined the business needs, requirements, and product boundaries for Walmart’s Spark, Kafka Connect, and Flink platforms, then established shared mechanisms for priorities, investment, adoption, reliability, and risk.

Executive summary

I created scorecards and “when to use” guidance so three streaming platforms addressed distinct tenant problems instead of developing overlapping capabilities. I also aligned engineering, product, program management, infrastructure, security, finance, observability, and tenants around shared decisions and accountability.

3 productsdistinct boundaries for Spark Streaming, Kafka Connect, and Flink
1 → 45engineers participating in design documents across the broader Streaming organization, 2023–2024
≈ 30%defect reduction during the measured design-process adoption period
94%team engagement; belonging was 83% and intent to stay was 100%

My role and scope

  • Defined business needs, requirements, scorecards, and platform-selection guidance.
  • Supplied product leadership before a formal product owner joined in 2021, then served as accountable engineering leader and product partner.
  • Resolved overlap and priority conflicts with product, program, technical leads, tenants, and executives.
  • Connected objectives and feedback to quarterly commitments and engineering milestones.
  • Influenced shared infrastructure roadmaps and developed durable technical ownership.
  • Secured alignment across senior vice president, assistant vice president, director, domain product, Security, Kubernetes, and Observability stakeholders.

The business problem

When decisions begin with technology, streaming products can duplicate functionality, fragment expertise, confuse tenants, and spread limited capacity across competing roadmaps. The portfolio needed a repeatable method to identify each tenant problem, choose the appropriate platform, and share common capabilities rather than rebuild them.

Strategy and product decisions

Spark Streaming served mature Kubernetes-based microbatch processing, ML pipelines, and Medallion and Lambda architectures where minutes of latency were acceptable. General batch processing was outside the product scope.

Kafka Connect served low-latency, lower-cost, lower-complexity integration and CDC. Complex transformations through extensive Single Message Transforms were out of scope. Walmart Global Tech’s Messaging Proxy Service article documents one Kafka Connect sink pattern.

Flink served complex, stateful, low-latency streaming, but carried greater operational and staffing risk and therefore required selective adoption.

Illustrative public scorecard

This reconstructs the decision model for public use; it is not a verbatim copy of Walmart’s internal document.

DimensionSpark StreamingKafka ConnectFlink
Preferred workloadComplex microbatch and ML/data architecturesFocused integration and CDCComplex, stateful, low-latency streaming
LatencyMinutes acceptableLowLowest-latency needs
TransformationHigh complexityLow; avoid complex SMT chainsHigh complexity
Portfolio constraintStreaming, not general batchNot a general processing engineSelective use when staffing and maturity support it

Organizational and cross-functional leadership

Weekly newsletters and recurring reviews exposed adoption, cost, commitments, pages, tickets, vulnerabilities, dependencies, and delivery risk. RFCs and design documents scaled review to the risk of each change; early demos and working groups invited Product, Security, Kubernetes, Observability, UX, and tenants to shape decisions before implementation.

During measured adoption, defects decreased approximately 30%, rework approximately 50%, and delivery success increased approximately 40%. Engineers using the process also received stronger evaluations overall; this is correlation, not proof that documents alone caused the difference.

Execution and operating model

Shared capabilities

The products shared Kubernetes, managed ArgoCD, GitOps, secrets and certificates, observability, vulnerability remediation, disaster recovery, storage, workload tools, and infrastructure upgrades. Managed ArgoCD reached 100% streaming adoption, later expanded to three other Data Platform organizations, and helped Dev Tools establish a supported offering.

Decision rights and accountability

When a proposed control-plane reorganization was declined, I created RACI documentation and required control-plane participation in paging. Control-plane pages initially increased approximately 50% as ownership shifted; data-plane pages declined approximately 20%. Systemic fixes then reduced the control-plane increase to about 10%.

Controlled Flink launch

I challenged the product requirement and the risk of broad availability without enough expertise. The organization delivered production MVP capabilities, launched first pilots in FYE25, onboarded two production pipelines, and limited rollout after an operational incident showed the need for more automation and expertise.

Measured results and people outcomes

  • Reduced overlap and made workload placement, onboarding, roadmap, and support expectations more consistent; no single cost or headcount measure was recorded for this qualitative impact.
  • Four of seven direct reports received above-average evaluations, including one exceptional rating.
  • Every team member used structured 30/60/90-day plans; onboarding included end-user platform use, on-call shadowing, and progress toward PagerDuty participation.

Leadership lessons

Technology choice is not the strategy. Begin with tenant and business outcomes, give each product a distinct purpose, make prioritization transparent, and create operating mechanisms that preserve accountability even when the preferred organization design is not adopted.

Continue the conversation

Review my LinkedIn profile or contact contact@bengesoftwarellc.com.